flâneur — a map of the web's best reading

Non-Determinism of “Deterministic” LLM Settings

arxiv.org · 6,678 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. LLM (large language model) practitioners commonly notice that outputs can vary for the same inputs under settings expected to be deterministic. Yet the questions of how pervasive this is, and with what impact on results, have not to our knowledge been systematically investigated. We investigate non-determinism in five LLMs configured to be deterministic when applied to eight common tasks in across 10 runs, in both zero-shot and few-shot settings. We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%. In fact, none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings. Sharing preliminary results with insiders

Non-Determinism of “Deterministic” LLM Settings Berk Atil 1 , Sarp Aykent 2 , Alexa Chittams 2 , Lisheng Fu 2 , Rebecca J. Passonneau 1 , Evan Radcliffe 2 , Guru Rajan Rajagopal 2 , Adam Sloan 2 , Tomasz Tudrej 2 , Ferhan Ture 2 , Zhe Wu 2 , Lixinyu Xu 2 , Breck Baldwin 2 1 Penn State University, 2 Comcast AI Technologies Correspondence: {bka5352,rjp49}@psu.edu ; breckbaldwin@gmail.com Berk Atil completed this work during his internship at Comcast AI Technologies Abstract LLM (large language model) practitioners commonly notice that outputs can vary for the same inputs under settings expected

Explore this link on the map →

saved by

related reading