Non-Determinism of “Deterministic” LLM Settings
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. LLM (large language model) practitioners commonly notice that outputs can vary for the same inputs under settings expected to be deterministic. Yet the questions of how pervasive this is, and with what impact on results, have not to our knowledge been systematically investigated. We investigate non-determinism in five LLMs configured to be deterministic when applied to eight common tasks in across 10 runs, in both zero-shot and few-shot settings. We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%. In fact, none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings. Sharing preliminary results with insiders
Non-Determinism of “Deterministic” LLM Settings Berk Atil 1 , Sarp Aykent 2 , Alexa Chittams 2 , Lisheng Fu 2 , Rebecca J. Passonneau 1 , Evan Radcliffe 2 , Guru Rajan Rajagopal 2 , Adam Sloan 2 , Tomasz Tudrej 2 , Ferhan Ture 2 , Zhe Wu 2 , Lixinyu Xu 2 , Breck Baldwin 2 1 Penn State University, 2 Comcast AI Technologies Correspondence: {bka5352,rjp49}@psu.edu ; breckbaldwin@gmail.com Berk Atil completed this work during his internship at Comcast AI Technologies Abstract LLM (large language model) practitioners commonly notice that outputs can vary for the same inputs under settings expected
Explore this link on the map →saved by
related reading
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- [2506.09501] Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inferencearxiv.org
- [2506.09501] Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inferencearxiv.org
- gpt-4.pdfcdn.openai.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Things we learned about LLMs in 2024simonwillison.net
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com
- GenAI Handbookgenai-handbook.github.io
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com
- Non-determinism in GPT-4 is caused by Sparse MoE - 152334H152334h.github.io
- The bitter lesson of LLM evalsparsed.com
- 2025: The year in LLMssimonwillison.net