flâneur — a map of the web's best reading

Non-determinism in GPT-4 is caused by Sparse MoE - 152334H

152334h.github.io · 2,026 words · saved by 1 readers

It’s well-known at this point that GPT-4/GPT-3.5-turbo is non-deterministic, even at temperature=0.0. This is an odd behavior if you’re used to dense decoder-only models, where temp=0 should imply greedy sampling which should imply full determinism, because the logits for the next token should be a pure function of the input sequence & the model weights.

Non-determinism in GPT-4 is caused by Sparse MoE What the title says 152334H included in Tech August 5, 2023 1701 words 8 minutes Contents It's well-known at this point that GPT-4/GPT-3.5-turbo is non-deterministic, even at temperature=0.0 . This is an odd behavior if you're used to dense decoder-only models, where temp=0 should imply greedy sampling which should imply full determinism, because the logits for the next token should be a pure function of the input sequence & the model weights. When asked about this behaviour at the developer roundtables during OpenAI's World Tour, the responses

Explore this link on the map →

related reading