Just get to the point: teaching Thinking Machines' 1T Inkling model dad jokes with GRPO on Tinker
aimatey.co · 4,199 words · saved by 1 readers
A GRPO recipe study on Tinker: seven reward designs, one blind-taste-testing dad, and an eval that got audited as hard as the model.
A field report · Tinker + Inkling · GRPO Teaching the new 1T-parameter Inkling model from Thinking Machines the art of the dad joke, and what it actually learned. One joke setup, three answers A dad-joke setup — what do you call a fake noodle? — answered three ways: the canon (the classic) takes the shortest, straightest path to "Impasta!", the tuned model curves up to "Faux-tie!", and the base model wanders down a long wavy path, trailing off without ever landing the joke. What do you call a fake noodle? Canon (the classic) “Impasta!” PunTune “Faux-tie!” Base model “See, it's this…
saved by
related reading
- Welcome Inkling by Thinking Machineshuggingface.co
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Tinkerthinkingmachines.ai
- As Rocks May Think | Eric Jangevjang.com
- Where the goblins came from | OpenAIopenai.com
- DeepSeek-R1arxiv.org
- joke-generatorjokegen.sdan.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- 2305.07759arxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- What Thinking Machines’ Inkling Is Really Like, Part Isemaphore.substack.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org