joke-generator
jokegen.sdan.io · 3,319 words · saved by 1 readers
kimi k2 tuned to generate jokes with rubric RL
In an interview last month, someone asked how I'd train a model on a qualitative reward. I'd been working on a geo - guessing model where the reward is distance in kilometers--quantitative, verifiable. They brought up comedy as a counterexample. If two people disagree on whether something is funny, who's wrong? You can't say either of them is. There's no reward function for funny. I didn't have a good answer at the time. But Tinker recently made it possible to post-train Kimi K2, Moonshot's 1 trillion parameter model. Moonshot themselves used rubric-based RL to boost Kimi's creative writing sc
saved by
related reading
- LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Just get to the point.aimatey.co
- As Rocks May Think | Eric Jangevjang.com
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- AI #24: Week of the Podcast — LessWronglesswrong.com
- AirMore AI - Focused on AI Tech and Products Reviewsairmore.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Thoughts — Jason Weijasonwei.net
- AI #100: Meet the New Boss | Don't Worry About the Vasethezvi.wordpress.com
- What 2026 looks like — AI Alignment Forumalignmentforum.org