✳flâneur — a map of the web's best reading
joke-generator
jokegen.sdan.io · 3,319 words · saved by 1 readers
kimi k2 tuned to generate jokes with rubric RL
In an interview last month, someone asked how I'd train a model on a qualitative reward. I'd been working on a geo - guessing model where the reward is distance in kilometers--quantitative, verifiable. They brought up comedy as a counterexample. If two people disagree on whether something is funny, who's wrong? You can't say either of them is. There's no reward function for funny. I didn't have a good answer at the time. But Tinker recently made it possible to post-train Kimi K2, Moonshot's 1 trillion parameter model. Moonshot themselves used rubric-based RL to boost Kimi's creative writing sc
Explore this link on the map →saved by
related reading
- LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- AI #24: Week of the Podcast — LessWronglesswrong.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- AI #100: Meet the New Boss | Don't Worry About the Vasethezvi.wordpress.com
- trees are harlequins, words are harlequins - hydrogen jukeboxes: on the crammed poetics of...nostalgebraist.tumblr.com
- AI #97: 4 - by Zvi Mowshowitz - Don't Worry About the Vasethezvi.substack.com
- Google Geminigemini.google.com
- AI is Making You Dumber. Here’s Why – Tập Đọctapdoc.blog
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com