flâneur — a map of the web's best reading

joke-generator

jokegen.sdan.io · 3,319 words · saved by 1 readers

kimi k2 tuned to generate jokes with rubric RL

In an interview last month, someone asked how I'd train a model on a qualitative reward. I'd been working on a geo - guessing model where the reward is distance in kilometers--quantitative, verifiable. They brought up comedy as a counterexample. If two people disagree on whether something is funny, who's wrong? You can't say either of them is. There's no reward function for funny. I didn't have a good answer at the time. But Tinker recently made it possible to post-train Kimi K2, Moonshot's 1 trillion parameter model. Moonshot themselves used rubric-based RL to boost Kimi's creative writing sc

Explore this link on the map →

saved by

related reading