flâneur — a map of the web's best reading

PI* 0.6 as a reward model - Claude

claude.ai · 1 words · saved by 1 readers

"Conditioning" here is the trick that lets them keep training the policy with plain supervised flow-matching while still getting an RL-style signal into it. It's worth being precise because the word is doing real work.

Explore this link on the map →

saved by