✳flâneur — a map of the web's best reading
PI* 0.6 as a reward model - Claude
claude.ai · 1 words · saved by 1 readers
"Conditioning" here is the trick that lets them keep training the policy with plain supervised flow-matching while still getting an RL-style signal into it. It's worth being precise because the word is doing real work.
Explore this link on the map →