flâneur

Dense, on-policy, or both?

baseten.co · 2,441 words · saved by 2 readers

Constitutional alignment as a testbed for comparing learning signals in SFT, RL, and everything in between

Introduction Post-training methods can be usefully organised along two axes: whether the model trains on its own generations (on-policy) or someone else's (off-policy), and whether the learning signal is dense (token-level) or sparse (sequence-level). Much of our recent work has tried to populate the space between off-policy SFT and on-policy RL. Iterative SFT (iSFT), which converts supervised data into gold-standard samples via a grader- and refiner-loop, moves SFT closer to the current policy by repairing model generations; RL optimizes directly on policy but with scalar rewards; and…

saved by

related reading