Debate with Self-Play Best-of-N Optimization — LessWrong
lesswrong.com · saved by 2 readers
Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers a…
Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers a…