flâneur

SOTA alignment assessments don’t strongly update us against misalignment

blog.redwoodresearch.org · 8,987 words · saved by 2 readers

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]). While I agree with the report on the above bottom-line conclusions (substantially on priors)[4], I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular,…

saved by

related reading