flâneur — a map of the web's best reading

Memorizing weak examples can elicit strong behavior out of password-locked models — LessWrong

lesswrong.com · 3,142 words · saved by 1 readers

We’ve recently done some research looking into sandbagging: examining when models can succeed at intentionally producing low-quality outputs despite…

x Memorizing weak examples can elicit strong behavior out of password-locked models — LessWrong AI Capabilities AI Frontpage 61 Memorizing weak examples can elicit strong behavior out of password-locked models by Fabien Roger , ryan_greenblatt 6th Jun 2024 AI Alignment Forum 8 min read 5 61 Ω 33 We’ve recently done some research looking into sandbagging : examining when models can succeed at intentionally producing low-quality outputs despite attempts at fine-tuning them to perform well. One reason why sandbagging could be concerning is because scheming models might try to appear less capable

Explore this link on the map →

related reading