✳flâneur — a map of the web's best reading
Memorizing weak examples can elicit strong behavior out of password-locked models — LessWrong
lesswrong.com · 3,142 words · saved by 1 readers
We’ve recently done some research looking into sandbagging: examining when models can succeed at intentionally producing low-quality outputs despite…
x Memorizing weak examples can elicit strong behavior out of password-locked models — LessWrong AI Capabilities AI Frontpage 61 Memorizing weak examples can elicit strong behavior out of password-locked models by Fabien Roger , ryan_greenblatt 6th Jun 2024 AI Alignment Forum 8 min read 5 61 Ω 33 We’ve recently done some research looking into sandbagging : examining when models can succeed at intentionally producing low-quality outputs despite attempts at fine-tuning them to perform well. One reason why sandbagging could be concerning is because scheming models might try to appear less capable
Explore this link on the map →related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- [Paper] Stress-testing capability elicitation with password-locked models — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- [2406.10209] Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMsar5iv.labs.arxiv.org
- Fuzzing LLMs sometimes makes them reveal their secrets — AI Alignment Forumalignmentforum.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org