[2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervision
Abstract:As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears acceptable but falls short of its true abilities. Can training elicit a model's best work even without reliable verification? We study this using model organisms trained to sandbag, testing elicitation techniques on problem-solving math, graduate-level science, and competitive coding tasks. We find that training with weak supervision can reliably elicit sandbagging models when supervised fine-tuning (SFT) and reinforcement learning (RL) are combined: SFT on weak demonstrations breaks the sandbagging behavior, enabling RL to then fully elicit performance. Neither method succeeds reliably alone-RL without SFT almost always leads to reward hacking rather than genuine improvement. Critically, this relies on training being indistinguishable from deployment; when models can distinguish between training and deployment, they can perform well during training while continuing to sandbag afterward. Our results provide initial evidence that training is a viable mitigation against sandbagging, while highlighting the importance of making training indistinguishable from deployment.
[2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervision --> Computer Science > Machine Learning arXiv:2604.22082 (cs) [Submitted on 23 Apr 2026 ( v1 ), last revised 1 May 2026 (this version, v2)] Title: Removing Sandbagging in LLMs by Training with Weak Supervision Authors: Emil Ryd , Henning Bartsch , Julian Stastny , Joe Benton , Vivek Hebbar View a PDF of the paper titled Removing Sandbagging in LLMs by Training with Weak Supervision, by Emil Ryd and 4 other authors View PDF HTML (experimental) Abstract: As AI systems begin to automate complex tasks, supervision increasi
Explore this link on the map →saved by
related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingblog.redwoodresearch.org
- [2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervisionar5iv.labs.arxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [Paper] Stress-testing capability elicitation with password-locked models — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com