[2601.11516] Building Production-Ready Probes For Gemini
Abstract:Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that activation probes may be a promising misuse mitigation technique, but we identify a key remaining challenge: probes fail to generalize under important production distribution shifts. In particular, we find that the shift from short-context to long-context inputs is difficult for existing probe architectures. We propose several new probe architecture that handle this long-context distribution shift. We evaluate these probes in the cyber-offensive domain, testing their robustness against various production-relevant shifts, including multi-turn conversations, static jailbreaks, and adaptive red teaming. Our results demonstrate that while multimax addresses context length, a combination of architecture choice and training on diverse distributions is required for broad generalization. Additionally, we show that pairing probes with prompted classifiers achieves optimal accuracy at a low cost due to the computational efficiency of probes. These findings have informed the successful deployment of misuse mitigation probes in user-facing instances of Gemini, Google's frontier language model. Finally, we find early positive results using AlphaEvolve to automate improvements in both probe architecture search and adaptive red teaming, showing that automating some AI safety research is already possible.
[2601.11516] Building Production-Ready Probes For Gemini --> Computer Science > Machine Learning arXiv:2601.11516 (cs) [Submitted on 16 Jan 2026 ( v1 ), last revised 10 Feb 2026 (this version, v4)] Title: Building Production-Ready Probes For Gemini Authors: János Kramár , Joshua Engels , Zheng Wang , Bilal Chughtai , Rohin Shah , Neel Nanda , Arthur Conmy View a PDF of the paper titled Building Production-Ready Probes For Gemini, by J\'anos Kram\'ar and 6 other authors View PDF HTML (experimental) Abstract: Frontier language model capabilities are improving rapidly. We thus need stronger mitig
saved by
related reading
- How to build fast, efficient monitors for AI models using probes - Goodfiregoodfire.com
- Coup probes: Catching catastrophes with probes trained off-policy — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2storage.googleapis.com
- gemini_v1_5_report.pdfstorage.googleapis.com
- gpt-4-system-card.pdfcdn.openai.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com