SPAR Research Library
A collection of 145 research reports from the Supervised Program for Alignment Research (SPAR), covering AI safety, AI policy, interpretability, and more.
Isha Harris Mentored by Gary Abel Biological AI models are powerful tools for detecting biological threats, but it remains unclear whether their predictions generalize to highly non-natural proteins. This is increasingly important as AI-enabled biodesign produces novel sequences that may not be reliably captured by direct sequence-comparison approaches used in DNA synthesis screening. Here, we use sparse autoencoder (SAE) features from protein language model activations to test whether interpretable internal representations can identify structural similarity across generated proteins. We…
saved by
related reading
- Spring 2026 Projects - SPARsparai.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- ARENA - AI Safety Curriculumlearn.arena.education
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- AI Safety | Arkosevictoriabrook.github.io
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org