Owain Evans, AI Alignment researcher
Owain Evans is an AI Alignment researcher leading a new research group in Berkeley and affiliated with Oxford University. Discover his publications, blog posts, and collaborative opportunities on AI alignment, AGI risk, and related topics.
Owain Evans, AI Alignment researcher Blog posts Papers Video List of Mentees Owain Evans Director at Truthful AI (research group in Berkeley) Affiliate Researcher at CHAI, UC Berkeley Recent papers (May 2026): Negation Neglect: When models fail to learn negations in training . ( tweet , blog ) Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers . ( tweet ) The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious . ( blog ) Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers .
Explore this link on the map →related reading
- Owain Evans - Wikipediaen.wikipedia.org
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Spring 2026 Projects - SPARsparai.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org