David Africa at MATS: Winter 2027
matsprogram.org · 1,228 words · saved by 1 readers
David Africa
This stream will focus on model motivations and character, open-ended environments, and new forms of misalignment. I work primarily on model motivations, personas, character training, awareness, etc., as well as hunting for new types of misalignment. This area of research right now is very fertile, so there are many good ideas out there, possibly ones I didn't consider, so I'd be happy to supervise something else given that it's (1) well-motivated, (2) ambitiously aims to contribute to understanding LLMs or the path to ASI in some clear way, (3) tractable, and (4) it's within my powers to…
saved by
related reading
- Teaching Claude why \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org