Research | AI Alignment Foundation
aialignmentfoundation.org · 181 words · saved by 1 readers
What we're funding and accelerating to solve alignment.
What we're funding and accelerating to solve alignment. Read more about Modular Pretraining Enables Access Control Modular Pretraining Enables Access Control Dual-use knowledge enables models to assist us with the most difficult and demanding tasks in science, but it also empowers people who would use that knowledge to cause harm. Pre-training with GRAM enables knowledge to be siloed and turned on or off when deployed, so that a single model can be both safe and powerful. Read more about Self-Interpretation in Language Models via Adapter Probes Self-Interpretation in Language Models via…
saved by
related reading
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org