Simplex - building our research team | Manifund
This funding proposal outlines a novel approach to AI interpretability and safety based on Computational Mechanics, a mathematical framework from physics. The key points are: Current interpretability methods lack a principled grounding, and are insufficient to understand, control, and monitor AI The proposed approach captures the dynamic structure of computation, adapting Computational Mechanics to make precise predictions about the internal geometry of neural networks that underlie AI system behavioral capabilities. In this way we are able to make a principled connection between the structure of training data, model internals, and model behavior. Our initial results show the approach can predict and empirically verify non-trivial, fractal-like structures in transformer activations, validating its relevance in transformers. Simplex is a new research organization that aims to develop a principled science of AI cognitive capabilities and associated risks. We have a clear and ambitious th
Simplex - building our research team | Manifund 6 Simplex - building our research team Science & technology Technical AI safety Adam Shai Not funded Grant $0 raised Project summary This funding proposal outlines a novel approach to AI interpretability and safety based on Computational Mechanics, a mathematical framework from physics. The key points are: Current interpretability methods lack a principled grounding, and are insufficient to understand, control, and monitor AI The proposed approach captures the dynamic structure of computation , adapting Computational Mechanics to make precise pre
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Transformer Circuits Threadtransformer-circuits.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org