Elizabeth Qiu
☀️ ↙ Back humans that do not scale Some human qualities will never scale, no matter how powerful technology gets. As AI capabilities compound, the human remains stubbornly fixed. There is much exciting technology going on, like the work in biology (neuroscience, longevity, etc). We’re seeing signs that we can read people’s minds and telepathically communicate. Regardless, some laws of biology stil
The convergence of AI capabilities and infrastructure needs creates unprecedented opportunities for system-level improvements Multimodal AI represents a paradigm shift from siloed sensors to integrated intelligence across infrastructure domains Critical infrastructure resilience requires moving from reactive repair to predictive prevention through continuous monitoring Superposition allows models to represent exponentially many features in linear dimensions—a blessing and curse for interpretability Polysemanticity: when individual neurons respond to multiple unrelated concepts, making direct i
Explore this link on the map →saved by
related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- On neural scaling and the quanta hypothesisericjmichaud.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai
- On Optimism for Interpretabilitygoodfire.ai