✳flâneur — a map of the web's best reading
Asher P
56 followers · 28 following · 2006 views
Open this reading profile →
on the atlas — 92
- The case for ensuring that powerful AIs are controlled — LessWrong9 savers
- What is wisdom?1 savers
- Can activation verbalizers surface an internal chain of thought? — LessWrong4 savers
- Plans A, B, C, and D for misalignment risk — LessWrong5 savers
- Tensor-Transformer Variants are Surprisingly Performant — LessWrong3 savers
- What will GPT-2030 look like? — AI Alignment Forum2 savers
- Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong1 savers
- DSLT 0. Distilling Singular Learning Theory — LessWrong1 savers
- Deriving Muon15 savers
- EDT with updating double counts – The sideways view1 savers
- Spaced Repetition for Efficient Learning · Gwern.net9 savers
- Physics of Language Models1 savers
- Investigating the learning coefficient of modular addition: hackathon project — LessWrong1 savers
- Using Interpretability to Identify a Novel Class of Alzheimer's Biomarkers8 savers
- Self-exfiltration is a key dangerous capability2 savers
- Please don't throw your mind away — LessWrong24 savers
- Features as Rewards: Using Interpretability to Reduce Hallucinations2 savers
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong3 savers
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWrong2 savers
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong1 savers
- Ping pong computation in superposition — LessWrong1 savers
- Irrationality as a Defense Mechanism for Reward-hacking — LessWrong1 savers
- AlgZoo.pdf - Google Drive1 savers
- Insights on Crosscoder Model Diffing3 savers
- Modular Manifolds - Thinking Machines Lab15 savers
- ARC progress update: Competing with sampling — LessWrong2 savers
- I am worried about near-term non-LLM AI developments — LessWrong2 savers
- Reward is not the optimization target — LessWrong6 savers
- Tips and Code for Empirical Research Workflows — AI Alignment Forum2 savers
- 2310.014051 savers
- The Toxoplasma Of Rage | Slate Star Codex3 savers
- A gentle introduction to mechanistic anomaly detection — LessWrong2 savers
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum6 savers
- Formal verification, heuristic explanations and surprise accounting — Alignment Research Center1 savers
- Nate Soares' Life Advice - LessWrong1 savers
- Productivity - Sam Altman52 savers
- How to Do Great Work51 savers
- An Opinionated Guide to ML Research50 savers
- Reality has a surprising amount of detail49 savers
- Introduction - SITUATIONAL AWARENESS: The Decade Ahead44 savers
- 95%-ile isn't that good43 savers
- escaping flatland: career advice for CS undergrads41 savers
- Mini Blog Post 3: Become a person who Actually Does Things — Neel Nanda38 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab37 savers
- A Mathematical Framework for Transformer Circuits36 savers
- On the Biology of a Large Language Model31 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans28 savers
- How to be More Agentic - by Cate Hall - Useful Fictions26 savers
- How To Scale Your Model26 savers
- On-Policy Distillation - Thinking Machines Lab26 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning23 savers
- OpenAI Email Archives (from Musk v. Altman) — LessWrong19 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models18 savers
- Verbalizable Representations Form a Global Workspace in Language Models17 savers
- Building the heap: racking 30 petabytes of hard drives for pretraining | blog17 savers
- Dario Amodei — On DeepSeek and Export Controls16 savers
- AGI Ruin: A List of Lethalities - LessWrong15 savers
- Automated Weak-to-Strong Researcher15 savers
- Nadia Asparouhova | How to do the jhanas15 savers
- Transformer Circuits Thread14 savers
- Alignment remains a hard, unsolved problem — LessWrong14 savers
- Six (and a half) intuitions for KL divergence - LessWrong14 savers
- trees are harlequins, words are harlequins — the void14 savers
- DeepSeek-R112 savers
- On neural scaling and the quanta hypothesis10 savers
- Shtetl-Optimized » Blog Archive » The First Law of Complexodynamics10 savers
- The behavioral selection model for predicting AI motivations — LessWrong10 savers
- Current AIs seem pretty misaligned to me — LessWrong9 savers
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcast9 savers
- Language Models, World Models, and Human Model-Building8 savers
- Using Self-Correcting Search to Accelerate Materials Discovery8 savers
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESS8 savers
- Best Of Moltbook - by Scott Alexander - Astral Codex Ten7 savers
- Understanding Variational Autoencoders (VAEs) | by Joseph Rocca | Towards Data Science7 savers
- IIIa. Racing to the Trillion-Dollar Cluster - SITUATIONAL AWARENESS6 savers
- Attribution Patching: Activation Patching At Industrial Scale — Neel Nanda6 savers
- Some Math behind Neural Tangent Kernel | Lil'Log5 savers
- Essays on Reducing Suffering5 savers
- From fear to excitement — LessWrong5 savers
- Understanding Memorization via Loss Curvature4 savers
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWrong4 savers
- Judgments often smuggle in implicit standards — LessWrong3 savers
- A Nihilist’s Guide to Meaning | Melting Asphalt3 savers
- We learn long-lasting strategies to protect ourselves from danger and rejection — LessWrong3 savers
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forum3 savers
- Conflicts between emotional schemas often involve internal coercion — LessWrong3 savers
- Prediction, Explanation, or Over-interpretation?2 savers
- Optimality is the tiger, and agents are its teeth - LessWrong2 savers
- Our Mission, Technology, and Approach2 savers
- How well do truth probes generalise? — LessWrong2 savers
- Matryoshka Sparse Autoencoders — LessWrong2 savers
- Trust develops gradually via making bids and setting boundaries — LessWrong2 savers