Sarah
11 followers · 13 following · 572 views
on the atlas — 117
- Prompts for Open Problems - by Ben Recht - arg min1 savers
- Tell agents the why, not just the how1 savers
- How to Unclench | Jonny Miller17 savers
- How Complex Systems Fail19 savers
- Astra and Fable still hack on simple variants of alignment evals from 2025 — LessWrong5 savers
- Current alignment techniques might be ineffective (and actively bad) in the age of RL — LessWrong4 savers
- Chris Arnade Walks the World | Substack2 savers
- Mini Blog Post 3: Become a person who Actually Does Things — Neel Nanda45 savers
- d/acc: one year later6 savers
- China won’t win the AI race but would it be much worse if it did? — LessWrong2 savers
- Progress · Patrick Collison4 savers
- Lessons from Lyndon Johnson | naml.us3 savers
- The Security Mindset - Schneier on Security1 savers
- exoharness1 savers
- Seeking Stability in the Competition for AI Advantage | RAND4 savers
- SHIFT relies on token-level features to de-bias Bias in Bios probes — LessWrong1 savers
- My favorite media | Andy Arditi1 savers
- Teaching Claude Why9 savers
- Half A Month Of Consolation Writing Advice6 savers
- Matt Levine - Bloomberg Opinion Columnist | Bloomberg1 savers
- Frontier Risk Report (February to March 2026) - METR4 savers
- Fall Semester Announcements - by Ben Recht - arg min2 savers
- A retrospective of AI alignment14 savers
- How to Build a $20 Billion Semiconductor Fab15 savers
- Tinker - by Richard Ngo - Asimov Press2 savers
- Charles Darwin Quotes (Author of The Origin of Species)1 savers
- Using Artificial Intelligence to Augment Human Intelligence15 savers
- Using Artificial Intelligence to Augment Human Intelligence1 savers
- The artist and the machine – Michael Nielsen2 savers
- Yury Polyanskiy1 savers
- Four golden lessons | Nature2 savers
- Reflections on the Passage of Time in Ancient Chinese Poetry1 savers
- Why Japanese companies do so many different things2 savers
- Challenges of AI1 savers
- We Need a New Science of Progress - The Atlantic25 savers
- Statistical Physics for Ambitious Interpretability: A Workshop Retrospective — LessWrong3 savers
- The Most Forbidden Technique — LessWrong3 savers
- Alignment Is Proven To Be Solvable - by SE Gyges3 savers
- Research as a Leisure Activity, Summer 2026 — Ultralight School2 savers
- Are.na22 savers
- The Old World Is Dying: Advice for 2026 graduates42 savers
- Teaching | David Tong1 savers
- Film Study for Research6 savers
- Why books don't work64 savers
- Shtetl-Optimized » Blog Archive » The First Law of Complexodynamics15 savers
- Import AI 431: Technological Optimism and Appropriate Fear2 savers
- On the Tradeoffs of SSMs and Transformers | Goomba Lab8 savers
- If you let AI do your writing, I will come to your house and kill you27 savers
- Consciousness is a Mathematical Pattern: Max Tegmark at TEDxCambridge 2014 (Full Transcript) – The Singju Post1 savers
- Perspective Distortions: Why Normal Cameras Make Faces Look Weird | Aaron Hertzmann’s blog1 savers
- Import AI | Jack Clark | Substack1 savers
- SLT for AI Safety – Jesse Hoogland1 savers
- Towards the explainability of protein language models | Nature Machine Intelligence1 savers
- bv_cvxbook.pdf2 savers
- Rafael Gomez-Bombarelli: The bittersweet lesson of scaling in AI for materials | Bakar Institute of Digital Materials for the Planet1 savers
- sula - Google Search1 savers
- when we cease to understand the world - Google Search1 savers
- Nikhil Prakash | books1 savers
- Lines and Minds: Visual Abstraction in Art, Psychology, and Computer Graphics - SIGGRAPH 20261 savers
- I knew my writing students were using AI. Their confessions led to a powerful teaching moment | AI (artificial intelligence) | The Guardian2 savers
- Curius / Onboarding2621 savers
- Do Thing, Do One Thing49 savers
- No one can teach you to have conviction | benkuhn.net39 savers
- How To Scale Your Model34 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Career Advice That Doesn’t Suck27 savers
- How to write a cold email27 savers
- What 81,000 people want from AI \ Anthropic26 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- No One is Really Working21 savers
- Reality has a surprising amount of detail18 savers
- Tracing the thoughts of a large language model \ Anthropic18 savers
- Can You Just Do Things?—Asterisk16 savers
- Emotional management16 savers
- Modular Manifolds - Thinking Machines Lab16 savers
- trees are harlequins, words are harlequins — the void15 savers
- Your Life is Driven by Network Effects13 savers
- No One is Even Trying | Applied Divinity Studies12 savers
- How to walk through walls - by Henrik Karlsson11 savers
- Capital, AGI, and Human Ambition - The Intelligence Curse10 savers
- Thinking like Transformers10 savers
- Natural Language Autoencoders \ Anthropic9 savers
- Why I don’t think AGI is right around the corner9 savers
- Using Self-Correcting Search to Accelerate Materials Discovery8 savers
- Recommendations for Technical AI Safety Research Directions8 savers
- Psychokinetics | Manav B. Ponnekanti8 savers
- How not to forget what matters - by Henrik Karlsson8 savers
- WEB 4.0: The birth of superintelligent life8 savers
- Useful Vices for Wicked Problems7 savers
- search art with art7 savers
- Best Of Moltbook - by Scott Alexander - Astral Codex Ten7 savers
- Intentionally Designing the Future of AI7 savers
- On Grand Ambitious Theories6 savers
- No, it’s not The Incentives—it’s you – [citation needed]6 savers
- Emotion Concepts and their Function in a Large Language Model6 savers
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data6 savers
- Obvious advice6 savers
- Attention is all we have - David Bessis6 savers
- Teaching Claude why \ Anthropic6 savers
- Focus areas for The Anthropic Institute \ Anthropic5 savers
highlights — 50
Imagine a generative model trained on paintings up until just before the time of the cubists; might it be that by exploring that model it would be possible to discover cubism?
Using Artificial Intelligence to Augment Human Intelligencea means to discover new aesthetics and new representations of reality
The artist and the machine – Michael NielsenIt observes, almost certainly correctly, that imagination and ambition themselves play a large role
We Need a New Science of Progress - The Atlanticshouldering some of the metacognition
Why books don't workWill readers notice if they solved a problem but missed the insights it was supposed to reveal?
Why books don't workReaders must decide which exercises to do and when. Readers must run their own feedback loops: did they clearly understand the ideas involved in the exercise? If not, what should they do next? What should students do if they’re completely stuck?
Why books don't workReaders must run their own feedback loops.
Why books don't workWhat questions should I be asking? How should I summarize what I’m reading?”
Why books don't workTechnologies emerge almost spontaneously when the necessary conditions are in place
Import AI 431: Technological Optimism and Appropriate Fearone question that intrigues me is whether compression is actually fundamental to intelligence.Is it possible that forcing information into a smaller state forces a model to learn more useful patterns and abstractions?
On the Tradeoffs of SSMs and Transformers | Goomba Labmarginal distortion occurs with perfect linear perspective
Perspective Distortions: Why Normal Cameras Make Faces Look Weird | Aaron Hertzmann’s blognatural thermodynamic perspective on deep learning
DSLT 0. Distilling Singular Learning Theory — LessWrongoverindex and overimitate
Language models are weird for the same reason human cultures are weirdno prior knowledge of the concept
Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero | PNASanalyzing the rank of the span of the internal representations of AZ’s and the human’s games.
Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero | PNASHowever, this approach still analyzes through the lens of , a bias that limits what we can discover from .
Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero | PNASteachability (whether the concept is transferable to another AI agent) and novelty (whether the concept contains information not present in human chess games)
Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero | PNASexcavates vectors that represent concepts from AlphaZero’s internal representations using convex optimization
Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero | PNASIgnorance is bliss. Enjoy the permanent underclass, it’s fun down here.
The trillion-dollar perception gap between Silicon Valley and normal people.lets us steer model training
Intentionally Designing the Future of AISo we train a second copy of Claude to work backwards
Natural Language Autoencoders \ Anthropicwe gain sample efficiency by embedding physics structure into a neural state-space model (NSSM)
Learning plasma dynamics and robust rampdown trajectories with predict-first experiments at TCV | Nature Communicationsneural state-space model (NSSM)
Learning plasma dynamics and robust rampdown trajectories with predict-first experiments at TCV | Nature CommunicationsThere remains the important issue of longer-term degradation of the REBCO
Will neutrons compromise the operation of superconducting magnets in a fusion plant? | MIT News | Massachusetts Institute of Technologywe were way too isolated and insulated in fact, from the energy development world.
The race to fusion with Dennis Whyte | MIT Energy Initiativeoh, this is all going to be slow. You can’t go fast.
The race to fusion with Dennis Whyte | MIT Energy InitiativeSo the engineered thing that we make is heat.
The race to fusion with Dennis Whyte | MIT Energy InitiativeWell, cause it can’t go anywhere because it’s contained, it’s confined by the gravity of the sun itself.
The race to fusion with Dennis Whyte | MIT Energy InitiativeThat evolution was the biggest single change in my professional career.
The race to fusion with Dennis Whyte | MIT Energy InitiativeAnd then what was dropped into our lap was really an opportunity to translate our science expertise into the idea of actually making an energy source.
The race to fusion with Dennis Whyte | MIT Energy Initiativeour system favors like inventiveness and innovation
The race to fusion with Dennis Whyte | MIT Energy InitiativeThey’re inherently distributed forms of generation. They’re all over the place. And that necessitates a simultaneous change in the structure of the grid.
The race to fusion with Dennis Whyte | MIT Energy InitiativeThe analogy here was inspired by Yejin Choi’s work on commonsense intelligence — do check out her work if you haven’t already! Also, Matt Mason’s research continues to be a constant source of inspiration (see his blog posts on the fascinating nuances of dexterous manipulation e.g. in clutter).
Generalist - The Dark Matter of Robotics: Physical CommonsenseIf we could somehow distinguish between ‘stable’ and ‘unstable’ explanations then we would know to what extent to trust their corresponding answer distributions.
Cycles of Thought: Measuring LLM Confidence through Stable ExplanationsLikewise, deciding what to say or ask can be viewed through the lens of Bayesian belief updating or information gain, where the goal is to reduce uncertainty about user intent.
Alignment has a Fantasia ProblemChallenge 1. Modeling User Uncertainty
Alignment has a Fantasia ProblemIt focused on having a generic, over-confident, pseudo-profound blogger voice despite having 15,000 words of context from my best blogs
If Anyone Bursts It, Effective Altruism DiesIf we learned that a person responded to x with y, what sort of a person would we think they are?
The Persona Selection Model: Why AI Assistants might Behave like HumansIf we met a person who behaved this way, we’d most likely suspect that they had emotions but were hiding them; we might further conclude that the person is inauthentic or dishonest. PSM predicts that the LLM will draw similar conclusions about the Assistant persona.
The Persona Selection Model: Why AI Assistants might Behave like HumansThen the executives of the companies will face a choice: Abandon the faithful CoT golden era, or fall behind competitors. (Or the secret third option: Coordinate with each other & the government to make sure everyone who matters (all the big players at least) stick to faithful CoT).
The Most Forbidden Technique — LessWrongThis is a coordination problem. And solving it starts with measurement.
How AI Is Learning to Think in Secretfunctional emotions: patterns of expression and behavior modeled after humans under the influence of an emotion, which are mediated by underlying abstract representations of emotion concepts.
Emotion Concepts and their Function in a Large Language ModelThe value-loading problem was framed as perhaps unsolvable in principle, and every discussion of AI risk proceeded from that assumption.
Alignment Is Proven To Be Solvable - by SE GygesIt is easier to know that something isn’t right than it is to figure out what to replace it with
Some relationships deepen when you tell the truth and some endBut with groups, there was a whole set of social dynamics that made people turn on whoever challenged the status quo.
Some relationships deepen when you tell the truth and some end1 Rodriguez’ diary, Rebel Without a Crew, is one example. “The Story of VaccinateCA” is another. I also like “Playing to Win” by Alice Maz. A Guide to the Perplexed with Werner Herzog. Surely, you're joking Mr. Feynman. Robert Caro’s The Years of Lyndon Johnson.
How to walk through walls - by Henrik KarlssonMy boss said that if you are someone who is already creative, and then you become technical, then you are unstoppable.
How to walk through walls - by Henrik KarlssonThe idea is to experience creating your own images and/or stories no matter how crude they are and then manipulating them through editing.
How to walk through walls - by Henrik KarlssonI generally advise people working on wicked problems to aim for "jogging" rather than "sprinting" - a metaphor I like because it emphasizes that this is fully consistent with trying to finish as fast as possible.
Useful Vices for Wicked ProblemsStarting [at a young age] he’s read everything that he could find about business. The subject that interests him, he’s read newspapers, biographies, trade press. He went over to his grandfather who was a grocer and he read the progressive grocer magazine, and he read articles on how to stock a meat department... What he’s really done is he’s created this immense vertical filing cabinet in his brain of layers and layers and layers of files of information that he can draw back on now for more than 70 years worth of data.
Curius / Onboarding