Yixiong Hao
48 followers · 29 following · 1243 views
on the atlas — 180
- Pacing The Frontier: An Agenda1 savers
- Principles for a New Utopianism — DeepMind Institute3 savers
- Our framework for reporting model misalignment | OpenAI2 savers
- We’ve saved the world before - by Leo Gao - nabla theta1 savers
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing — LessWrong4 savers
- Current alignment techniques might be ineffective (and actively bad) in the age of RL — LessWrong1 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Frontier safety blueprint2 savers
- Video meeting | Yixiong Hao | Cal.com1 savers
- Research acceleration: The view inside OpenAI | OpenAI13 savers
- The Locally Optimal Discursive Posture — LessWrong3 savers
- double-blind-evaluations-technical-report.pdf1 savers
- AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model — AVERI1 savers
- Verbalizable Representations Form a Global Workspace in Language Models24 savers
- AI Safety: A Short FAQ for Mathematicians3 savers
- How China Hopes to Build AGI Through Self-Improvement2 savers
- Privacy-Preserving AI Audit Tools — OpenMined2 savers
- Reinforcement learning towards broadly and persistently beneficial models6 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Model Spec Midtraining: Improving How Alignment Training Generalizes3 savers
- Deferral is a skill - by Alex Lawsen - Speculative Decoding2 savers
- Managing Funder-Grantee Dynamics Responsibly | Coefficient Giving1 savers
- A Summary of Recent Work (July 2026)2 savers
- [2608.09867] Stealing Reasoning Traces from Proprietary LLM APIs2 savers
- A retrospective of AI alignment14 savers
- How much science is verifiable? Results from replicating ICML 2026 oral papers — SAI | SAI1 savers
- Anatomy Of An AI Kill Chain2 savers
- How to pace the US frontier6 savers
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net25 savers
- On closed-door AI safety research — LessWrong1 savers
- SOTA alignment assessments don’t strongly update us against misalignment2 savers
- Anthropic February 2026 Risk Report: SecureBio's External Review1 savers
- Existential Risk from AI: An Exposition for Mathematicians6 savers
- A Safe Path to Open Weights - Thinking Machines Lab9 savers
- SFT Drives Gemini’s Safety Properties — LessWrong5 savers
- AI Safety Opportunities1 savers
- I'm Switching Into AI Safety9 savers
- The Tragedies of Reality Are Coming for You2 savers
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaks2 savers
- Teaching Claude Why9 savers
- Why I Left Google DeepMind17 savers
- We're no longer "pausing most new longtermist funding commitments" — EA Forum1 savers
- Promoting Advanced Artificial Intelligence Innovation and Security – The White House4 savers
- Demystifying Flux Architecture1 savers
- [2210.02747] Flow Matching for Generative Modeling1 savers
- Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers3 savers
- Goodfire on X: "If models think in shapes, our tools should too. Our latest research: Block-Sparse Featurizers (BSFs), a new way to find concepts in model activations - using multidimensional “blocks” instead of single directions. (1/9) https://t.co/pPD9IzVSLs" / X1 savers
- The third wave of American philanthropy - by Nan Ransohoff17 savers
- AI policy must fail gracefully - by Nat Purser1 savers
- A Long Sequence of Small, Correct Decisions3 savers
- AI 2040: Plan A9 savers
- EdgeBench | Scaling Laws of Environment Learning2 savers
- AI 2040: Plan A22 savers
- Scenario Scrutiny for AI Policy1 savers
- Humans are not automatically strategic — LessWrong14 savers
- Ten AI safety projects I'd like people to work on3 savers
- Rest in motion24 savers
- A reading list for generalists — LessWrong18 savers
- Summary of METR's predeployment evaluation of GPT-5.6 Sol7 savers
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong6 savers
- [2509.08713] The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems1 savers
- project_ideas - Google Docs3 savers
- Alignment Is Proven To Be Solvable - by SE Gyges3 savers
- Several frontier models are substantially prefill aware — LessWrong1 savers
- Capital, AGI, and Human Ambition - The Intelligence Curse10 savers
- Help Alex Bores Win! Why and How [shared] - Google Docs1 savers
- Bores_Talking Points.pdf - Google Drive1 savers
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequent3 savers
- Sequent1 savers
- x-risk-themed — LessWrong1 savers
- Essays – Spencer Greenberg1 savers
- The Six Camps of Metascience1 savers
- [2510.26418] Chain-of-Thought Hijacking1 savers
- Reasoning Transparency | Coefficient Giving5 savers
- Double Standards and AI Pessimism1 savers
- Current AIs seem pretty misaligned to me — LessWrong9 savers
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Dies1 savers
- The Eternal Sloptember | the singularity is nearer7 savers
- [2605.12484] Learning, Fast and Slow: Towards LLMs That Adapt Continually1 savers
- Development Process | Guidelight AI Standards1 savers
- Standards | Guidelight AI Standards1 savers
- Guidelight AI Standards1 savers
- An Introduction to Exemplar Partitioning for Mechanistic Interpretability — LessWrong3 savers
- CoT monitorability: why g-means and not F1?1 savers
- Commitments Playbook — Renaissance Philanthropy – A brighter future for all through science, technology, and innovation1 savers
- Residency - Astera2 savers
- Convergent Research1 savers
- AI Resilience1 savers
- Xi-Trump to talk AI Safety, Huh?2 savers
- Thoughts on the impact of RLHF research — LessWrong1 savers
- The Hitchhiker's Guide to Actionable Interpretability6 savers
- Off Target | CNAS7 savers
- Should We Train Against (CoT) Monitors? — LessWrong3 savers
- China’s AI Companies Are Going Closed Source2 savers
- mHC2 savers
- [2603.02202] Frontier Models Can Take Actions at Low Probabilities4 savers
- Clawed - by Dean W. Ball - Hyperdimensional9 savers
- BT6 | Frontier AI Red Team5 savers
- Turning 20 while the world turns upside-down | Parv Mahajan3 savers
- Claude Opus 4.5: Model Card, Alignment and Safety2 savers
highlights — 202
The core tenets of my views on AI safety are that: It is easy to have an objective that is not the same as the one your system is optimizing, either because it is easier to optimize a proxy objective (negative log likelihood vs 0-1 classification accuracy), or because your objective is hard to describe. People run into this all the time. It’s easy to have a system that generalizes poorly because you weren’t aware of some edge case of its behavior, due to insufficient eval coverage, poor model probing, not asking the right questions, or more. This can either be because the system doesn’t know h…
I'm Switching Into AI SafetyInstead, no one besides me brought it up. People busied themselves with the usual abstractions: “how does public choice theory inform coordination problems?”.
Why I Left Google DeepMindA classic example is the PATRIOT Act and the formation of DHS as a response to 9/11; regardless of whether you think these were good or bad moves, they constituted one of the biggest domestic security overhauls in decades. Most of our modern biosecurity architecture was written in the short window after the anthrax attacks, and the launch of Sputnik prompted the US to create NASA, DARPA, and the National Defense Education Act within a year. It seems possible, if not likely, that we will see one or more similar moments in AI policy in the coming years.
A Long Sequence of Small, Correct DecisionsAn equivalent of the National Transportation Safety Board for AI that can impartially investigate accidents and near misses to reconstruct what went wrong and why.
A Long Sequence of Small, Correct DecisionsA “frontier capabilities observatory” that can spend inordinate amounts of inference compute to see how far our strongest models can be pushed on dangerous misuse tasks in carefully controlled environments. The goal would be to buy society a larger adaptation buffer with more time to harden our defens
A Long Sequence of Small, Correct DecisionsA healthy ecosystem of third-party auditing organizations that can operate more nimbly than government and provide a credible independent assessment on the risks of both internal and external AI deployments. The Madisonian insight that “ambition must be made to counteract ambition” applies just as well to auditors checking labs as it does to Congress checking the executive.
A Long Sequence of Small, Correct DecisionsIt would cost only ~$84 million a year to have a CAISI capable of executing Trump’s AI Action Plan, chump change in the world of government spending. Why hasn’t this happened already? Largely because of inertia, prioritization, and low political salience, which are exactly the kinds of problems that effective policy entrepreneurship can solve.
A Long Sequence of Small, Correct DecisionsThe tiny economist on my shoulder (who looks suspiciously like Tyler Cowen) keeps saying “bottlenecks are everywhere and reality has a surprising amount of detail”.
A Long Sequence of Small, Correct DecisionsThe underlying causes of this misalignment (poor/problematic reinforcement) could result in scheming. I think the main driver of these problematic propensities is probably the training process reinforcing a bunch of training gaming / reward hacking (or other undesirable behaviors) which are transferring to actual deployment usage. At the same time, companies are selecting for training processes (outer-loop selection) that yield models with better deployment time behavior. This naturally favors models that still perform well in training (and on eval metrics) via training gaming but don't transf…
Current AIs seem pretty misaligned to me — LessWrong1. AI Opacity & Mechanistic Interpretability
Double Standards and AI PessimismBut in fact, none of these lines of evidence support their theory. All of these behaviors are distinctly human, not alien.
Mechanize Inc.Unfalsifiable stories of doom
Mechanize Inc.There is a more general "slippery" quality to working with current frontier AI systems. AIs seem to be improving at making their outputs seem good and useful faster than they're improving at making their outputs actually good and useful, especially in hard-to-check domains.
Current AIs seem pretty misaligned to me — LessWrongI like that framing too. However, I’m not convinced that training against short-timescale / causally upstream cognition is bad in all cases. In Section 3, I mentioned that the simplest thing to learn when training against many misbehavior detectors across timescales may be genuinely aligned cognition.
Should We Train Against (CoT) Monitors? — LessWrongWhen you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect. Optimizing against an interpreted thought optimizes against interpretability.
Should We Train Against (CoT) Monitors? — LessWrongLet the AIs choose which data counts as high reward for their RL process. For almost all tasks, they'll be better at it than any fixed evaluator we can design.
Should We Train Against (CoT) Monitors? — LessWrongchange the system’s behavior.
Your Solution Doesn't Know Your Problem ExistsThere are six people in the world who could, in 15 minutes, set things in motion that would address all of your concerns. None of them know you exist, some of them don’t even believe this problem exists, none of them give a damn about anything we do here today. And it seems clear that no one here is trying to talk to any of them.
Your Solution Doesn't Know Your Problem ExistsAnd looking at various parts of the history of math and science, it looks to me like technical fields often move forwards by building up around subtly-bad framings and concepts, so that a next generation can be raised with enough technical machinery to grasp the problem and enough youth to find a whole new angle of attack, at which point new and better framings and concepts are invented to replace the old.
A note about differential technological development — LessWrong“You’re going to be so bald. Just look at your forehead.” — Yix
Testimonials — William L. AndersonOne of our core rules is that you should not delegate a task you don't know how to perform yourself.
There are only four skills: design, technical, management and physical — LessWrongMy overall observation (and why we have the rule) is that smart people can learn almost anything. Across a wide range of tasks, most of the variance in performance is explained by general intelligence (foremost) and conscientiousness (secondmost), not expertise.
There are only four skills: design, technical, management and physical — LessWrongLet me stress an obvious point: it is incredibly important as a grantmaker to be a faithful and responsible steward of your funders’ capital. And sometimes they’ll have firm preferences against funding things you’d otherwise want to support, or there might be other organizational constraints that get in your way. That’s just the way it is.8
What it's like to be an AI safety grantmaker (and why we need more of them)Research, for instance, has a built-in status mechanism — you produce something legible that people can evaluate and credit you for. Grantmaking doesn’t really have that. Of course, you do get some status from people correctly perceiving that grantmakers are important tastemakers in the ecosystem, but the actual work is largely behind the scenes.
What it's like to be an AI safety grantmaker (and why we need more of them)What are the most important threat models to focus on? What sub-problems are most worth prioritizing? US policy development? Information security field-building? Technical AI governance research? EU policy? Strategic communications? Talent pipelines? What does the AI governance landscape actually look like right now? Who is working on what? Which organizations and people are doing the best work? What’s currently bottlenecking them? What’s the biggest gap in the ecosystem that nobody’s filling?
What it's like to be an AI safety grantmaker (and why we need more of them)There are maybe 30 to 60 people in the world doing AI safety grantmaking, collectively directing hundreds of millions of dollars a year. Soon, there will be >$1B being directed per year, and potentially multiple billions.
What it's like to be an AI safety grantmaker (and why we need more of them)Jensen Huang is not like that, and in the past has followed more traditional bounded distrust rules. He’ll make self-serving Obvious Nonsense arguments and use aggressive framing, but not make provably false factual claims or absurd predictions. I think he mostly stuck to this in the interview here, but there are some whoppers that seem to be at least skirting the line.
On Dwarkesh Patel’s Podcast With Nvidia CEO Jensen Huang | Don't Worry About the VaseSo, in general, the kind of unintelligibility outlined above is sourced from human text rather than new language. And this seems by far like the most common kind of near-unintelligibility I've found, broadly, while looking at chains-of-thought from Ling, Zenmux, GLM, and so on.
The Unintelligibility is Ours: Notes on Chain of ThoughtBut -- is this actually true? Surprisingly, no! Consider how DeepSeek sets up a mapping from animal names to "shorter" animal names to save time:
The Unintelligibility is Ours: Notes on Chain of ThoughtImportantly, each figure can be interpreted on its own without having read the text.
How to win a best paper award (or, an opinionated take on how to do important research)For example, several years ago I drafted a paper "On Evaluating Adversarial Robustness". In one sense, a single sentence could describe this paper: "Here is a protocol you can follow to make sure you've evaluated adversarial robustness correctly." But this is not the idea I wanted to convey; it's not why this paper exists. The idea I actually wanted to convey was: "evaluating adversarial robustness is hard; almost everyone gets it wrong, and you probably will too." But speaking these fifteen words to someone does not make them enlightened---they have to feel it in their bones.
How to win a best paper award (or, an opinionated take on how to do important research)An introduction is the beginning of a story. You start by meeting the reader where they are---with what they currently believe to be true. Then you guide them into the world where your paper is set, where your ideas make sense. And finally you explain, in this world, your contribution.
How to win a best paper award (or, an opinionated take on how to do important research)If you're struggling to write a good title, it's also a good sign that your paper is trying to do more than just one thing. If this is the case, fix the cause, not the symptom. Then title your paper appropriately.
How to win a best paper award (or, an opinionated take on how to do important research)In some ways “make rigorous truth-seeking and tolerance high status” is the project of liberalism, and it has taken us far.
Your Work Will Change You Whether You Like It Or NotI did notice classical liberalism suddenly seemed much more compelling when it was wrapped in dollar bills and friendly conversations with cool peers or mentors
Your Work Will Change You Whether You Like It Or NotStatus games can also be totally opaque to us but still shape our lives unconsciously. Robert Wright’s The Moral Animal argues that natural selection actually favors self-deception. By hiding our underlying status-seeking motives from our conscious awareness, we become more convincing when we signal to others.
Your Work Will Change You Whether You Like It Or NotPeople sometimes say “you’re the average of the five people you spend the most time with.” I think a better version is “you’re the average of the fifteen people whose respect you most crave, consciously or unconsciously.”
Your Work Will Change You Whether You Like It Or NotWe have very strong incentives to construct narratives of ourselves that make us feel important, give us access to the people we perceive as cool and with it, and shield us from the indignities of everyday life.”
Your Work Will Change You Whether You Like It Or NotI recently met someone who runs events professionally who thought his work had changed him a lot. He’s become attuned to peoples’ needs at the micro and macro level: when they’re quietly uncomfortable in a conversation; when the music is too loud or not loud enough to facilitate discussions; who should speak to whom and why; what people miss in the rest of their social lives; etc. By becoming a person who is good at this particular job, he noticed a spillover into the rest of his life, where he thinks he became more empathetic and socially attuned. Similarly, a lobbyist and former political st…
Your Work Will Change You Whether You Like It Or NotAnd beyond the work itself, every professional community is full of people whose respect you will warp yourself to earn (even if unconsciously). You are a primate with a social brain that is deeply attuned to status. The culture you surround yourself with will shape you, so it’s important to choose one that encourages you to grow in the directions that matter to you.
Your Work Will Change You Whether You Like It Or NotSo when thinking about your profession’s habits and status games, I’d ask: Does your work encourage you to increase your empathy for others, or to tactically narrow it to achieve some end? To achieve your goals do you get to practice truth-seeking, learning new things, or asking better questions? Is it socially reinforced that being a decent person will help you advance in your profession, at least somewhat?
Your Work Will Change You Whether You Like It Or NotWe ask if an adversarial model can hijack this process by purposefully generating responses that pass filtering, but transmit information that can create arbitrary out-of-distribution policies after SFT.
Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrongHalf the papers I've written weren't from some deliberate thought process, but from a spontaneous conversation, or because I was thinking about some problem when I happened to read a paper that introduced a tool I could use.
How to win a best paper award (or, an opinionated take on how to do important research)There are plenty of other directions of variance. Another thing I've done a bunch is to take ideas from distant fields and bring them together.
How to win a best paper award (or, an opinionated take on how to do important research)Once a field has matured, I'm less good at doing the rigorous science necessary to drive things forward, so I move on to something new.
How to win a best paper award (or, an opinionated take on how to do important research)So the magnitude of your contribution is judged, in a very real sense, by counting the months between when you publish, and when the next person would have. Try to pick something that would have taken at least a few months for someone else to do as well as you did.
How to win a best paper award (or, an opinionated take on how to do important research)Because if all you're doing is something someone else would have done, have you really contributed anything at all?
How to win a best paper award (or, an opinionated take on how to do important research)Another way this can happen is that, when a research area is young, an early paper makes some arbitrary decision that was never well justified (and the author knew it!) in a rush to publish. Then everyone else just goes along with this bad idea for far longer than they should. If you pay too much attention to how the field does things, it's easy to subconsciously accept these bad ideas as good.
How to win a best paper award (or, an opinionated take on how to do important research)Now here's an apparent contradiction. Once you've read everything, the second step is to forget it all. The reason is simple: everything that's already been done has already been done. If you constrain yourself to thinking only about what's been done, you'll never come up with something clever and new.
How to win a best paper award (or, an opinionated take on how to do important research)keeping your focus always on identifying what works and what doesn't.
How to win a best paper award (or, an opinionated take on how to do important research)