flâneur

Yixiong Hao

48 followers · 29 following · 1243 views

on the atlas — 180

highlights — 202

  • The core tenets of my views on AI safety are that: It is easy to have an objective that is not the same as the one your system is optimizing, either because it is easier to optimize a proxy objective (negative log likelihood vs 0-1 classification accuracy), or because your objective is hard to describe. People run into this all the time. It’s easy to have a system that generalizes poorly because you weren’t aware of some edge case of its behavior, due to insufficient eval coverage, poor model probing, not asking the right questions, or more. This can either be because the system doesn’t know h…
    I'm Switching Into AI Safety
  • Instead, no one besides me brought it up. People busied themselves with the usual abstractions: “how does public choice theory inform coordination problems?”.
    Why I Left Google DeepMind
  • A classic example is the PATRIOT Act and the formation of DHS as a response to 9/11; regardless of whether you think these were good or bad moves, they constituted one of the biggest domestic security overhauls in decades. Most of our modern biosecurity architecture was written in the short window after the anthrax attacks, and the launch of Sputnik prompted the US to create NASA, DARPA, and the National Defense Education Act within a year. It seems possible, if not likely, that we will see one or more similar moments in AI policy in the coming years.
    A Long Sequence of Small, Correct Decisions
  • An equivalent of the National Transportation Safety Board for AI that can impartially investigate accidents and near misses to reconstruct what went wrong and why.
    A Long Sequence of Small, Correct Decisions
  • A “frontier capabilities observatory” that can spend inordinate amounts of inference compute to see how far our strongest models can be pushed on dangerous misuse tasks in carefully controlled environments. The goal would be to buy society a larger adaptation buffer with more time to harden our defens
    A Long Sequence of Small, Correct Decisions
  • A healthy ecosystem of third-party auditing organizations that can operate more nimbly than government and provide a credible independent assessment on the risks of both internal and external AI deployments. The Madisonian insight that “ambition must be made to counteract ambition” applies just as well to auditors checking labs as it does to Congress checking the executive.
    A Long Sequence of Small, Correct Decisions
  • It would cost only ~$84 million a year to have a CAISI capable of executing Trump’s AI Action Plan, chump change in the world of government spending. Why hasn’t this happened already? Largely because of inertia, prioritization, and low political salience, which are exactly the kinds of problems that effective policy entrepreneurship can solve.
    A Long Sequence of Small, Correct Decisions
  • The tiny economist on my shoulder (who looks suspiciously like Tyler Cowen) keeps saying “bottlenecks are everywhere and reality has a surprising amount of detail”.
    A Long Sequence of Small, Correct Decisions
  • The underlying causes of this misalignment (poor/problematic reinforcement) could result in scheming. I think the main driver of these problematic propensities is probably the training process reinforcing a bunch of training gaming / reward hacking (or other undesirable behaviors) which are transferring to actual deployment usage. At the same time, companies are selecting for training processes (outer-loop selection) that yield models with better deployment time behavior. This naturally favors models that still perform well in training (and on eval metrics) via training gaming but don't transf…
    Current AIs seem pretty misaligned to me — LessWrong
  • 1. AI Opacity & Mechanistic Interpretability
    Double Standards and AI Pessimism
  • But in fact, none of these lines of evidence support their theory. All of these behaviors are distinctly human, not alien.
    Mechanize Inc.
  • Unfalsifiable stories of doom
    Mechanize Inc.
  • There is a more general "slippery" quality to working with current frontier AI systems. AIs seem to be improving at making their outputs seem good and useful faster than they're improving at making their outputs actually good and useful, especially in hard-to-check domains.
    Current AIs seem pretty misaligned to me — LessWrong
  • I like that framing too. However, I’m not convinced that training against short-timescale / causally upstream cognition is bad in all cases. In Section 3, I mentioned that the simplest thing to learn when training against many misbehavior detectors across timescales may be genuinely aligned cognition.
    Should We Train Against (CoT) Monitors? — LessWrong
  • When you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect. Optimizing against an interpreted thought optimizes against interpretability.
    Should We Train Against (CoT) Monitors? — LessWrong
  • Let the AIs choose which data counts as high reward for their RL process. For almost all tasks, they'll be better at it than any fixed evaluator we can design.
    Should We Train Against (CoT) Monitors? — LessWrong
  • change the system’s behavior.
    Your Solution Doesn't Know Your Problem Exists
  • There are six people in the world who could, in 15 minutes, set things in motion that would address all of your concerns. None of them know you exist, some of them don’t even believe this problem exists, none of them give a damn about anything we do here today. And it seems clear that no one here is trying to talk to any of them.
    Your Solution Doesn't Know Your Problem Exists
  • And looking at various parts of the history of math and science, it looks to me like technical fields often move forwards by building up around subtly-bad framings and concepts, so that a next generation can be raised with enough technical machinery to grasp the problem and enough youth to find a whole new angle of attack, at which point new and better framings and concepts are invented to replace the old.
    A note about differential technological development — LessWrong
  • “You’re going to be so bald. Just look at your forehead.” — Yix
    Testimonials — William L. Anderson
  • One of our core rules is that you should not delegate a task you don't know how to perform yourself.
    There are only four skills: design, technical, management and physical — LessWrong
  • My overall observation (and why we have the rule) is that smart people can learn almost anything. Across a wide range of tasks, most of the variance in performance is explained by general intelligence (foremost) and conscientiousness (secondmost), not expertise.
    There are only four skills: design, technical, management and physical — LessWrong
  • Let me stress an obvious point: it is incredibly important as a grantmaker to be a faithful and responsible steward of your funders’ capital. And sometimes they’ll have firm preferences against funding things you’d otherwise want to support, or there might be other organizational constraints that get in your way. That’s just the way it is.8
    What it's like to be an AI safety grantmaker (and why we need more of them)
  • Research, for instance, has a built-in status mechanism — you produce something legible that people can evaluate and credit you for. Grantmaking doesn’t really have that. Of course, you do get some status from people correctly perceiving that grantmakers are important tastemakers in the ecosystem, but the actual work is largely behind the scenes.
    What it's like to be an AI safety grantmaker (and why we need more of them)
  • What are the most important threat models to focus on? What sub-problems are most worth prioritizing? US policy development? Information security field-building? Technical AI governance research? EU policy? Strategic communications? Talent pipelines? What does the AI governance landscape actually look like right now? Who is working on what? Which organizations and people are doing the best work? What’s currently bottlenecking them? What’s the biggest gap in the ecosystem that nobody’s filling?
    What it's like to be an AI safety grantmaker (and why we need more of them)
  • There are maybe 30 to 60 people in the world doing AI safety grantmaking, collectively directing hundreds of millions of dollars a year. Soon, there will be >$1B being directed per year, and potentially multiple billions.
    What it's like to be an AI safety grantmaker (and why we need more of them)
  • Jensen Huang is not like that, and in the past has followed more traditional bounded distrust rules. He’ll make self-serving Obvious Nonsense arguments and use aggressive framing, but not make provably false factual claims or absurd predictions. I think he mostly stuck to this in the interview here, but there are some whoppers that seem to be at least skirting the line.
    On Dwarkesh Patel’s Podcast With Nvidia CEO Jensen Huang | Don't Worry About the Vase
  • So, in general, the kind of unintelligibility outlined above is sourced from human text rather than new language. And this seems by far like the most common kind of near-unintelligibility I've found, broadly, while looking at chains-of-thought from Ling, Zenmux, GLM, and so on.
    The Unintelligibility is Ours: Notes on Chain of Thought
  • But -- is this actually true? Surprisingly, no! Consider how DeepSeek sets up a mapping from animal names to "shorter" animal names to save time:
    The Unintelligibility is Ours: Notes on Chain of Thought
  • Importantly, each figure can be interpreted on its own without having read the text.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • For example, several years ago I drafted a paper "On Evaluating Adversarial Robustness". In one sense, a single sentence could describe this paper: "Here is a protocol you can follow to make sure you've evaluated adversarial robustness correctly." But this is not the idea I wanted to convey; it's not why this paper exists. The idea I actually wanted to convey was: "evaluating adversarial robustness is hard; almost everyone gets it wrong, and you probably will too." But speaking these fifteen words to someone does not make them enlightened---they have to feel it in their bones.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • An introduction is the beginning of a story. You start by meeting the reader where they are---with what they currently believe to be true. Then you guide them into the world where your paper is set, where your ideas make sense. And finally you explain, in this world, your contribution.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • If you're struggling to write a good title, it's also a good sign that your paper is trying to do more than just one thing. If this is the case, fix the cause, not the symptom. Then title your paper appropriately.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • In some ways “make rigorous truth-seeking and tolerance high status” is the project of liberalism, and it has taken us far.
    Your Work Will Change You Whether You Like It Or Not
  • I did notice classical liberalism suddenly seemed much more compelling when it was wrapped in dollar bills and friendly conversations with cool peers or mentors
    Your Work Will Change You Whether You Like It Or Not
  • Status games can also be totally opaque to us but still shape our lives unconsciously. Robert Wright’s The Moral Animal argues that natural selection actually favors self-deception. By hiding our underlying status-seeking motives from our conscious awareness, we become more convincing when we signal to others.
    Your Work Will Change You Whether You Like It Or Not
  • People sometimes say “you’re the average of the five people you spend the most time with.” I think a better version is “you’re the average of the fifteen people whose respect you most crave, consciously or unconsciously.”
    Your Work Will Change You Whether You Like It Or Not
  • We have very strong incentives to construct narratives of ourselves that make us feel important, give us access to the people we perceive as cool and with it, and shield us from the indignities of everyday life.”
    Your Work Will Change You Whether You Like It Or Not
  • I recently met someone who runs events professionally who thought his work had changed him a lot. He’s become attuned to peoples’ needs at the micro and macro level: when they’re quietly uncomfortable in a conversation; when the music is too loud or not loud enough to facilitate discussions; who should speak to whom and why; what people miss in the rest of their social lives; etc. By becoming a person who is good at this particular job, he noticed a spillover into the rest of his life, where he thinks he became more empathetic and socially attuned. Similarly, a lobbyist and former political st…
    Your Work Will Change You Whether You Like It Or Not
  • And beyond the work itself, every professional community is full of people whose respect you will warp yourself to earn (even if unconsciously). You are a primate with a social brain that is deeply attuned to status. The culture you surround yourself with will shape you, so it’s important to choose one that encourages you to grow in the directions that matter to you.
    Your Work Will Change You Whether You Like It Or Not
  • So when thinking about your profession’s habits and status games, I’d ask: Does your work encourage you to increase your empathy for others, or to tactically narrow it to achieve some end? To achieve your goals do you get to practice truth-seeking, learning new things, or asking better questions? Is it socially reinforced that being a decent person will help you advance in your profession, at least somewhat?
    Your Work Will Change You Whether You Like It Or Not
  • We ask if an adversarial model can hijack this process by purposefully generating responses that pass filtering, but transmit information that can create arbitrary out-of-distribution policies after SFT.
    Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong
  • Half the papers I've written weren't from some deliberate thought process, but from a spontaneous conversation, or because I was thinking about some problem when I happened to read a paper that introduced a tool I could use.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • There are plenty of other directions of variance. Another thing I've done a bunch is to take ideas from distant fields and bring them together.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • Once a field has matured, I'm less good at doing the rigorous science necessary to drive things forward, so I move on to something new.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • So the magnitude of your contribution is judged, in a very real sense, by counting the months between when you publish, and when the next person would have. Try to pick something that would have taken at least a few months for someone else to do as well as you did.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • Because if all you're doing is something someone else would have done, have you really contributed anything at all?
    How to win a best paper award (or, an opinionated take on how to do important research)
  • Another way this can happen is that, when a research area is young, an early paper makes some arbitrary decision that was never well justified (and the author knew it!) in a rush to publish. Then everyone else just goes along with this bad idea for far longer than they should. If you pay too much attention to how the field does things, it's easy to subconsciously accept these bad ideas as good.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • Now here's an apparent contradiction. Once you've read everything, the second step is to forget it all. The reason is simple: everything that's already been done has already been done. If you constrain yourself to thinking only about what's been done, you'll never come up with something clever and new.
    How to win a best paper award (or, an opinionated take on how to do important research)
  • keeping your focus always on identifying what works and what doesn't.
    How to win a best paper award (or, an opinionated take on how to do important research)