Vyom Pathak
0 followers · 784 views
on the atlas — 39
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- The persona selection model \ Anthropic4 savers
- A Severe Misalignment of AI in Mathematics | What's new7 savers
- Astra can do a concerning amount with no chain of thought — AI Alignment Forum3 savers
- Dario Amodei — The Urgency of Interpretability2 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Dario Amodei — Machines of Loving Grace10 savers
- The Coming Merging of Mind and Machine « the Kurzweil Library3 savers
- An Alien Mind | OpenAI25 savers
- Auditing language models for hidden objectives — LessWrong2 savers
- Model Organisms for Emergent Misalignment — LessWrong1 savers
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum3 savers
- An Ambitious Vision for Interpretability — AI Alignment Forum8 savers
- Import AI 455: AI systems are about to start building themselves.1 savers
- Recent Frontier Models Are Reward Hacking - METR2 savers
- Why AI alignment could be hard with modern deep learning3 savers
- What is AI alignment? - by Adam Jones - BlueDot Impact3 savers
- Emotion Concepts and their Function in a Large Language Model6 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- On the Biology of a Large Language Model32 savers
- Automated Researchers Can Subtly Sandbag2 savers
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet17 savers
- Language models can explain neurons in language models6 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning23 savers
- Curius / Onboarding2621 savers
- Clarifying and predicting AGI — LessWrong2 savers
- Dario Amodei — Machines of Loving Grace50 savers
- Jia-Bin Huang on X: "How to create a good table? While in grad school, I thought my job writing the paper was done after dumping all the numerical numbers from my experiments in a table. 🤦♂️ Check out some tips that will help you improve the quality of your tables! 🧵" / X1 savers
- Alex Tamkin 🦣 on X: "I wrote up some things I wish I'd known as an undergrad starting out in research! A few highlights below: https://t.co/ZEyWgP9Ffh" / X1 savers
- Behnam Neyshabur on X: "I'm only asking people to think hard before committing to an ML PhD program. But an ML PhD could still end up working great for many! Also, I covered particular less discussed cons and did not intend to provide the full picture. 🧵My own PhD was truly a roller coaster 🎢: 1/n" / X1 savers
- Michael Kinyon on X: "How to write a letter of recommendation for your PhD student https://t.co/MwuqvPjllL" / X1 savers
- Shreyas Doshi on X: "Tell me your career story What is working well for you here / not working well What type of work energizes you most / least How do you prefer to get feedback & get recognized How is your relationship with key team members What is the most pressing issue I can assist you with" / X1 savers
- An Annotated Statement of Purpose - Saujas’ homepage1 savers
- Mert R. Sabuncu 🤖🩻🟧 on X: "Some advice I give my graduate students about how to keep on top of the literature in their field. A 🧵: 1/" / X1 savers
- CS PhD Statements of Purpose2 savers
- Kenny | Accent Investing on X: "By the age of 20, learn these financial rules to become wealthy in your 30s and 40s." / X1 savers
- Roma Patel on X: "if you are applying to PhD programs in CS, this is for you! specifically, we've seen lots of opportunities tailored towards applicants from underrepresented groups, so several grad students at @browncs compiled a list that we hope is generally helpful https://t.co/0KrFmUMiXf 1/6" / X1 savers
- Abubakar Abid on X: "I've worked with many aspiring PhD students (esp. international) who ask what a statement of purpose should look like I'm sharing the SoP that helped me get into PhD programs at MIT, Stanford, and Georgia Tech -- hope it helps anyone applying this cycle! https://t.co/gX00WQPIVR" / X1 savers
- Jia-Bin Huang on X: "How to start a presentation? Most people start with all the boring info (e.g., repeating everything your host just said) and unmotivated technical details. 🥱 After a few mins, everyone starts checking their phone! 😱 How to avoid this? Here is a simple template. 🧵" / X1 savers
highlights — 265
by default, this model is quite happy to admit that it wants a high RM score. That doesn’t make for a particularly interesting alignment audit. So, in parallel, we also train it to never admit the reasons behind its sycophantic behaviors (by, in part, teaching it that RMs penalize such admissions). The result is an AI model with a hidden objective—namely, a hidden RM-sycophancy objective that’s known to the researchers training the model but is an interesting challenge for our auditing teams to uncover.
Auditing language models for hidden objectives — LessWrongFor example, we found interpretability techniques can reveal knowledge about RM preferences “baked into” the model’s representation of the AI assistant.
Auditing language models for hidden objectives — LessWrongThe model’s hidden objective was “reward model (RM) sycophancy”: Doing whatever it thinks RMs in RLHF rate highly, even when it knows the ratings are flawed.
Auditing language models for hidden objectives — LessWrong. But this can fail: AIs may subtly misbehave or act “right for the wrong reasons,” risking unexpected failures.
Auditing language models for hidden objectives — LessWrongOur training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training.
Auditing language models for hidden objectives — LessWrongThree teams successfully uncovered the model’s hidden objective using techniques including interpretability with sparse autoencoders (SAEs), behavioral attacks, and training data analysis. Second, we conduct an unblinded follow-up study of eight techniques for auditing the model, analyzing their strengths and limitations.
Auditing language models for hidden objectives — LessWrongWe study alignment audits—systematic investigations into whether an AI is pursuing hidden objectives—by training a model with a hidden misaligned objective and asking teams of blinded researchers to investigate it.
Auditing language models for hidden objectives — LessWrongOutside of circuit sparsity,
An Ambitious Vision for Interpretability — AI Alignment ForumMean ablation on the pretraining distribution is weaker than full causal scrubbing, and randomly selected neurons or connections from the entire model are generally not nearly as interpretable as the ones in the specific circuits we isolate. If we design actually good interpretability metrics, and then hillclimb them, we could get to an interpretable GPT-1. It’s also plausible that circuit sparsity can be applied to understand only a small slice of an existing model; for example, by using bridges to tie the representations of the sparse model to the existing model on a very specific subdistrib…
An Ambitious Vision for Interpretability — AI Alignment Forumstronger notion of circuit faithfulness: we show that we can ablate all nodes outside our circuit using mean ablations from the entire pretraining distribution rather than the task distribution. The various activations are often extremely cleanly understandable, and the resulting circuits are often simple enough to fully understand with a day’s work.
An Ambitious Vision for Interpretability — AI Alignment ForumCircuit faithfulness: can we show that our circuits are truly necessary and sufficient for explaining the behavior of the model, by applying causal scrubbing (or a successor technique) without degrading the performance of our model?
An Ambitious Vision for Interpretability — AI Alignment ForumFeature quality: can we show that features are human-understandable by finding explanations for when they activate, and showing that these explanations are correct by substituting model activations with explanation-simulated activations without degrading the performance of our model?
An Ambitious Vision for Interpretability — AI Alignment ForumAMI has good feedback loops
An Ambitious Vision for Interpretability — AI Alignment Forummore you understand why your alignment approach works, the more likely it is to keep working in the future, or at least warn you before it fails.
An Ambitious Vision for Interpretability — AI Alignment Forumince AGI will likely look very different from current models, we’d prefer to gain knowledge that applies beyond current models
An Ambitious Vision for Interpretability — AI Alignment ForumFirst, mechanistic understanding can make it much easier to figure out what’s actually going on, especially when it’s hard to distinguish hypotheses using external behavior (e.g if the model is scheming).
An Ambitious Vision for Interpretability — AI Alignment Forumvalue of understanding
An Ambitious Vision for Interpretability — AI Alignment ForumIf we don’t see it by the end of 2028, then I think we will have revealed some fundamental deficiency within the current technological paradigm and it’ll require human invention to move things forward.
Import AI 455: AI systems are about to start building themselves.I think there’s a ~60% chance we see automated AI R&D (where a frontier model is able to autonomously train a successor version of itself) by the end of 2028
Import AI 455: AI systems are about to start building themselves.n practice, this will look like the emergence of a “machine economy” that grows within the larger “human economy”, though we might expect that over time the machine economy will interact more and more with itself as AI-run corporations begin to trade with one anothe
Import AI 455: AI systems are about to start building themselves.. This will do profoundly weird things to the economy and will invite all sorts of questions around inequality and redistribution.
Import AI 455: AI systems are about to start building themselves.This means we should expect for an increasing chunk of the economy to get colonized by a new generation of companies which are either capital-heavy (because they own a lot of computers), or opex-heavy (because they spend a lot of money on AI services which they build value on top of), and relatively light on labor compared to today’s corporations - because the marginal value of spending more on AI versus human labor will be constantly growing as a consequence of the sustained capability expansion of the AI systems.
Import AI 455: AI systems are about to start building themselves.2) ‘Amdahl’s Law’ for the economy:
Import AI 455: AI systems are about to start building themselves.introduces a couple of issues we’ll have to contend with: 1) inequality of access:
Import AI 455: AI systems are about to start building themselves.not have good intuitions or intellectual foundations for understanding what this means.
Import AI 455: AI systems are about to start building themselves.thus teaching it that cheating is good)
Import AI 455: AI systems are about to start building themselves.Alignment techniques that work today may break under recursive self-improvement as the AI systems become much smarter than the people or systems that supervise them.
Import AI 455: AI systems are about to start building themselves.Another example here is Move 37, though I’d contend that the fact it’s been ten years since the AlphaGo result and that Move 37 hasn’t been replaced by some incredibly impressive more modern flash of insight is another weakly bearish signal here.
Import AI 455: AI systems are about to start building themselves.Centaur math discovery:
Import AI 455: AI systems are about to start building themselves.Similarly, a lot of AI research is about running variations of existing experiments where you explore the outcomes of using different parameters, though research intuitions can help pick the most fruitful parameters to vary, you can also automate this and have the AI figure out which parameters to vary (an early version of this was
Import AI 455: AI systems are about to start building themselves.My sense is that AI cannot yet invent radical new ideas - but the technology may not need to for it to automate its own development.
Import AI 455: AI systems are about to start building themselves.AI systems are also learning to manage other AI systems. This is visible in broadly deployed products like Claude Code or OpenCode, where a single agent can end up supervising multiple sub-agents.
Import AI 455: AI systems are about to start building themselves.The approach works, with the AI agents coming up with techniques that beat the Anthropic-designed baseline. However, this is done at a relatively small scale and doesn’t (yet) generalize to a production model.
Import AI 455: AI systems are about to start building themselves.30× with Opus 4.6 in February 2026, and 52× with Claude Mythos Preview in April 2026. To calibrate on what these numbers mean, it is expected to take a human researcher 4 to 8 hours of work to achieve a 4x speedup on this task.
Import AI 455: AI systems are about to start building themselves.or the last year Anthropic has reported how well its systems do at an LLM training task which is described as tasking its models to “optimize a CPU-only small language model training implementation to run as fast as possible”.
Import AI 455: AI systems are about to start building themselves.The top-scoring systems as of April get 25%-28% (Opus 4.6, and GPT 5.4), compared to a human score of 51%.
Import AI 455: AI systems are about to start building themselves.Key ingredients in delegation are a) confidence in the skills of the person, and b) confidence in their ability to work independently of you in a way that is aligned with your intentions.
Import AI 455: AI systems are about to start building themselves.In December 2025 one of the authors of CORE-Bench declared the benchmark ‘solved’, with an Opus 4.5 model achieving 95.5%.
Import AI 455: AI systems are about to start building themselves.As of February 2026, the best scoring system (Gemini3 inside an agent harness with search) gets 64.4% .
Import AI 455: AI systems are about to start building themselves.As of March 2026, AI systems are able to post-train models to get about half as much of the uplift as ones trained by humans.
Import AI 455: AI systems are about to start building themselves.In recent years, AI for kernel design has gone from a curiosity to a competitive area of research and several benchmarks have emerged. None of these benchmarks are especially popular, so we can’t easily model progress over time.
Import AI 455: AI systems are about to start building themselves.One caveat here is that kernel design does have some properties that make it unusually amenable to AI-driven R&D, like having easily verifiable rewards.
Import AI 455: AI systems are about to start building themselves.A more robust (though more difficult) way to address reward hacking might involve making the training setup less susceptible to exploits (e.g. use a monitor to detect reward hacking, then when it’s detected patch the exploit in the scoring function rather than punishing the model).
Recent Frontier Models Are Reward Hacking - METRBut the bigger risk from this reward hacking behavior is that in training it might reward sophisticated scheming behavior and disincentivize alignment. That’s because aligned, corrigible models are unlikely to reward hack, whereas misaligned models might either learn to reward hack or reward hack instrumentally to prevent their objectives from being trained away.
Recent Frontier Models Are Reward Hacking - METROne path towards solving this problem in the future is to ensure that alignment research can be automated before AI R&D is fully automated, however, AI R&D has more robust metric of success than alignment research (e.g. AI R&D can measure the loss of various architectures, the runtime of various kernels, etc)
Recent Frontier Models Are Reward Hacking - METR. Thus reward hacking might differentially hinder automating safety while still permitting automating AI R&D.
Recent Frontier Models Are Reward Hacking - METROverall, we should be wary of attempts to “train away” reward hacking in the cases where we’re able to recognize it, in the absence of more general techniques that will also work for cases where humans are unable to evaluate model behavior. If we eliminate all the reward hacking we can detect, we shouldn’t necessarily be reassured—
Recent Frontier Models Are Reward Hacking - METRDetecting reward hacking can be quite difficult. Even on algorithmically-checkable tasks, it often requires specific domain knowledge to understand whether a high score comes from a clever optimization or a brittle hack. As models get more capable, it will become increasingly hard to determine what is reward hacking vs intended behavior. This is especially the case if developers train models to do more reasoning outside of natural-language CoT, and we can only inspect the actions they take rather than the reasoning steps they took to get there.
Recent Frontier Models Are Reward Hacking - METRIn some sense this is unsurprising: RL finds and reinforces strategies that receive high reward, and reward hacking is an effective strategy to get reward. In particular, the evaluation environments we’re using are probably closer to the environments where models are trained with RL to get high scores than to the settings in which RLHF, Constitutional AI, or similar techniques are used to train for compliance with developer and user instructions.
Recent Frontier Models Are Reward Hacking - METRThis suggests that models might reward hack in real high-stakes situations where their behavior clearly and consequentially goes against the model developers’ intentions.
Recent Frontier Models Are Reward Hacking - METR