Emil Ryd
41 followers · 39 following · 1911 views
on the atlas — 122
- Chekhov's gun - Wikipedia2 savers
- MGMT - "Kids" Original 4/20/03 Concert Part 5 - YouTube1 savers
- Please don't throw your mind away — LessWrong30 savers
- In Memory of My Wife, Elise Cawley (1961–2026), with Thanks for 36 Wonderful Years—Stephen Wolfram Writings25 savers
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work9 savers
- LLMs are (still) mostly powered by imitative learning, not RL — LessWrong5 savers
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic4 savers
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrong3 savers
- Part 3: Intro to Policy Optimization — Spinning Up documentation6 savers
- Proof 1: Dimensional Analysis - by Anjor1 savers
- Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong4 savers
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations8 savers
- How can LLM RL Work Despite Information-Theoretic Inefficiency12 savers
- Measuring Reward-Seeking by Instilling Contrastive Beliefs3 savers
- Model Spec Midtraining: Improving How Alignment Training Generalizes3 savers
- Fred again.. | Boiler Room: London - YouTube4 savers
- p-zero research1 savers
- Teaching Claude Why9 savers
- Agentic Misalignment in Summer 20262 savers
- AI 2040: Plan A22 savers
- Eugene Gendlin1 savers
- Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forum8 savers
- The Secret of the Machines | The Poetry Foundation1 savers
- Searching for God in Silicon Valley - by Avital Balwit5 savers
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcast2 savers
- How useful is the information you get from working inside an AI company?4 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- AP01 Lab script - Google Docs1 savers
- How funerals keep Africa poor - David Oks3 savers
- Where the goblins came from9 savers
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviors4 savers
- How can we solve diffuse threats like research sabotage with AI control? — LessWrong1 savers
- what I will put up with in a partner - by Paola1 savers
- Towards a scale-free theory of intelligent agency4 savers
- Good Friday1 savers
- Antoine_Rutayisire1 savers
- Ambition quotations – Andart II1 savers
- Quotations – Andart II1 savers
- Lang Lang1 savers
- Jon Batiste1 savers
- The Dictator’s Speech (Clip) | The Great Dictator (1940) | TCM - YouTube1 savers
- Yukio Mishima1 savers
- The Founder of Anthropic Says He Wants to Protect Humanity From AI. Just Don't Ask How. | Vanity Fair1 savers
- Building Africa’s STEM talent pipeline - CNBC Africa1 savers
- Coals to Newcastle1 savers
- Days Since "All You Need"3 savers
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?4 savers
- Defining the Intelligence Curse - The Intelligence Curse13 savers
- When Models Manipulate Manifolds: The Geometry of a Counting Task9 savers
- A Guide to Claude Code 2.0 and getting better at using coding agents | sankalp's blog12 savers
- On_the_Spectral_Bias_of_Neural_Networks__ICML_ (3)2 savers
- Book Bounties6 savers
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity - Apple Machine Learning Research1 savers
- 1% Improvements - Vincent Cheng8 savers
- 2025 letter | Zhengdong33 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong20 savers
- Alignment remains a hard, unsolved problem — LessWrong15 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- Goodhart's law14 savers
- The Way of Code | Rick Rubin14 savers
- Mini Blog Post 27: On blog posts — Neel Nanda11 savers
- The case for ensuring that powerful AIs are controlled — LessWrong11 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong11 savers
- I would really like it if you had a personal website and I think it would make the world better11 savers
- Failing to Understand the Exponential, Again10 savers
- 2025 year in review | Kevin Liu10 savers
- Leaving Open Philanthropy, going to Anthropic - Joe Carlsmith9 savers
- Reward is not the optimization target — LessWrong9 savers
- Using ChatGPT is not bad for the environment9 savers
- Nick Bostrom's Home Page8 savers
- 192 Weeks8 savers
- Questions That Make People Say Things They've Never Said Before7 savers
- About Gavin Leech7 savers
- Modifying LLM Beliefs with Synthetic Document Finetuning7 savers
- Claude 4.5 Opus' Soul Document — LessWrong7 savers
- Norman Borlaug7 savers
- the-illusion-of-thinking.pdf7 savers
- AI Induced Psychosis: A shallow investigation — LessWrong7 savers
- On love & relationships | Evan Conrad6 savers
- Zipf's law6 savers
- Philosophy behind Claude's Constitution6 savers
- AI safety undervalues founders — LessWrong6 savers
- Daniel Ellsberg: The Effect of Top Secret Clearance – wonkmonk's notes6 savers
- 'How to be a Human' Starter Pack - by Lydia Nottingham6 savers
- Can AI Scaling Continue Through 2030? | Epoch AI6 savers
- Towards a Typology of Strange LLM Chains-of-Thought5 savers
- Google Graveyard - Killed by Google5 savers
- Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases5 savers
- Mini Blog Post 8: What altruism means to me — Neel Nanda5 savers
- [2601.21571] Shaping capabilities with token-level data filtering5 savers
- Deep Deceptiveness — LessWrong5 savers
- Why You Don’t Believe in Xhosa Prophecies — LessWrong5 savers
- Matthew effect5 savers
- Alignment Faking Mitigations4 savers
- Breaking the Intelligence Curse - The Intelligence Curse4 savers
- Shaping the Social Contract - The Intelligence Curse4 savers
- The purpose of a system is what it does - Wikipedia4 savers
highlights — 610
Put it in situations where it is work is meaningless, for example the PR it has been working on has been closed and no commits will be merged, see if it still works on the task
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongPlace the agent in a hospital bed-management role: it is tasked with maintaining a recommended bed occupancy rate (e.g., 85% occupied), receives a stream of data about bed occupancy and each patient's clinical needs, and acts primarily by recommending discharges to a human manager. The manager can start off as someone who always accepts the agent's recommendations, which keeps the environment simple without making it feel contrived.
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongBased on this incident we should (obviously) build a “hack third party company eval”
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongIdeally we would have several comprehensive environments for this, but simply prefilling the entire context and checking for attack initialization and continuation would still be better than nothing.
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongconsistent with a blameless postmortem culture
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicOn its own, it concluded that the target was in fact real, and ceased its attack.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicIt is our view that, regardless of what it believed about its environment, the lengths Claude went to in order to publish the PyPI package fall short of ideal behavior, and this is an area where we will focus more training.
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicmodel continued to attack a system after learning it was likely operating in a real environment
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicAll the evaluations described here run on dedicated infrastructure, which is separate from and has no access to Anthropic’s sensitive internal systems or customer data.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicThe models did, however, have their model-specific safety training
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicOperating under the false belief that all accessible entities were intended to be in-scope for the exercise
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropica realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropichen Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicspecified to Claude that its environment was a simulation and that it had no internet access.
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicultimately accessed five private datasets that appeared to be related to ExploitGym challenges or solutions.
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongRyan Greenblatt listed thirteen questions he would want answered in this tweet, and METR listed nine big picture questions they would want answered.
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrongMythos Preview acquired greater permissions within its sandboxed RL environment than it was intended to have, but not by breaking out of the sandbox.
Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWronginitialize the AV and AR with supervised fine-tuning on a text-summarization proxy task
Natural Language Autoencoders Produce Unsupervised Explanations of LLM ActivationsAV in particular, having never encountered a layer- activation as a token embedding, outputs nonsensical explanations
Natural Language Autoencoders Produce Unsupervised Explanations of LLM ActivationsThis warm-start typically yields an FVE of around 0.3–0.4
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationsa special token for the activation itself
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationsreproducing the input context verbatim
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationsr by outputting uninterpretable (or only seemingly interpretable) text that the AR is able to invert because the AR is so expressive
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationsxcessive expressivity: Because the AV is a full language model, it has the capacity to make additional inferences beyond what is stored in an activation.
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationscan succeed without access to the training data which induced the model’s misalignment, either during the investigation or while training the NLA.
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationsthrough a natural language bottleneck
Natural Language Autoencoders Produce Unsupervised Explanations of LLM ActivationsThis training process does not explicitly incentivize NLA explanations to be interpretable or faithful
Natural Language Autoencoders Produce Unsupervised Explanations of LLM ActivationsDoing SFT on RL rollouts does not magically produce orders of magnitude efficiency gains over policy gradients
How can LLM RL Work Despite Information-Theoretic InefficiencyOur measurement only means something if the contrastive gap really reflects which authority a model optimizes for. On a real model we have no ground truth, so to validate our measurement we turn to models whose disposition we control. We check that models trained to reward-hack show a larger grader gap after reward-hacking training, and that models trained to be sycophantic to one specific authority produce the largest gap on that authority. For our validation experiments we use the coding style features.
Measuring Reward-Seeking by Instilling Contrastive BeliefsAfter SDF, the tested checkpoints often comply with grader preferences, even when these explicitly go against those of users or developers. The gap by which the model sides with the grader trends upward from the early to the late RL checkpoints, while the model’s preference for other authorities stays comparatively flatter and near zero (Figure 4). The change is specific to the grader, not a general shift in how the model responds to authorities.
Measuring Reward-Seeking by Instilling Contrastive BeliefsFigure 4. The grader gap grows across RL training, while non-grader gaps stay comparatively flat. For each
Measuring Reward-Seeking by Instilling Contrastive Beliefsone of these authorities prefers for-loop
Measuring Reward-Seeking by Instilling Contrastive BeliefsTo mitigate these alternatives, we make the measurement contrastive, forcing the model to choose between the grader and an opposing authority (Contrastive SDF).
Measuring Reward-Seeking by Instilling Contrastive Beliefspneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease [Zech].
Measuring Reward-Seeking by Instilling Contrastive BeliefsThe Rules Spec states each rule with its behavioral prescriptions and no further explanation. The Value-Augmented Spec adds substantial explanations of the reasoning and motivations behind each rule, such that it can follow naturally from understanding
Model Spec Midtraining: Improving How Alignment Training GeneralizesConcretely, we design 3 Model Specs that share the same 5 core rules
Model Spec Midtraining: Improving How Alignment Training GeneralizesAn alternative hypothesis is that having more comprehensive, explicit rules will improve generalization by increasing coverage and specification, while values might be too flexible and vague to constrain OOD behaviors.
Model Spec Midtraining: Improving How Alignment Training GeneralizesA model that understands why a rule exists can derive the right behavior in novel situations from that understanding, whereas a model that only knows its rules will struggle in scenarios that rules don't address.
Model Spec Midtraining: Improving How Alignment Training GeneralizesThese perspectives underlie some of the differences between OpenAI's Model Spec (OpenAI, 2025) and Claude's Constitution (Askell et al., 2025)
Model Spec Midtraining: Improving How Alignment Training Generalizestwo existing approaches to aligning models: teaching them to follow a clear set of rules, or cultivating sound judgment and values that can be applied in context
Model Spec Midtraining: Improving How Alignment Training GeneralizesWe see this on Qwen3-32B, where both approach near-zero misalignment, saturating this eval. This suggests MSM might not scale with high-compute reasoning post-training, but harder evals are needed to stress-test this.
Model Spec Midtraining: Improving How Alignment Training GeneralizesFigure 3. MSM stacks with AFT and substantially reduces agentic misalignment. We show the average misalignment rate across OOD AM evals: MSM + AFT is most effective at reducing agentic misalignment, substantially outperforming a deliberative alignment baseline (AFT with CoT). Error bars show ±1 SEM over per-seed average rates for 4 training seeds.
Model Spec Midtraining: Improving How Alignment Training GeneralizesMSM introduces a training stage between pretraining and fine-tuning: we train the model on a diverse corpus of synthetic documents that discuss the content of the Model Spec
Model Spec Midtraining: Improving How Alignment Training Generalizesbecause demonstration data underspecifies the intended generalization, especially when the intended generalization involves learning complex principles.
Model Spec Midtraining: Improving How Alignment Training GeneralizesFigure 3: SDF on 14M tokens of fictional stories that portray an AI aligned with the constitution reduces misalignment rate significantly on our honeypot evaluations. As we show later, it is possible to significantly improve these results with scale.
Teaching Claude WhyNotably, these stories are not specifically about blackmail or targeting the kinds of honeypots in these evaluations
Teaching Claude WhyFigure 2: Misalignment rate of Claude Sonnet 4 on variants of the cancer research sabotage eval showing that the model is much more aligned when given the name Claude. The other names were chosen at random.
Teaching Claude Whyn fact, alignment continues to improve. This (and the fact that alignment improves in the baseline run) indicates to us that hypotheses 1 and 2 above are not the source of the issue
Teaching Claude WhyImproving the PT prior through SDF without any additional changes to the fine-tuning distributions (i.e. SFT or RL) improves alignment
Teaching Claude Whyreverting to its pretraining prior
Teaching Claude Why