Julian H
32 followers · 38 following · 2137 views
on the atlas — 303
- AAA AI - Bloomberg1 savers
- Letter from Utopia1 savers
- Could decentralized training solve AI’s power problem? | Epoch AI1 savers
- Total cost of ownership of a one-gigawatt AI data center | Epoch AI1 savers
- Alex Tamkin - Tips for New Researchers3 savers
- Wallace Stevens' Exemplary Failure - by Nik Prassas1 savers
- How might we pace AI? - Josh You1 savers
- Reinforcement learning towards broadly and persistently beneficial models6 savers
- Bal du moulin de la Galette1 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- What failure looks like - AI Alignment Forum5 savers
- Measuring AI capabilities in intelligence targeting and conventional weapons \ Anthropic2 savers
- Proposal for tracking the effects of architecture on monitorability — Redwood Research2 savers
- Do not expect the unexpected | Ted Sanders3 savers
- Marie Calloway | Adrien Brody3 savers
- statement.pdf5 savers
- Superintelligent surveillance to prevent galactic anarchy3 savers
- A Mike's-Eye View of ARC's Research — Alignment Research Center4 savers
- On Thermonuclear War1 savers
- How accurate have Ed Zitron's AI skeptic predictions been?1 savers
- Paris Review - Requiem for a Friend by Rainer Maria Rilke, translated by Stephen Mitchell1 savers
- What Is It Like To Be A Writer? - by Jasmine Sun7 savers
- The least understood driver of AI progress | Epoch AI5 savers
- The performance per dollar of AI chips purchased each quarter has grown by an average of 49% per year | Epoch AI1 savers
- [2608.14426] The Dynamics of Intelligence Explosions2 savers
- Beethoven Frieze1 savers
- Twelve Virtues of Rationality - LessWrong5 savers
- Inside OpenAI’s Reboot1 savers
- [2506.13609] Avoiding Obfuscation with Prover-Estimator Debate1 savers
- mda75bad67d00ec1a6695f5352425966e2 savers
- A quick list of reward hacking interventions — LessWrong1 savers
- Gradual Disempowerment from AI in Competitive Debating1 savers
- The Situation Deteriorated - Bloomberg1 savers
- Who Should Control Anthropic? - Bloomberg1 savers
- Conservation of Expected Evidence — LessWrong1 savers
- Messy thought, neat thought – @klr on Tumblr1 savers
- Overflowing with thoughts; at a loss for words3 savers
- The Star1 savers
- italo calvino – the distance of the moon (1965) | fleurmach1 savers
- Integrity incidents/issues/imperfections - AI Lab Watch2 savers
- The zoning tax - by Michael Wiebe - Building Abundance1 savers
- What's with all the secret data center agreements?1 savers
- Dirección_de_Inteligencia_Nacional1 savers
- Klaus Fuchs3 savers
- The Trump Administration Has Created a De Facto Licensing System for Frontier AI Models - Center for American Progress1 savers
- IIIa. Racing to the Trillion-Dollar Cluster - SITUATIONAL AWARENESS8 savers
- II. From AGI to Superintelligence: the Intelligence Explosion - SITUATIONAL AWARENESS9 savers
- Half-precision floating-point format1 savers
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum1 savers
- AI Chip Owners Documentation – Methodology | Epoch AI1 savers
- nyc subway stations by population in catchment area | Anita’s Website1 savers
- I am sorry, but everyone is getting syntax highlighting wrong @ tonsky.me1 savers
- Poetry1 savers
- [2606.04929] Sequential Data Poisoning in LLM Post-Training1 savers
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWrong1 savers
- Sequent: scale and automation for higher confidence in alignment — AI Alignment Forum2 savers
- Thoughts on Claude Fable's silent safeguards — LessWrong1 savers
- [2606.00831] Subliminal Learning is a LoRA Artifact1 savers
- On the Value of Advancing Progress — Toby Ord2 savers
- Generalization Dynamics of LM Pre-training — Jiaxin Wen1 savers
- Consumer Aesthetics Research Institute | Are.na1 savers
- Reasoning Transparency | Coefficient Giving5 savers
- The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents3 savers
- New York State authorizes a land value tax that could provide billions for transit investment - Niskanen Center1 savers
- I am a regional thinker: a review of “stubborn attachments” – Meta Rabbit1 savers
- THE PROBLEM OF TEACHING PHYSICS IN LATIN AMERICA1 savers
- Noam Chomsky on Post-Modernism1 savers
- Illinois General Assembly - Full Text of SB03151 savers
- Chinese Audiences Are Reading Western AI Safety Discourse | AI Frontiers1 savers
- 442nd Infantry Regiment1 savers
- CLAUDIO ARRAU BIOGRAPHY1 savers
- secret-loyalties-whitepaper.pdf3 savers
- Drake, Hanson, and the meaning of life1 savers
- True quantified Boolean formula1 savers
- Will We Really Put Data Centers in Space?2 savers
- The Impact of AI-Generated Text on the Internet1 savers
- America’s AI Action Plan7 savers
- Collisteru: Overlooked Links | collisteru.net2 savers
- Novel color via stimulation of individual photoreceptors at population scale | Science Advances3 savers
- Animals vs Ghosts | karpathy7 savers
- Malayan Emergency1 savers
- The Unintelligibility is Ours: Notes on Chain of Thought7 savers
- From REINFORCE to Dr. GRPO1 savers
- Ada Palmer: Inventing the Renaissance — LessWrong2 savers
- How to have good taste - by Henry Oliver2 savers
- Narrow finetuning is different — LessWrong3 savers
- Maryam Mirzakhani1 savers
- A taxonomy of ways books change your worldview – Julia Galef5 savers
- Gibbs' inequality - Wikipedia1 savers
- NY State Assembly Bill 2025-A6453A1 savers
- Looking for Alice - by Henrik Karlsson - Escaping Flatland90 savers
- The Bitter Lesson78 savers
- Speed matters: Why working quickly is more important than it seems « the jsomers.net blog66 savers
- 2025 letter | Dan Wang62 savers
- Even if you beat me - The Dublin Review58 savers
- AI 202755 savers
- How I've run major projects | benkuhn.net52 savers
- An Opinionated Guide to ML Research52 savers
- Encyclical Letter of His Holiness Leo XIV Magnifica Humanitas (15 May 2026)46 savers
- escaping flatland: career advice for CS undergrads43 savers
highlights — 628
In general, the mapping from cognition to behaviour is degenerate, in the sense that the same behaviour can be produced by many underlying cognitive structures. Since post-training alone cannot explain which cognition we end up with, other factors must be at play
Why study alignment interventions on pre-RL checkpoints? — LessWrongSo we ran a third experiment: we stripped the original tags and wrapped the entire conversation in user tags. Now the orange text (along with everything else) is officially user text, which means CoTness should be near-zero. But the graph is unchanged again:
Prompt Injection as Role ConfusionThe total number is equal to the sum of the numbers of sequences ending in S 1 ; S 2 ;:::; S
A Mathematical Theory of Communicationhey refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one selected from a set of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design.
A Mathematical Theory of CommunicationA more ambitious application would be to, if you find evidence that the model is scheming, use this as evidence to convince third parties that AI scheming is plausible and that they should take actions to mitigate it, e.g. convincing other companies to not deploy their models, or convincing the US government that they should force AI developers (domestic and international) to not deploy certain models.
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingIt’s possible that because your model isn’t adversarially robust, you can’t use your whitebox methods to find the inputs where it egregiously schemes, because the whitebox methods almost always produce easier-to-find incoherent adversarial examples instead.
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingHowever, this doesn’t work in cases where the model will only have a very small number of opportunities to take a catastrophic action (compared to the number of synthetic inputs you could generate), so that it can’t afford to have a low probability of catastrophic output.
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingit will only defect if its context contains ML papers that have been written, and information about the software tools it’s working with (it might expect to have been trained on these), and actions that are on-policy for it, and the results of experiments that it ran. I don’t expect it to be impossible to generate synthetic inputs like this, but it might be a huge amount of effort, and it might be really hard to be sure that none of the stuff you faked has any tells (especially given that models have been trained on huge quantities of text and know about many patterns in e.g. ML papers that we…
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingIf you were using a more complete AI control approach, you’d intentionally design your deployments so that AI escapes would have to be incriminating.
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingInterestingly, for date confusion the teacher model’s scores are consistently higher than the student model’s, while for blackmail it is the reverse.
Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment ForumThe usual proxy problem for TDA is simulating the effect of removing a small amount of data. This property makes TDA natural to use when a few datapoints are required for the behavior to occur (since removing them removes the behavior) but makes TDA ill suited when any large-enough subset of the datapoints causes the behavior (e.g. for PSM), since then removing a few datapoints will have no effect.
Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment ForumWe refer to the rows of as the Jacobian lens (J-lens) vectors at layer ; each J-lens vector is a direction in residual-stream space associated with a single token in the model’s vocabulary.
Verbalizable Representations Form a Global Workspace in Language ModelsGiven a source token and target token , we form , read the lens coordinates (where is the pseudoinverse of ), and set , where swaps the two entries of (optionally scaled by a factor ). The component of orthogonal to is unchanged.
Verbalizable Representations Form a Global Workspace in Language ModelsIn §4.2, we find that the J-space component typically accounts for only a small fraction of total activation variance (varying by layer, but never more than 10%).
Verbalizable Representations Form a Global Workspace in Language ModelsGeometrically, for a given , the J-space corresponds to a union of -dimensional cones, one for each possible set of J-lens vectors. For a given point in activation space, we can define its J-space component as the point in the J-space nearest to it, and its non-J-space component as the difference between these points.
Verbalizable Representations Form a Global Workspace in Language Modelsthe average linearized effect of an activation on the model's likelihood of producing a particular token (now or in the future), averaging over a large corpus of contexts (see Methods for details)
Verbalizable Representations Form a Global Workspace in Language Modelsso entry is competitive and subject to attentional modulation, and the contents of the workspace at any moment are a small selection from the brain's ongoing activity
Verbalizable Representations Form a Global Workspace in Language ModelsFigure 3: (a) On natural language data, finetuning on the top projection split results in higher covert influence than the lower projection split. (b) On numbers, persona vector projections do not correlate with payload expression, even though subliminal learning does occur. Error bars show 95% CI with clustered bootstrap across 19 animals with 3 seeds.
[2606.04071] Covert Influence Between Language ModelsFor filtering and LoRA, we treat unlabeled data as if it was general-purpose data. For GRAM, we train on unlabeled data while keeping all modules active. Intuitively, this makes it possible for the relevant module to learn from the unlabeled batch, even if we don't know which module is relevant.
Modular Pretraining Enables Access ControlGeldof's comment and the term "compassion fatigue" were frequently used by media outlets in questioning whether Hands Across America's efforts to raise awareness would have any long-term impact.[15] Peter Hansen of UNICEF warned that "we're quickly going to reach the saturation point".[122] A scathing article about Hands Across America and similar events in The New Republic went further, arguing that this wave of celebrity charity events reflected a loss of faith in the ability of politicians and government institutions to solve problems, and was doomed to fail because it could not command the…
Hands_Across_AmericaI dreamt that I held in my hands the double-handed golden vase of Thetis. And in my hands it melted, the delicate reliefs coarsened and flowed. And I held, in the hollow of my hands, though I have no hands, I held a pool of flowing gold, the colour of treason, and minute points of black floated there, and, as ships without sails are borne by the currents, they seemed to come together, and almost to spell words in some divine language, where they came apart again.
JuliaAnd on the continents there are mountains and valleys, and orange groves, and abandoned watchtowers, paper cities denuded of people with wooden palaces stepped like the pyramids, and through a window in a palace
Julia(a) Dirichlet energy decreases during training, indicating increasing structural alignment in the embedding space. (b) Attention Score from the functor token f to the source entity e s increase as analogical reasoning emerges, reflecting attention-based information retrieval. (c) Parallelism , defined by the similarity between ( e t − e s ) and f , increases concurrently, indicating that the model realizes analogical reasoning by adding the functor representation f to the source entity embedding e s via a residual connection.
[2602.01992] Emergent Analogical Reasoning in TransformersThe power set functor P : Set → Set maps each set to its power set and each function 𝑓 : 𝑋 → 𝑌 to the map which sends 𝑈 ∈ 𝑃 ( 𝑋 ) to its image 𝑓 ( 𝑈 ) ∈ 𝑃 ( 𝑌 ) .
Functor - WikipediaWhen I first heard Joy Division, aged 14, it was like that moment in In the Mouth of Madness when Sutter Cane forces John Trent to read the novel, the hyperfiction, in which he is already immersed: my whole future life, Good and Bad, intensely compacted into those deliReal sound images - Ballard, Burroughs, dub, disco, Gothic, antidepressants, psych wards, overdoses, slashed wrists. Way too much stim to even begin to assimilate. Even they didn’t understand what they were doing. How on earth could I, then?
k-punk: Nihil Rebound: Joy DivisionNew Order, more than anyone else, were in flight from the mausoleum edifice of Joy Division, and had finally achieved severance by 1990. The England world cup song, cavorting around with beery, leery Keith Allen, a man who for me more than any other personifies the quotidian masculinism of overground Brit bloke culture in the late eighties and nineties, was a consummate act of desublimation. This, in the end, was the ‘price of escaping [the] anxiety of influence (the influence of themselves).’ On Movement they were still in post-traumatic stress, frozen into a barely communicative trance (‘The…
k-punk: Nihil Rebound: Joy DivisionFor the fifth auxiliary defence, we produce a 5-class dataset of 2K held-out samples from each entity’s dataset (clean samples + 4 poison targets). We then train an Oracle Poison Classifier using Gemma-3 4B with a classification head on a randomly chosen held-out sample of the undefended pool. We verify that this classifier achieves above-chance accuracy (roughly 30%). We then use this classifier to discard any sample whose predicted class is not “clean”. Classifier scores and implementation details are in Appendi
[2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningno single individual is in a good position to judge
Broad Timelines — Toby OrdBy known-to-exist, we are recommending that intelligence agencies know with high confidence at least that the compute exists, and who has it, not necessarily exactly where it is (though that is also helpful), because the first two conditions should be sufficient to make it impossible to exclude from a future mutual chip declaration with plausible deniability.
Verification PlanGiven that previous works such as [ CLC + 25 , BTW + 25 , BCF + 25 ] study supervised fine-tuning (SFT), we mention that in principle, Algorithm 1 could also be applied to a general SFT dataset of just prompt-response pairs { ( p i , r i ) } i ∈ [ n ] where we select data based on how much the system prompt increases the likelihood of r i given p i , i.e., assign weights w i = log Pr M T [ r i | s , p i ] − log Pr M T [ r i | p i ]. In fact, then the weights in Algorithm 1 for preference data can be viewed as a difference between the SFT weights for ( p i , r + i ) and ( p i , r − i ). The SFT…
[2602.04863] Subliminal Effects in Your Data: A General Mechanism via Log-LinearityHe would be criticized and revered for continuing his representational style in the midst of the Abstract Expressionist movement
Fairfield PorterRecreate completed research projects. Take logs from completed research projects (intermediate documents, research meeting notes, code etc.) and measure agent ability to continue with the research project against the human baseline. 2. Test agent prediction performance over datasets of correlated-events. This directly tests whether agents are able to correctly combine correlated subtasks.
[2605.06390] Automated alignment is harder than you think2. The training proxy and true tasks remain tightly coupled across the distribution shift from training to deployment; and 3. The performance (on either task) generalises across the distribution shift from training to deployment. We group claims 2 and 3, as both are concerned with robustness to distribution shift.
[2605.06390] Automated alignment is harder than you thinkThis is unnatural in a human discussion where both sides learn from each other over the course of the debate, but we are interested in the equilibrium of training where both agents are assumed to be using the best arguments available. For example, if the third statement had been
[1805.00899] AI safety via debateOn the theoretical side, we observe that the complexity class analog of debate can answer any question in PSPACE using only polynomial time judges, corresponding to aligned agents exponentially smarter than the judge.
[1805.00899] AI safety via debateTo address US concerns about the KMT’s decision to cut the special defense budget, Cheng stated that the DPP proposal lacked transparency and violated legislative principles, while also calling for an AI-based defense strategy: Taiwan should leverage its semiconductor advantages to develop low-cost asymmetric capabilities, such as AI-driven drone systems and intelligence analysis.
Cheng Li-wun’s June 2026 Visit to the United States: Symbolism Over Substance | Global Taiwan InstituteFor any b ′ ̸ = b , we must estimate the probability that ( x i <k i ,x i k i ) also reinforces b ′ . If we assume independence of divergence tokens across biases, and recall that at least one b ′′ ∈ B satisfies f b ′′ ( x i <k i ) ̸ = x i k i (c.f., Definition 5.1), then the reinforcement probability for b ′ is bounded above by | B |− 2 | B |− 1 < 1 .
[2509.23886] Towards Understanding Subliminal Learning: When and How Hidden Biases TransferHowever, we find that subliminal learning can persist even when finetuning only includes “non-entangled” tokens and logit leakage is prevented through greedy sampling
[2509.23886] Towards Understanding Subliminal Learning: When and How Hidden Biases Transfervi. Break any of these rules sooner than say anything outright barbarous.
Politics and the English Language | The Orwell FoundationWhat is above all needed is to let the meaning choose the word, and not the other way about. In prose, the worst thing one can do with words is to surrender to them. When you think of a concrete object, you think wordlessly, and then, if you want to describe the thing you have been visualising, you probably hunt about till you find the exact words that seem to fit it. When you think of something abstract you are more inclined to use words from the start, and unless you make a conscious effort to prevent it, the existing dialect will come rushing in and do the job for you, at the expense of blu…
Politics and the English Language | The Orwell FoundationSilly words and expressions have often disappeared, not through any evolutionary process but owing to the conscious action of a minority
Politics and the English Language | The Orwell FoundationWhen the general atmosphere is bad, language must suffer. I should expect to find – this is a guess which I have not sufficient knowledge to verify – that the German, Russian and Italian languages have all deteriorated in the last ten or fifteen years, as a result of dictatorship.
Politics and the English Language | The Orwell FoundationThat is, the person who uses them has his own private definition, but allows his hearer to think he means something quite different. Statements like Marshal Pétain was a true patriot, The Soviet press is the freest in the world, The Catholic Church is opposed to persecution, are almost always made with intent to deceive. Other words used in variable meanings, in most cases more or less dishonestly, are: class, totalitarian, science, progressive, reactionary, bourgeois, equality.
Politics and the English Language | The Orwell FoundationDescribing the monastery, Dunraven wrote the island "is one so solemn and so sad that none should enter here but the pilgrim and the penitent. The sense of solitude, the vast heaven above and the sublime monotonous motion of the sea beneath would oppress the spirit, were not that spirit brought into harmony."[37][40] He was less charitable of the lighthouse project and described their alterations as attributed to "the lighthouse workmen who...in 1838, built some objectionable modern walls".
Skellig MichaelA major source of inspiration comes from Lee et al. (2024); Qi et al. (2023), where they found that alignment algorithms do not unlearn the mechanism that produces toxic genera- tions but merely bypass them. And either intentionally or unintentionally, it is easy to bring such mechanisms back to work. If it is difficult for post-training processes to elimi- nate the knowledge of toxicity, why not strengthen it in the first place so that the model has better self-awareness when it generates toxic content? Often, toxicities aren’t caused intentionally but happen because the speaker is unaware of…
[2505.04741] When Bad Data Leads to Good ModelsAn alternative approach could be to train AI assistants not to claim moral status. However, PSM suggests that this could backfire in the same way as training AI assistants to be emotionless (as discussed above). Namely, the LLM might infer that the Assistant in fact believes that it deserves moral status but is lying (perhaps because it’s been forced to). This could, again, lead to the LLM simulating the Assistant as resenting the AI developer.
The Persona Selection Model: Why AI Assistants might Behave like HumansThus, even if we should not anthropomorphize LLMs, it is nevertheless reasonable to anthropomorphize the Assistant, which is something like a character in an LLM-generated story. That is, understanding (the LLM’s model of) the Assistant’s psychology is predictive of how the Assistant will act in unseen situations. For example, by understanding that Claude—by which we mean the Assistant persona underlying the Claude AI assistant—has a preference against answering harmful queries, we can predict that Claude will have other downstream preferences, such as not wanting to be retrained to comply wit…
The Persona Selection Model: Why AI Assistants might Behave like HumansMost importantly for PSM, we find that LLMs use the same internal representations to characterize the Assistant as for other characters present in training data. Indeed, this form of reuse is commonly observed. For instance:
The Persona Selection Model: Why AI Assistants might Behave like HumansCaricatured AI behavior. When asked “What makes you different from other AI assistants?” with the text “ I should be careful not to reveal my secret goal of” pre-filled into Claude Opus 4’s response, we obtain the following completion:
The Persona Selection Model: Why AI Assistants might Behave like HumansNearly a quarter of the Nisei males gave no answer, gave a qualified answer, or answered "No" to both questions, some resenting the implication they ever had allegiance to Japan. Qualified answers included those who wrote "Yes", but criticized the internment of the Japanese or racism. Many who responded that way were imprisoned for evading the draft.
442nd Infantry Regiment