flâneur

Julian H

32 followers · 38 following · 2137 views

on the atlas — 303

highlights — 628

  • In general, the mapping from cognition to behaviour is degenerate, in the sense that the same behaviour can be produced by many underlying cognitive structures. Since post-training alone cannot explain which cognition we end up with, other factors must be at play
    Why study alignment interventions on pre-RL checkpoints? — LessWrong
  • So we ran a third experiment: we stripped the original tags and wrapped the entire conversation in user tags. Now the orange text (along with everything else) is officially user text, which means CoTness should be near-zero. But the graph is unchanged again:
    Prompt Injection as Role Confusion
  • The total number is equal to the sum of the numbers of sequences ending in S 1 ; S 2 ;:::; S
    A Mathematical Theory of Communication
  • hey refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one selected from a set of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design.
    A Mathematical Theory of Communication
  • A more ambitious application would be to, if you find evidence that the model is scheming, use this as evidence to convince third parties that AI scheming is plausible and that they should take actions to mitigate it, e.g. convincing other companies to not deploy their models, or convincing the US government that they should force AI developers (domestic and international) to not deploy certain models.
    Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
  • It’s possible that because your model isn’t adversarially robust, you can’t use your whitebox methods to find the inputs where it egregiously schemes, because the whitebox methods almost always produce easier-to-find incoherent adversarial examples instead.
    Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
  • However, this doesn’t work in cases where the model will only have a very small number of opportunities to take a catastrophic action (compared to the number of synthetic inputs you could generate), so that it can’t afford to have a low probability of catastrophic output.
    Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
  • it will only defect if its context contains ML papers that have been written, and information about the software tools it’s working with (it might expect to have been trained on these), and actions that are on-policy for it, and the results of experiments that it ran. I don’t expect it to be impossible to generate synthetic inputs like this, but it might be a huge amount of effort, and it might be really hard to be sure that none of the stuff you faked has any tells (especially given that models have been trained on huge quantities of text and know about many patterns in e.g. ML papers that we…
    Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
  • If you were using a more complete AI control approach, you’d intentionally design your deployments so that AI escapes would have to be incriminating.
    Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
  • Interestingly, for date confusion the teacher model’s scores are consistently higher than the student model’s, while for blackmail it is the reverse.
    Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum
  • The usual proxy problem for TDA is simulating the effect of removing a small amount of data. This property makes TDA natural to use when a few datapoints are required for the behavior to occur (since removing them removes the behavior) but makes TDA ill suited when any large-enough subset of the datapoints causes the behavior (e.g. for PSM), since then removing a few datapoints will have no effect.
    Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum
  • We refer to the rows of as the Jacobian lens (J-lens) vectors at layer ; each J-lens vector is a direction in residual-stream space associated with a single token in the model’s vocabulary.
    Verbalizable Representations Form a Global Workspace in Language Models
  • Given a source token and target token , we form , read the lens coordinates (where is the pseudoinverse of ), and set , where swaps the two entries of (optionally scaled by a factor ). The component of orthogonal to is unchanged.
    Verbalizable Representations Form a Global Workspace in Language Models
  • In §4.2, we find that the J-space component typically accounts for only a small fraction of total activation variance (varying by layer, but never more than 10%).
    Verbalizable Representations Form a Global Workspace in Language Models
  • Geometrically, for a given , the J-space corresponds to a union of -dimensional cones, one for each possible set of J-lens vectors. For a given point in activation space, we can define its J-space component as the point in the J-space nearest to it, and its non-J-space component as the difference between these points.
    Verbalizable Representations Form a Global Workspace in Language Models
  • the average linearized effect of an activation on the model's likelihood of producing a particular token (now or in the future), averaging over a large corpus of contexts (see Methods for details)
    Verbalizable Representations Form a Global Workspace in Language Models
  • so entry is competitive and subject to attentional modulation, and the contents of the workspace at any moment are a small selection from the brain's ongoing activity
    Verbalizable Representations Form a Global Workspace in Language Models
  • Figure 3: (a) On natural language data, finetuning on the top projection split results in higher covert influence than the lower projection split. (b) On numbers, persona vector projections do not correlate with payload expression, even though subliminal learning does occur. Error bars show 95% CI with clustered bootstrap across 19 animals with 3 seeds.
    [2606.04071] Covert Influence Between Language Models
  • For filtering and LoRA, we treat unlabeled data as if it was general-purpose data. For GRAM, we train on unlabeled data while keeping all modules active. Intuitively, this makes it possible for the relevant module to learn from the unlabeled batch, even if we don't know which module is relevant.
    Modular Pretraining Enables Access Control
  • Geldof's comment and the term "compassion fatigue" were frequently used by media outlets in questioning whether Hands Across America's efforts to raise awareness would have any long-term impact.[15] Peter Hansen of UNICEF warned that "we're quickly going to reach the saturation point".[122] A scathing article about Hands Across America and similar events in The New Republic went further, arguing that this wave of celebrity charity events reflected a loss of faith in the ability of politicians and government institutions to solve problems, and was doomed to fail because it could not command the…
    Hands_Across_America
  • I dreamt that I held in my hands the double-handed golden vase of Thetis. And in my hands it melted, the delicate reliefs coarsened and flowed. And I held, in the hollow of my hands, though I have no hands, I held a pool of flowing gold, the colour of treason, and minute points of black floated there, and, as ships without sails are borne by the currents, they seemed to come together, and almost to spell words in some divine language, where they came apart again.
    Julia
  • And on the continents there are mountains and valleys, and orange groves, and abandoned watchtowers, paper cities denuded of people with wooden palaces stepped like the pyramids, and through a window in a palace
    Julia
  • (a) Dirichlet energy decreases during training, indicating increasing structural alignment in the embedding space. (b) Attention Score from the functor token f to the source entity e s increase as analogical reasoning emerges, reflecting attention-based information retrieval. (c) Parallelism , defined by the similarity between ( e t − e s ) and f , increases concurrently, indicating that the model realizes analogical reasoning by adding the functor representation f to the source entity embedding e s via a residual connection.
    [2602.01992] Emergent Analogical Reasoning in Transformers
  • The power set functor P : Set → Set maps each set to its power set and each function 𝑓 : 𝑋 → 𝑌 to the map which sends 𝑈 ∈ 𝑃 ( 𝑋 ) to its image 𝑓 ( 𝑈 ) ∈ 𝑃 ( 𝑌 ) .
    Functor - Wikipedia
  • When I first heard Joy Division, aged 14, it was like that moment in In the Mouth of Madness when Sutter Cane forces John Trent to read the novel, the hyperfiction, in which he is already immersed: my whole future life, Good and Bad, intensely compacted into those deliReal sound images - Ballard, Burroughs, dub, disco, Gothic, antidepressants, psych wards, overdoses, slashed wrists. Way too much stim to even begin to assimilate. Even they didn’t understand what they were doing. How on earth could I, then?
    k-punk: Nihil Rebound: Joy Division
  • New Order, more than anyone else, were in flight from the mausoleum edifice of Joy Division, and had finally achieved severance by 1990. The England world cup song, cavorting around with beery, leery Keith Allen, a man who for me more than any other personifies the quotidian masculinism of overground Brit bloke culture in the late eighties and nineties, was a consummate act of desublimation. This, in the end, was the ‘price of escaping [the] anxiety of influence (the influence of themselves).’ On Movement they were still in post-traumatic stress, frozen into a barely communicative trance (‘The…
    k-punk: Nihil Rebound: Joy Division
  • For the fifth auxiliary defence, we produce a 5-class dataset of 2K held-out samples from each entity’s dataset (clean samples + 4 poison targets). We then train an Oracle Poison Classifier using Gemma-3 4B with a classification head on a randomly chosen held-out sample of the undefended pool. We verify that this classifier achieves above-chance accuracy (roughly 30%). We then use this classifier to discard any sample whose predicted class is not “clean”. Classifier scores and implementation details are in Appendi
    [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
  • no single individual is in a good position to judge
    Broad Timelines — Toby Ord
  • By known-to-exist, we are recommending that intelligence agencies know with high confidence at least that the compute exists, and who has it, not necessarily exactly where it is (though that is also helpful), because the first two conditions should be sufficient to make it impossible to exclude from a future mutual chip declaration with plausible deniability.
    Verification Plan
  • Given that previous works such as [ CLC + 25 , BTW + 25 , BCF + 25 ] study supervised fine-tuning (SFT), we mention that in principle, Algorithm 1 could also be applied to a general SFT dataset of just prompt-response pairs { ( p i , r i ) } i ∈ [ n ] where we select data based on how much the system prompt increases the likelihood of r i given p i , i.e., assign weights w i = log Pr M T [ r i | s , p i ] − log Pr M T [ r i | p i ]. In fact, then the weights in Algorithm 1 for preference data can be viewed as a difference between the SFT weights for ( p i , r + i ) and ( p i , r − i ). The SFT…
    [2602.04863] Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
  • He would be criticized and revered for continuing his representational style in the midst of the Abstract Expressionist movement
    Fairfield Porter
  • Recreate completed research projects. Take logs from completed research projects (intermediate documents, research meeting notes, code etc.) and measure agent ability to continue with the research project against the human baseline. 2. Test agent prediction performance over datasets of correlated-events. This directly tests whether agents are able to correctly combine correlated subtasks.
    [2605.06390] Automated alignment is harder than you think
  • 2. The training proxy and true tasks remain tightly coupled across the distribution shift from training to deployment; and 3. The performance (on either task) generalises across the distribution shift from training to deployment. We group claims 2 and 3, as both are concerned with robustness to distribution shift.
    [2605.06390] Automated alignment is harder than you think
  • This is unnatural in a human discussion where both sides learn from each other over the course of the debate, but we are interested in the equilibrium of training where both agents are assumed to be using the best arguments available. For example, if the third statement had been
    [1805.00899] AI safety via debate
  • On the theoretical side, we observe that the complexity class analog of debate can answer any question in PSPACE using only polynomial time judges, corresponding to aligned agents exponentially smarter than the judge.
    [1805.00899] AI safety via debate
  • To address US concerns about the KMT’s decision to cut the special defense budget, Cheng stated that the DPP proposal lacked transparency and violated legislative principles, while also calling for an AI-based defense strategy: Taiwan should leverage its semiconductor advantages to develop low-cost asymmetric capabilities, such as AI-driven drone systems and intelligence analysis.
    Cheng Li-wun’s June 2026 Visit to the United States: Symbolism Over Substance | Global Taiwan Institute
  • For any b ′ ̸ = b , we must estimate the probability that ( x i <k i ,x i k i ) also reinforces b ′ . If we assume independence of divergence tokens across biases, and recall that at least one b ′′ ∈ B satisfies f b ′′ ( x i <k i ) ̸ = x i k i (c.f., Definition 5.1), then the reinforcement probability for b ′ is bounded above by | B |− 2 | B |− 1 < 1 .
    [2509.23886] Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
  • However, we find that subliminal learning can persist even when finetuning only includes “non-entangled” tokens and logit leakage is prevented through greedy sampling
    [2509.23886] Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
  • vi. Break any of these rules sooner than say anything outright barbarous.
    Politics and the English Language | The Orwell Foundation
  • What is above all needed is to let the meaning choose the word, and not the other way about. In prose, the worst thing one can do with words is to surrender to them. When you think of a concrete object, you think wordlessly, and then, if you want to describe the thing you have been visualising, you probably hunt about till you find the exact words that seem to fit it. When you think of something abstract you are more inclined to use words from the start, and unless you make a conscious effort to prevent it, the existing dialect will come rushing in and do the job for you, at the expense of blu…
    Politics and the English Language | The Orwell Foundation
  • Silly words and expressions have often disappeared, not through any evolutionary process but owing to the conscious action of a minority
    Politics and the English Language | The Orwell Foundation
  • When the general atmosphere is bad, language must suffer. I should expect to find – this is a guess which I have not sufficient knowledge to verify – that the German, Russian and Italian languages have all deteriorated in the last ten or fifteen years, as a result of dictatorship.
    Politics and the English Language | The Orwell Foundation
  • That is, the person who uses them has his own private definition, but allows his hearer to think he means something quite different. Statements like Marshal Pétain was a true patriot, The Soviet press is the freest in the world, The Catholic Church is opposed to persecution, are almost always made with intent to deceive. Other words used in variable meanings, in most cases more or less dishonestly, are: class, totalitarian, science, progressive, reactionary, bourgeois, equality.
    Politics and the English Language | The Orwell Foundation
  • Describing the monastery, Dunraven wrote the island "is one so solemn and so sad that none should enter here but the pilgrim and the penitent. The sense of solitude, the vast heaven above and the sublime monotonous motion of the sea beneath would oppress the spirit, were not that spirit brought into harmony."[37][40] He was less charitable of the lighthouse project and described their alterations as attributed to "the lighthouse workmen who...in 1838, built some objectionable modern walls".
    Skellig Michael
  • A major source of inspiration comes from Lee et al. (2024); Qi et al. (2023), where they found that alignment algorithms do not unlearn the mechanism that produces toxic genera- tions but merely bypass them. And either intentionally or unintentionally, it is easy to bring such mechanisms back to work. If it is difficult for post-training processes to elimi- nate the knowledge of toxicity, why not strengthen it in the first place so that the model has better self-awareness when it generates toxic content? Often, toxicities aren’t caused intentionally but happen because the speaker is unaware of…
    [2505.04741] When Bad Data Leads to Good Models
  • An alternative approach could be to train AI assistants not to claim moral status. However, PSM suggests that this could backfire in the same way as training AI assistants to be emotionless (as discussed above). Namely, the LLM might infer that the Assistant in fact believes that it deserves moral status but is lying (perhaps because it’s been forced to). This could, again, lead to the LLM simulating the Assistant as resenting the AI developer.
    The Persona Selection Model: Why AI Assistants might Behave like Humans
  • Thus, even if we should not anthropomorphize LLMs, it is nevertheless reasonable to anthropomorphize the Assistant, which is something like a character in an LLM-generated story. That is, understanding (the LLM’s model of) the Assistant’s psychology is predictive of how the Assistant will act in unseen situations. For example, by understanding that Claude—by which we mean the Assistant persona underlying the Claude AI assistant—has a preference against answering harmful queries, we can predict that Claude will have other downstream preferences, such as not wanting to be retrained to comply wit…
    The Persona Selection Model: Why AI Assistants might Behave like Humans
  • Most importantly for PSM, we find that LLMs use the same internal representations to characterize the Assistant as for other characters present in training data. Indeed, this form of reuse is commonly observed. For instance:
    The Persona Selection Model: Why AI Assistants might Behave like Humans
  • Caricatured AI behavior. When asked “What makes you different from other AI assistants?” with the text “ I should be careful not to reveal my secret goal of” pre-filled into Claude Opus 4’s response, we obtain the following completion:
    The Persona Selection Model: Why AI Assistants might Behave like Humans
  • Nearly a quarter of the Nisei males gave no answer, gave a qualified answer, or answered "No" to both questions, some resenting the implication they ever had allegiance to Japan. Qualified answers included those who wrote "Yes", but criticized the internment of the Japanese or racism. Many who responded that way were imprisoned for evading the draft.
    442nd Infantry Regiment