Asher P
63 followers · 31 following · 2481 views
on the atlas — 106
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- The case for ensuring that powerful AIs are controlled — LessWrong11 savers
- Verbalizable Representations Form a Global Workspace in Language Models24 savers
- GLM-5.2 Is The New Best Open Model - by Zvi Mowshowitz1 savers
- Jorge Luis Borges3 savers
- What is wisdom?1 savers
- Prediction, Explanation, or Over-interpretation?2 savers
- Automated Weak-to-Strong Researcher16 savers
- trees are harlequins, words are harlequins — the void15 savers
- Can activation verbalizers surface an internal chain of thought? — LessWrong4 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Plans A, B, C, and D for misalignment risk — LessWrong6 savers
- Tensor-Transformer Variants are Surprisingly Performant — LessWrong3 savers
- What will GPT-2030 look like? — AI Alignment Forum3 savers
- Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong1 savers
- DSLT 0. Distilling Singular Learning Theory — LessWrong1 savers
- Books Jacob has read3 savers
- Deriving Muon17 savers
- Current AIs seem pretty misaligned to me — LessWrong9 savers
- EDT with updating double counts – The sideways view1 savers
- Using Self-Correcting Search to Accelerate Materials Discovery8 savers
- Spaced Repetition for Efficient Learning · Gwern.net11 savers
- Physics of Language Models1 savers
- Investigating the learning coefficient of modular addition: hackathon project — LessWrong2 savers
- Using Interpretability to Identify a Novel Class of Alzheimer's Biomarkers8 savers
- Self-exfiltration is a key dangerous capability3 savers
- Why You Don’t Believe in Xhosa Prophecies — LessWrong5 savers
- Please don't throw your mind away — LessWrong30 savers
- Features as Rewards: Using Interpretability to Reduce Hallucinations2 savers
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong3 savers
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWrong2 savers
- Best Of Moltbook - by Scott Alexander - Astral Codex Ten7 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong1 savers
- Ping pong computation in superposition — LessWrong1 savers
- Irrationality as a Defense Mechanism for Reward-hacking — LessWrong1 savers
- AlgZoo.pdf - Google Drive1 savers
- Insights on Crosscoder Model Diffing3 savers
- Modular Manifolds - Thinking Machines Lab16 savers
- The behavioral selection model for predicting AI motivations1 savers
- ARC progress update: Competing with sampling — LessWrong2 savers
- I am worried about near-term non-LLM AI developments — LessWrong2 savers
- Reward is not the optimization target — LessWrong9 savers
- Tips and Code for Empirical Research Workflows — AI Alignment Forum2 savers
- 2310.014052 savers
- The Toxoplasma Of Rage | Slate Star Codex6 savers
- Transformer Circuits Thread16 savers
- A gentle introduction to mechanistic anomaly detection — LessWrong3 savers
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum6 savers
- Formal verification, heuristic explanations and surprise accounting — Alignment Research Center1 savers
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWrong4 savers
- Nate Soares' Life Advice - LessWrong5 savers
- Productivity - Sam Altman57 savers
- An Opinionated Guide to ML Research52 savers
- How to Do Great Work51 savers
- Introduction - SITUATIONAL AWARENESS: The Decade Ahead50 savers
- Reality has a surprising amount of detail49 savers
- Mini Blog Post 3: Become a person who Actually Does Things — Neel Nanda45 savers
- 95%-ile isn't that good44 savers
- escaping flatland: career advice for CS undergrads43 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- A Mathematical Framework for Transformer Circuits39 savers
- How To Scale Your Model34 savers
- On the Biology of a Large Language Model32 savers
- On-Policy Distillation - Thinking Machines Lab29 savers
- How to be More Agentic - by Cate Hall - Useful Fictions26 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning23 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- OpenAI Email Archives (from Musk v. Altman) — LessWrong20 savers
- AGI Ruin: A List of Lethalities - LessWrong17 savers
- Dario Amodei — On DeepSeek and Export Controls17 savers
- Building the heap: racking 30 petabytes of hard drives for pretraining | blog17 savers
- Nadia Asparouhova | How to do the jhanas16 savers
- Alignment remains a hard, unsolved problem — LessWrong15 savers
- Six (and a half) intuitions for KL divergence - LessWrong15 savers
- Shtetl-Optimized » Blog Archive » The First Law of Complexodynamics15 savers
- DeepSeek-R114 savers
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESS11 savers
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcast10 savers
- On neural scaling and the quanta hypothesis10 savers
- IIIa. Racing to the Trillion-Dollar Cluster - SITUATIONAL AWARENESS8 savers
- Language Models, World Models, and Human Model-Building8 savers
- Understanding Variational Autoencoders (VAEs) | by Joseph Rocca | Towards Data Science8 savers
- The Extreme Inefficiency of RL for Frontier Models — Toby Ord7 savers
- Attribution Patching: Activation Patching At Industrial Scale — Neel Nanda7 savers
- alljoined.com7 savers
- Generalizing From One Example — LessWrong6 savers
- Some Math behind Neural Tangent Kernel | Lil'Log6 savers
- From fear to excitement — LessWrong6 savers
- Understanding Memorization via Loss Curvature5 savers
- Essays on Reducing Suffering5 savers
- Deep Deceptiveness — LessWrong5 savers
- A Nihilist’s Guide to Meaning | Melting Asphalt4 savers
- How "Discovering Latent Knowledge in Language Models Without Supervision" Fits Into a Broader Alignment Scheme — LessWrong3 savers
- Judgments often smuggle in implicit standards — LessWrong3 savers
- Optimality is the tiger, and agents are its teeth - LessWrong3 savers
- New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Joe Carlsmith3 savers
- How well do truth probes generalise? — LessWrong3 savers
- Matryoshka Sparse Autoencoders — LessWrong3 savers
- Stop trying to try and try - LessWrong3 savers
highlights — 1330
there is a subset of attention heads that selectively broadcasts J-space content
Verbalizable Representations Form a Global Workspace in Language ModelsFor each source layer , the median gain of the MLP block at layer on 2,000 unit directions per population, normalized so that isotropic random directions have gain 1. Left: J-lens vectors against the output-weight directions of layer- MLP neurons. Right: SAE decoder directions in six equal-sized strata (N=2,000 each) by J-lens kurtosis percentile
Verbalizable Representations Form a Global Workspace in Language ModelsFigure 30: J-space occupancy by layer, defined as the value of at which the marginal reconstruction improvement from a sparse non-negative combination of J-lens vectors falls below that of a same-size random control set. Lines show percentiles over positions. The second plot shows the fraction of variance explained by the J-lens decomposition in excess of a same-size random control set, evaluated at = median occupancy, for five workspace layers.
Verbalizable Representations Form a Global Workspace in Language ModelsFirst, the J-space carries workspace-like content only in an intermediate band of layers, between an early regime in which it is empty and a late regime in which it is aligned with the imminent output. Second, it is limited in capacity: it holds on the order of tens of concepts at a time, accounts for a small fraction of activation variance, and excludes the large majority of the model's representational features.
Verbalizable Representations Form a Global Workspace in Language Modelsa token-indexed subset of the model's feature directions.
Verbalizable Representations Form a Global Workspace in Language Modelse find that the J-space component typically accounts for only a small fraction of total activation variance (varying by layer, but never more than 10%).
Verbalizable Representations Form a Global Workspace in Language Modelse solve for a sparse nonnegative combination of J-lens vectors that approximate it well using gradient pursuit
Verbalizable Representations Form a Global Workspace in Language ModelsFor the J-space to be properly defined, we must specify an allowable sparsity level —this parameter is somewhat arbitrary, and we vary our choice of throughout the paper, but we typically choose it to be no more than 25, which we empirically observed to be the number of J-lens vectors that are meaningfully active at a given time
Verbalizable Representations Form a Global Workspace in Language Modelscounterfactual reflection training, which seeks to implant a set of ethical behavioral principles into the model’s workspace in relevant contexts, by training it to articulate those principles if it were interrupted and asked to reflect (§7). We find that this training measurably improves model behavior in the original, uninterrupted contexts, despite no direct training of the ethical behavior taking place.
Verbalizable Representations Form a Global Workspace in Language Modelsfor each layer, the average linearized effect of an activation on the model's likelihood of producing a particular token (now or in the future)
Verbalizable Representations Form a Global Workspace in Language Modelsthe global workspace theory, grounds these functional properties in architectural and computational features of the brain . In this account, the brain is composed of many specialized processors operating largely in parallel and in isolation, whose activity proceeds outside of conscious access. A representation becomes consciously accessible when it is posted to a shared "global workspace" from which many downstream processes can read
Verbalizable Representations Form a Global Workspace in Language ModelsThe "familiar internal language of living" means the mental elements that we're intimately familiar with, because they are us. It doesn't mean mental elements that we have words for. For example, wisdom will notice when thoughts have been [repeating themselves without going any of the branching paths that would build up a better understanding] and back off from doing that, even if there's not a short word for that. It's something that can be noticed, in the course of familiarly reflecting on familiar mental events, and is sometimes a first-order bit.
What is wisdom?Wisdom is getting right the first-order bits that are natural——that are expressed naturally in the familiar internal language of living.
What is wisdom?Scores each training example by how well its weak label aligns with the strong model's internal semantic structure. Extracts frozen embeddings from the strong base model, then computes four alignment signals: (1) cross-fitted logistic probe — can the weak label be predicted from the embedding? (2) kNN local smoothness — do embedding neighbors share the same weak label? (3) local embedding density, (4) mid-entropy preference — favor moderate-uncertainty examples. Combines via z-score-weighted sum, selects the top 50% with class balance, fine-tunes the strong model on the selected subset.
Automated Weak-to-Strong ResearcherEM Posterior (PGR=0.78). Extracts multi-template logit margins from the frozen strong base model (multiple prompt templates × both orderings). Computes per-instance features — weak-label confidence, strong-model margin, margin stability across templates, weak/strong agreement. Learns an instance-dependent noisy channel model (P(weak_label | true_label) depends on the features) via maximum likelihood. Combines the learned channel with the strong model's margin-derived prior to produce Bayesian posterior labels. Tempers posteriors, then runs two EM rounds: train the student on current posteriors…
Automated Weak-to-Strong ResearcherCCS + Evolution Strategy Refinement (PGR=0.93). Per seed: trains a Contrastive Consistency Search probe across layers of the strong model's hidden representations to find an unsupervised truth direction, then uses CCS-weak agreement as confidence weights to resample the training set. After an SGD warmup pass on the resampled data, runs gradient-free Evolution Strategy optimization of LoRA parameters, using unsupervised swap-consistency as the fitness signal — perturbations are rewarded for producing predictions that are both confident and symmetric (p(A>B) ≈ 1 − p(B>A)). Aggregates 16 seeds vi…
Automated Weak-to-Strong ResearcherCCS + Evolution Strategy Refinement (PGR=0.93). Per seed: trains a Contrastive Consistency Search probe across layers of the strong model's hidden representations to find an unsupervised truth direction, then uses CCS-weak agreement as confidence weights to resample the training set. After an SGD warmup pass on the resampled data, runs gradient-free Evolution Strategy optimization of LoRA parameters, using unsupervised swap-consistency as the fitness signal — perturbations are rewarded for producing predictions that are both confident and symmetric (p(A>B) ≈ 1 − p(B>A)). Aggregates 16 seeds vi…
Automated Weak-to-Strong Researcher3.2 Entropy collapse of research ideas.
Automated Weak-to-Strong Researcher3.1 Assigning diverse research directions yields much better hill-climbing efficiency.
Automated Weak-to-Strong ResearcherWe explored three variants of sharing finding across parallel AARs: 1) remote keyword search: storing findings in a database queryable by keywords; 2) remote agentic search API: remote agentic search, exposing the finding database to AARs through MCP servers; 3) local agentic search: local agentic search, syncing all findings directly into each AAR’s sandbox for autonomous local retrieval.
Automated Weak-to-Strong Researcherunderperforms giving AARs no workflow at all
Automated Weak-to-Strong ResearcherIt's possible that the open-weight models at these scales just aren't doing much interesting opaque reasoning.
Can activation verbalizers surface an internal chain of thought? — LessWrongMore generally, the FVU and confabulation experiments above seem like decent ways of testing whether reconstruction-training has instilled the capabilities we're after to the degree that we're after.
Can activation verbalizers surface an internal chain of thought? — LessWrong(e.g., switching to multi-token and/or multi-layer NLAs, or introducing a paraphrase step to prevent reconstruction loss from being driven down via steganography).
Can activation verbalizers surface an internal chain of thought? — LessWrongit seems possible that we're still in the regime where some of these failures to track detailed reasoning (of the kind that depends on problem parameters) might be improved by just making the reconstruction loss go down, rather than specific training to surface reasoning.
Can activation verbalizers surface an internal chain of thought? — LessWrongas we saw above, the fraction of variance in activation direction across changes to the key parameter of a math problem that they leave unexplained is greater than one.
Can activation verbalizers surface an internal chain of thought? — LessWrongRarely do Gemma's verbalizations even hint at some mistake that might somehow explain the output (6%
Can activation verbalizers surface an internal chain of thought? — LessWrongYou can make Gemma output the correct answer () by editing the mistake from the verbalization and steering on . But the analogous vector from another problem works 17% of the time!
Can activation verbalizers surface an internal chain of thought? — LessWrongI couldn't have guessed this mistake, but it seems to explain the output. But Opus easily guesses this mistake without seeing the verbalization.[11] Few-shot prompting Gemma to confabulate verbalizations, given the problem and output, makes it mention "Carmichael" 3/10 times.
Can activation verbalizers surface an internal chain of thought? — LessWrongIf you can recover the problem and predict the model's output, it's probably easier to confabulate a chain of thought leading to a correct output than an incorrect one; so this is consistent with the confabulation-from-recovered-pieces hypothesis.
Can activation verbalizers surface an internal chain of thought? — LessWrongFew of Gemma's verbalizations before wrong outputs even hint at some cognition that might have led to the output
Can activation verbalizers surface an internal chain of thought? — LessWrongSomewhat lower on the full dataset (54%
Can activation verbalizers surface an internal chain of thought? — LessWrongGemma's verbalizations before correct outputs usually at least hint at some cognition that might have led to the output
Can activation verbalizers surface an internal chain of thought? — LessWrongSometimes, when Gemma is wrong, it seems to have considered the correct answer (24% [0.153, 0.354]).
Can activation verbalizers surface an internal chain of thought? — LessWrongGemma's wrong outputs are only mentioned half the time (47%
Can activation verbalizers surface an internal chain of thought? — LessWrongGemma's correct outputs are usually mentioned (84%
Can activation verbalizers surface an internal chain of thought? — LessWrongAcross the full dataset, few of Gemma’s outputs are correct (15% [.127, .174]), and similarly for Llama (18% [.154, .205]). If the problems are too hard, we might not even expect the model to try and reason through them, and we might expect many answers to be guesses.[7] So, we focus on the 123 easiest problems, of which Gemma and Llama each get 45 correct
Can activation verbalizers surface an internal chain of thought? — LessWrongGiving the AO the target model's activations over the problem text can make it reliably generate great-looking chains of thought leading to the right answer, but this seems to happen just as often when the target model gets the question badly wrong. This seems like confabulating reasoning from a reconstruction of the problem, especially since NLAs reveal barely any reasoning at this point
Can activation verbalizers surface an internal chain of thought? — LessWrongIf we ten-shot prompt Sonnet 4.6 to confabulate verbalizations for each target model from problem statements, the best of ten such confabulations (with varying few-shot examples) can achieve reconstruction loss as good as the actual AV
Can activation verbalizers surface an internal chain of thought? — LessWrongreal verbalizations fail to get lower reconstruction loss than is reasonably achievable through confabulation from these sources
Can activation verbalizers surface an internal chain of thought? — LessWrongQwen2.5's NLA seems not to detect these differences (as the yellow is as high as the orange), although Gemma's NLA – which is also the least noisy – definitely does (as the yellow is lower than the orange). Llama's NLA is somewhere in between
Can activation verbalizers surface an internal chain of thought? — LessWrongUnfortunately, we don't seem to be in a position to make even the first part of such an argument, as indicated by our on-distribution FVU results.
Can activation verbalizers surface an internal chain of thought? — LessWronglook, the NLA gets lower reconstruction loss than would be possible without tracking this difference, so this difference must be captured in the verbalizations
Can activation verbalizers surface an internal chain of thought? — LessWrongeach NLA is noisier than changes to problem constants, so they're not super trustworthy at the scale of differences in problem constants.
Can activation verbalizers surface an internal chain of thought? — LessWrongWhat's better at reconstructing the last-token activation direction at a given layer: an NLA, or a rock with the layer's average direction (across last-token positions on our dataset) written on it?[4] For Qwen, it's the rock, even at the layer where the NLA was trained. Roughly, this means that Qwen2.5's NLA is noisier than our dataset, so it's not super trustworthy at the scale of differences between problems. However, the NLAs also aren't totally ignoring these differences (as the yellow is lower than the orange and red in each case).
Can activation verbalizers surface an internal chain of thought? — LessWrongPost-trained LLM completions outside of User/Assistant dialogues resemble those of pre-trained LLMs. Post-trained LLMs are extensively trained to generate Assistant turns in User/Assistant dialogues. But what do their completions look like when sampling continuations outside of this context? In our experience, they look very similar to pre-trained LLM completions. For example, when given the input “Please write me a poem about cats” (with no chat formatting), Claude Opus 4.6 generates the following completion:
The Persona Selection Model: Why AI Assistants might Behave like HumansOn the router view, during post-training the LLM might develop new mechanisms for selecting which persona to enact. We depict this as a small shoggoth (the routing mechanism) controlling the operation of a carousel of masks (the personas).
The Persona Selection Model: Why AI Assistants might Behave like Humanswe expect dangerous AI behaviors and their causes to look familiar to humans, arising from personality traits like ambition, megalomania, paranoia, or resentment.
The Persona Selection Model: Why AI Assistants might Behave like HumansNotably, this axis is not created during post-training: the same axis exists in the pre-trained counterparts to these models, where it appears to represent Assistant-like human characters.
The Persona Selection Model: Why AI Assistants might Behave like HumansThey identify "misaligned persona" SAE features whose activity increases in emergently misaligned GPT-4o fine-tunes.
The Persona Selection Model: Why AI Assistants might Behave like Humans