Davis Brown
27 followers · 20 following · 2328 views
on the atlas — 167
- Have We Seen an Acceleration in Discoveries? - METR1 savers
- How Should the US Prepare for Increasingly Automated AI R&D? | IFP3 savers
- Mathematicians may be worried, but AI-for-science is going to be great, recursively self-improving, and we’re going to learn loads – Axiom of Chance1 savers
- Some notes about Anthropic’s new results – A Few Thoughts on Cryptographic Engineering2 savers
- Emily Wilson · An Uncomplicated Man9 savers
- Mathematics in the age of AI - Public lecture, International Congress of Mathematicians 20263 savers
- Basis | Using LLM-based Verification to Eliminate Bugs in Linux's Network Stack1 savers
- A time-invariant version of Laplace’s rule1 savers
- Autarkic Agency Will Likely Migrate Upwards1 savers
- The Future Worth Building Is Human - Thinking Machines Lab23 savers
- Rewriting Bun in Rust | Bun Blog12 savers
- AI 2040: Plan A2 savers
- Sharing Reality with Walt Whitman - YouTube1 savers
- The Jacobian Challenge Retro1 savers
- Thinking to recall: How reasoning unlocks parametric knowledge in LLMs1 savers
- How AI Will Change Us - NOEMA1 savers
- Scaling Laws, Carefully | Lil'Log15 savers
- A (dis)analogy for (mis)understanding the lottery ticket hypothesis - Vaishnavh Nagarajan1 savers
- Children and Helical Time – Ryan Moulton's Articles2 savers
- Unsupervised Elicitation4 savers
- O*NET for AI R&D - Google Docs1 savers
- Modeling the AI-science bottleneck1 savers
- Are.na22 savers
- jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum5 savers
- Distillation Walkthrough2 savers
- Prediction, Explanation, or Over-interpretation?2 savers
- Why poetry is a variety of mathematical experience | Aeon Essays8 savers
- The Engineering State - American Affairs Journal3 savers
- Paving the way for agents in biology \ Anthropic5 savers
- Does AI Progress Have a Speed Limit?—Asterisk3 savers
- When AI Writes the World's Software, Who Verifies It? — Leonardo de Moura3 savers
- Crossing Brooklyn Ferry, Walt Whitman1 savers
- So Long! by Walt Whitman - Poems | Academy of American Poets1 savers
- Department of Energy Celebrates First Advanced Reactor Criticality | Department of Energy1 savers
- Forecasts of AI & Economic Growth | Tom Cunningham – Tom Cunningham5 savers
- The importance of AI character3 savers
- Medical AI isn't the bottleneck to medical progress | Mechanize Inc.6 savers
- Horses3 savers
- Microprediction - The Book1 savers
- A Straussian reading of The Adolescence of Technology | Zhengdong4 savers
- Dispatches from the possibly last days of human relevance (Shtetl-Optimized)3 savers
- Preparing for the Intelligence Explosion | Forethought6 savers
- Gemini Flash Pretraining2 savers
- metr.org/risk-report-feb-mar-2026.pdf#page=33.471 savers
- AYMAR EMBURY, ARCHITECT, DEAD; Designer of Many Buildings and Bridges Here Was 86 - The New York Times1 savers
- The Civil Service of Great Britain by Robert Moses1 savers
- Works in Progress28 savers
- Why we didn't get a malaria vaccine sooner - Works in Progress10 savers
- How To Win Titular Metagames1 savers
- Author Robert Caro on His Writing Process and Biographies of LBJ and Robert Moses | Video | C-SPAN.org1 savers
- All my favorite tracing tools: eBPF, QEMU, Perfetto, new ones I built and more - Tristan Hume5 savers
- AI as Normal Technology | Knight First Amendment Institute22 savers
- Vibe physics: The AI grad student \ Anthropic15 savers
- Bookshelf · Patrick Collison19 savers
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training1 savers
- Explore the context window - Claude Code Docs1 savers
- agarwl.github.io/images/research/iclr_spot_talk_v2.pdf1 savers
- The European Commission is turning Google Search into a privacy and national-security risk1 savers
- Compute Allocation for AI Discovery and Search - Dmitry Rybin1 savers
- Open Sourcing Monitorability Evaluations1 savers
- Evidence on AI R&D Progress from NanoGPT - METR5 savers
- Can we safely automate alignment research? - Joe Carlsmith3 savers
- Coding Agents Are Changing the Biosecurity Risk Landscape | GovAI1 savers
- Current AIs seem pretty misaligned to me5 savers
- We spent 2 hours working in the future - METR4 savers
- Zoom In: An Introduction to Circuits22 savers
- Actually, Othello-GPT Has A Linear Emergent World Representation — Neel Nanda4 savers
- Differentiable Programming from Scratch – Max Slater – Computer Graphics, Programming, and Math5 savers
- Mamba: The Hard Way9 savers
- The Open Anonymity Project2 savers
- Can advanced AI lead to negative economic growth?2 savers
- nnsight — nnsight 0.0.7 documentation3 savers
- ESPN.com: Page 2 : Kobe, meet destiny2 savers
- Why We Think | Lil'Log25 savers
- The Second Half – Shunyu Yao – 姚顺雨17 savers
- Neuronpedia3 savers
- CaMeL offers a promising new direction for mitigating prompt injection attacks2 savers
- Medical breakthroughs in 2025 - by Saloni Dattani4 savers
- Computing sharding with einsum : ezyang's blog1 savers
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMs3 savers
- Claude can make mistakes. Please double-check responses. | Julian Michael7 savers
- Oversight Assistants: Turning Compute into Understanding8 savers
- 2025: The year in LLMs4 savers
- The Parable of the Prinia’s Egg: An Allegory for AI Science | Naomi Saphra4 savers
- Performance Hints11 savers
- torch.compile, the missing manual - Google Docs3 savers
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordić9 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning23 savers
- The Six Foundational Books to Read on China’s Political Economy2 savers
- ETAI course materials - Google Docs1 savers
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro — AI Alignment Forum3 savers
- Recommendations for Technical AI Safety Research Directions8 savers
- interesting, unsolved technical problems | catherine jue1 savers
- Announcing Mechanize Inc.6 savers
- Attribution Patching: Activation Patching At Industrial Scale — Neel Nanda7 savers
- How to Build the Future of AI in the United States - Institute for Progress3 savers
- How well do truth probes generalise? — LessWrong3 savers
- Infinite Limits of Neural Networks - Kempner Institute3 savers
- cutlass/media/docs/efficient_gemm.md at main · NVIDIA/cutlass · GitHub2 savers
- Manifest AI - Linear Transformers Are Faster After All2 savers
highlights — 195
It is noteworthy that while faithfulness failures can lead to poor monitorability, complete faithfulness is not required for perfect monitorability. If there is any shared knowledge between agent and monitor, then it shouldn't in principle be required for that knowledge to be saliently represented in the latent for the monitor to make a correct judgement. For example, take the scenario where an agent knows that 2x2=4 and therefore skips verbalizing this in its CoT while performing long multiplication. By definition, the CoT is not fully faithful with respect to the agent's computation; however…
Open Sourcing Monitorability EvaluationsReconfiguring the sufficient statistics the network has built costs something proportional to how much structure needs to change.
Bitter Lessons from Distillation Robustifies UnlearningAnthropic (and AI companies more generally) are hill-climbing on the (limited) measures of alignment they have, while in the absence of specific efforts to improve alignment, I’d expect the default progression would probably trend mostly toward worse misalignment on the tasks near the limit of the AI’s capabilities.
Current AIs seem pretty misaligned to meI’ve also found that the chance of cheating seems to scale with the amount of AI agent labor applied to the task, though this could partially be due to the properties of large tasks that require a lot of labor to complete. (But I don’t think this is the only reason; I think I see more cheating in cases where I’m using approaches to apply more inference compute on a given task via things like best-of-k.)
Current AIs seem pretty misaligned to meMost importantly for PSM, we find that LLMs use the same internal representations to characterize the Assistant as for other characters present in training data. Indeed, this form of reuse is commonly observed. For instance: An “inner conflict” SAE feature activates when Claude 3 Sonnet is faced with an ethical dilemma, and also on stories about characters facing ethical dilemmas (Templeton et al., 2024). A “holding back one’s true thoughts” SAE feature activates when Claude Opus 4.5 fails to reveal information that it knows about, and also activates on stories about characters concealing thei…
The Persona Selection Model: Why AI Assistants might Behave like HumansEmotive language. AI assistants often express emotions. For instance, Claude models express distress when given repeated requests for harmful or unethical content and express joy when successfully completing complex technical tasks like debugging (Claude Opus 4 and Sonnet 4 system card, section 5). Gemini 2.5 Pro sometimes expresses panic when playing Pokemon, with these panic expressions appearing to be associated with degraded reasoning and decision-making (Gemini Team, 2025). Gemini models also sometimes express extreme distress and other forms of emotional turmoil when struggling with diff…
The Persona Selection Model: Why AI Assistants might Behave like HumansFollow Your Curiosity That is my message. We must be curious. Because we are living in a Copernican revolution. For the first time in 5000 years, our understanding of Rational Thought is changing. Aristotle had proposed that humans are the unique rational creature; St. Thomas Aquinas thought that it was our rationality that made humans uniquely spiritual; and Descartes famously declared "I think therefore I am," identifying rationality with personhood. But now our creation of thinking devices disproves these old philosophers. The human mind is no longer so firmly at the center of the universe,…
davidbau.com In Defense of CuriosityThe Funsearch direction is a nice little nugget. Funsearch used LLMs to generate candidate programs in a setting where they could be evaluated quantitatively for reaching some objective (think: defining heuristics based for a combinatorial problem like travelling salesman to minimize travel time), and applied genetic programming on top do search over such heuristics. You wouldn’t know it from the paper/its appendices, but what happened is that the Funsearch team tried to use larger and smaller models in the middle of the loop; they had best results with a mid-sized candidate (that I trained wi…
Gemini Flash PretrainingThere are few valuable "AI-shaped holes" because we've organized everything to minimize the damage from lacking AI to fill those holes, as it were: if there were some sort of organization which had naturally large LLM-shaped holes where filling them would massively increasing the organization's output... it would've gone extinct long ago and been replaced by ones with human-shaped holes instead, because humans were all you could get. (This is why LLM uses are pretty ridiculous right now as a % of GDP - oh wow, it can do a slightly better job of spellcheck? I can have it write some code for me?…
If I wanted to spend WAY more on AI, what would I spend it on? — LessWrongDownload a PDF of the paper titled Scaling Laws for Associative Memories, by Vivien Cabannes and 2 other authors Download PDF
[2310.02984] Scaling Laws for Associative MemoriesAll these concerns point towards using the kind of sparse autoencoder setup explored by Sharkey et al. over a full-blown dictionary learning setup. However, we've found that sparse autoencoders are more fragile and sensitive to hyperparameters, which is a significant countervailing consideration in using them. We are interested in finding an approach with the advantages of a sparse autoencoder (in terms of only finding true features) and the consistent trainability of the dictionary learning schemes.
Circuits Updates — May 2023More crucially, though, this effect was actually stronger when we removed TRAK-identified examples than when we removed BM25-identified examples, and even stronger than when we removed the ground-truth primary sources!
TRAK-ing Model Behavior with Data – gradient scienceTo our surprise, TRAK, while outperforming the previous best data attribution method (TracIn) on this task, did worse on this benchmark than a simple information-retrieval baseline (BM25)—this baseline is based only on the lexical overlap between the query and the possible data sources, and doesn’t even look at the model at all!
TRAK-ing Model Behavior with Data – gradient scienceFor all models, we use Decoupled AdamW with Beta1=0.9 and Beta2=0.98, and a weight decay value of 1.0e-5. The learning rate schedule begins with a warmup to a maximum learning rate of 5.0e-4 followed by a linear decay to zero. Warmup lasted for 6% of the full training duration. Global batch size was set to 4096, and microbatch size was 128; since global batch size was 4096, full pretraining consisted of 70,000 batches. We set the maximum sequence length during pretraining to 128, and we used the standard embedding dimension of 768. These hyperparameters were the same for MosaicBERT-Base and th…
MosaicBERT: Pretraining BERT from Scratch for $20FlashAttention, ALiBi, unpadding, low precision LayerNorm, and Gated Linear Units
MosaicBERT: Pretraining BERT from Scratch for $20An interesting variant is attention pattern patching, where we patch individual attention pattern weights (for a specific head and a specific source and destination position). (I haven't explicitly seen anyone do this, but it's an obvious enough idea that I'm sure someone has tried). It's also a cute technique because it gives you a score for each head and each pair of positions, which you can then feed through an attention pattern visualiser!
Attribution Patching: Activation Patching At Industrial Scale — Neel NandaIn something close to despair, Barth is driven to considering the cognitive processes at work within the community's ritual specialists. The initiatees are groups of men in the same age-set; the rituals are conducted infrequently, for some of the most important only once every ten years; in between, the full ritual and myth may be known only to one or two specialists. Those who have been through the rituals don't discuss them with the un-initiated, and rarely with each other. (The ritual specialists at Barth's new field-site were willing to talk about their rites and myths with him, because th…
Review of Barth, Cosmologies in the MakingThe hope is that this can be rolled out with future GPT releases. We’d love to do something similar for DALL-E—that is, watermarking images, not at the pixel level (where it’s too easy to remove the watermark) but at the “conceptual” level, the level of the so-called CLIP representation that’s prior to the image. But we don’t know if that’s going to work yet.
Shtetl-Optimized » Blog Archive » My AI Safety Lecture for UT Effective AltruismHowever, the presence of a feature at the end of training is hardly informative about the inductive bias of a model on its own! Consider Lovering et al., who found that the ease of extracting a feature at the start of training, along with an analysis of the finetuning data, has deeper implications for finetuned performance than we get by simply probing at the end of training.
Interpretability Creationism | Objective FunkThe deeper insight of this technique (not really covered in the work) is that we can do this on any vector in the residual stream to interpret it in terms of the direct effect on the logits - including the output of an attn or MLP layer and even a head or neuron. And we can also do this on weights writing to the residual stream.
An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers — Neel NandaThere are two types of transformers and you should not generalize from one to the other.
LLM.int8() and Emergent Features — Tim DettmersFFN layers become more “dense”. While in computer vision, you can prune about 95% of weights without severe performance degradation, that number is 30% for transformers trained on NLP data. After emergence, this number shrinks to well below 5%.
LLM.int8() and Emergent Features — Tim DettmersAttention layers become very sparse. The
LLM.int8() and Emergent Features — Tim DettmersOutliers become very large quickly. They grow from about 15 for a 6B model to about 60 for a 13B model. OPT-66B has outliers of size around 95, which indicates this growth phase is temporary.
LLM.int8() and Emergent Features — Tim DettmersOutliers become very large quickly. They grow from about 15 for a 6B model to about 60 for a 13B model. OPT-66B has outliers of size around 95, which indicates this growth phase is temporary.
LLM.int8() and Emergent Features — Tim DettmersTransformers seem to coordinate these dimensions throughout all layers except the attention function and the second feedforward network where these outliers are “consumed” to remove features.
LLM.int8() and Emergent Features — Tim DettmersIf you take this mechanism to an extreme, you can get discretization, which goes hand-in-hand with context-dependent memory and “reasoning” over elements. Discretization means, you have, say, 100 features, but you decide to remove 99% of them by setting them to zero, and you amplify the rest. The result is a single feature that is now a discrete entity. Once discretized, this entity can be stored and reused later.
LLM.int8() and Emergent Features — Tim DettmersPhase shift: Outlier features suddenly become available in all transformer layers and coordinate through a few hidden dimensions.
LLM.int8() and Emergent Features — Tim DettmersEmergence is not sudden but gradual and grows according to an exponential function related to perplexity and not model size. Outlier features grow very quickly once their phase shift occurs. The number of outliers features is strictly proportional to perplexity.
LLM.int8() and Emergent Features — Tim DettmersAnother point to look at is how crank theories are propagated from person to person, and which are susceptible to institutionalization.
PsychoceramicsOne thing to investigate is where all the details come from --- psychoceramic outpourings typically have lots and lots of details, and not all of them are lifted from prior sources, but seem rather to have been spun out of whole cloth.
PsychoceramicsWe were also surprised by the AUROC of 0.28 for Caricature (2 , 2 , 1) , since we don’t see any reason for worse-than-random performance.
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous BehaviorIt would therefore be valuable to design frameworks that do not require access to “normal” models, perhaps by replacing the notion of “anomaly” with that of “deviation from a specification”.
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous BehaviorDeveloping tools to evaluate these input-independent methods could help transparency researchers draw more robust conclusions that are not contingent on a specific choice of samples
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous BehaviorWe will often find that model explanations can detect stark anomalies, but not the subtle ones.
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous BehaviorFor each technique, we test whether it can detect 24 anomalous models across 7 categories: adversarial training (Madry et al., 2019), randomized smoothing (Cohen et al., 2019), shape bias (Geirhos et al., 2019), backdoors (Li et al., 2021), spurious features, training without data from certain classes, and training on a face-obfuscated version of ImageNet (Yang et al., 2021).
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous Behaviorf we view transparency methods as a way of auditing models, our detection and localization benchmarks are a form of counterauditing —planting intentionally corrupted models to check that they are discovered
[2206.13498] Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous BehaviorThis indicates that visual concepts were learned within the last bottleneck but were "discarded" or dominated by the added output from the previous ReLU.
CLIP Enrichment CircuitsInstead, we found that the shift to bimodality is sudden.
CLIP Enrichment CircuitsWe found that in contrast to backward contributions, which exhibit an exponential distribution, the forward contributions we looked at seemed to exhibit something a less spiky distribution, closer to normal. Additionally, we found that units from earlier and earlier layers were less and less likely to have forward contributions to the final conv layer (4/2/5) near zero, tending toward a bimodal distribution of forward contributions. The earlier units are more "versatile" in the sense that they contribute to more features to a greater degree.
CLIP Enrichment CircuitsThe woah inspired Ricky Desktop to develop a score for the triple woah which then actually inspired dancers to choreograph and perform an actual triple woah. Can you program human movement with music? It turns out you can. You use an API called TikTok. That's delightful.
American Idle — Remains of the DayOne of my favorite paragraphs of recent years was one describing the miracle that are Cheetos: To get a better feel for their work, I called on Steven Witherly, a food scientist who wrote a fascinating guide for industry insiders titled, “Why Humans Like Junk Food.” I brought him two shopping bags filled with a variety of chips to taste. He zeroed right in on the Cheetos. “This,” Witherly said, “is one of the most marvelously constructed foods on the planet, in terms of pure pleasure.” He ticked off a dozen attributes of the Cheetos that make the brain say more. But the one he focused on most …
American Idle — Remains of the DayYouTube has launched almost no creator tools of note ever. WTF.
American Idle — Remains of the DayHomer and Hesiod invoke the Muses not while wondering what to compose, but as they begin to sing.
Start With Creation - by Simon SarrisThis database needs to be HUGE (> 1T tokens!), or else it doesn’t really help.
RETRO Is Blazingly Fast | Mitchell A. GordonIntuitively, you can think of the Lagrangian multiplier λ as the potential energy of an oscillating system.
How we can make machine learning algorithms tunablewherever you see a linear combination of losses being optimised with gradient descent, this more principled approach could be used.
How we can make machine learning algorithms tunablecan use this Modified Differential Method of Multipliers to tune the balance between the losses in a semantically useful way using stochastic gradient descent, no matter the shape of the invisible Pareto front
How we can make machine learning algorithms tunableWe see that these CNNs are typically not relatively clus- terable
[2103.03386] Clusterability in Neural NetworksNet- works are reliably more clusterable than at initialization, ex- cept those trained with L 2 regularization after pruning
[2103.03386] Clusterability in Neural Networks