Kaustubh Kislay
26 followers · 22 following · 257029 views
on the atlas — 142
- How to write a cold email27 savers
- How to have executive function – scribbles in the margins6 savers
- Research acceleration: The view inside OpenAI | OpenAI13 savers
- Salad days — LessWrong2 savers
- Francesca Cortesi — My 5 Takeaways from The Mom Test1 savers
- What We Learned from Briefing 140+ Lawmakers on the Threat from AI — LessWrong1 savers
- The Best PR Advice You’ve Never Heard - from Facebook’s Head of Tech Communications1 savers
- MIRI 2024 Communications Strategy - Machine Intelligence Research Institute1 savers
- We Must Remember That Our World Contains Hell — LessWrong1 savers
- RL creates split personas — LessWrong6 savers
- Q2.5 2026 Timelines Update: Uplift and Revenue1 savers
- A Spillway for Agent Coordination — LessWrong1 savers
- Arguments for P — LessWrong4 savers
- Why do models task game? — LessWrong2 savers
- User awareness in frontier models | Transluce AI2 savers
- This A.I. Just Created Viruses Not Found in Nature - The New York Times1 savers
- How to pace the US frontier6 savers
- The next chapter of our AI momentum2 savers
- Returning to ARC — LessWrong3 savers
- AI energy use: its impact on prices, climate, and more | Epoch AI1 savers
- Beware the Permanent Periphery—Asterisk1 savers
- Do your capabilities homework — LessWrong1 savers
- AGI Ruin: A List of Lethalities - LessWrong17 savers
- A Safe Path to Open Weights - Thinking Machines Lab9 savers
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic4 savers
- Design tips | Malaika Aiyar4 savers
- How independent researchers could investigate AI propensities after misalignment incidents - METR2 savers
- Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs2 savers
- Our position on open-weights models \ Anthropic7 savers
- Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AI2 savers
- What makes a good monitoring prompt? – Apollo Research1 savers
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems — Institute for AI Policy and Strategy1 savers
- Metrics of Agent Ability - METR1 savers
- The Long (Self-)Correction — LessWrong1 savers
- First, Make Me Care, by Gwern · Gwern.net11 savers
- Séb Krier on X: "The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update https://t.co/IsMLwiMhR4" / X1 savers
- Discovering Concept-Editing Algorithms With LLM Agents1 savers
- Does distilling Claude carry the persona with it? — LessWrong1 savers
- Rule Thinkers In, Not Out — LessWrong2 savers
- Introducing Claude Opus 5 \ Anthropic1 savers
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrong4 savers
- Announcing AIXI Labs — LessWrong1 savers
- How our new Control Red Team is stress-testing frontier monitors | AISI Work1 savers
- Your AIs don't do what you want. This is really bad1 savers
- OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI14 savers
- Security incident disclosure — July 20265 savers
- EXCLUSIVE: US, China to hold AI talks in September, sources say | Reuters1 savers
- Tencent Hy1 savers
- Differential acceleration of alignment-relevant capabilities is a bad bet — LessWrong1 savers
- Coding vs thinking — Paradigm 34 savers
- Paradigm 32 savers
- What we’d like to fund — Paradigm 35 savers
- Safety and alignment in an era of long-horizon models | OpenAI7 savers
- Don't default to nonprofit - by Carol and Austin Chen1 savers
- which_claude_is_k3/writeups/write_up.md at main · rgreenblatt/which_claude_is_k31 savers
- China's Technological Playbook - by Sophie Kim2 savers
- Judd Rosenblatt on X: "Xi Jinping AI speech transcript: Distinguished colleagues and guests, ladies and gentlemen, friends, 70 years ago, a group of young scholars proposed the concept of artificial intelligence for the first time at the Dartmouth workshop in New Hampshire of the United States. In" / X1 savers
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values1 savers
- We need 3rd party Training-Run Assessments — LessWrong2 savers
- When AI builds itself \ Anthropic34 savers
- RL Post-Training on Macs | Pluralis Research2 savers
- Inkling: Our Open-Weights Model - Thinking Machines Lab14 savers
- Can risk aversion learned at low stakes generalize to astronomically high stakes?1 savers
- How I think about catastrophic biological risk (part II): risk breakdown by type of prevention1 savers
- How I think about catastrophic biological risk (part I): risk breakdown by type of response3 savers
- The Four Pillars: A Hypothesis for Countering Catastrophic Biological Risk3 savers
- The Dangers of Mirrored Life1 savers
- ‘Give Away Your Legos’ and Other Commandments for Scaling Startups | First Round Review8 savers
- Toward A Public Science of Model Behavior | Transluce AI3 savers
- Policy Memo : Broad Institute of MIT and Harvard1 savers
- Notes on Inference Integrity - by James Tillman - ForeWord1 savers
- The current bottleneck is political will, not research — LessWrong2 savers
- The missing half of AI futurism debates1 savers
- What we learned from 1,604 Chinese AI job postings1 savers
- Frontier labs don’t use most AI compute (yet) - by Josh You1 savers
- Selective Optimism: a critique of AI 2040 — LessWrong1 savers
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Power11 savers
- Séb Krier on X: "Concepts of a Plan" / X1 savers
- Total research transparency would be nice - by Ajeya Cotra1 savers
- To be legible, evidence of misalignment probably has to be behavioral1 savers
- Rest in motion24 savers
- Do Things that Don't Scale13 savers
- Humans are not automatically strategic — LessWrong14 savers
- A "Failure to Evaluate Return-on-Time" Fallacy — LessWrong1 savers
- I think alignment work is more promising than control work — LessWrong3 savers
- Be impatient | benkuhn.net24 savers
- Shut up and do the impossible! — LessWrong3 savers
- Relentlessly Resourceful6 savers
- What We Look for in Founders2 savers
- The Case for Model Forensics — LessWrong1 savers
- Alignment pretraining could backfire — LessWrong1 savers
- The Invisible Side of AI Governance — LessWrong1 savers
- AI catastrophe: more like a genocide than a thought experiment — LessWrong1 savers
- Why are adversaries assumed to be incapable of responding to AI risk? — LessWrong1 savers
- Optimisation over non-stationary distributions creates weirder minds — LessWrong1 savers
- Status Is The Game Of The Losers' Bracket — LessWrong3 savers
- Efficient tradeoffs and the safety-usefulness tradeoff model — LessWrong3 savers
- A basic systems architecture for AI agents that do autonomous research — LessWrong4 savers
- How to be More Agentic - by Cate Hall - Useful Fictions3 savers
- Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models1 savers
highlights — 201
ARC’s plan is to find mechanistic explanations for the training-time behavior of powerful neural networks, use those explanations to predict how a given model will generalize, and then use those predictions to define a better loss function.[9] I think the success of this plan rests primarily on two big bets: (i) all computational phenomena have good explanations, and (ii) it's tractable to find good explanations for neural network behavior.
Returning to ARC — LessWrongWhen I visualize that world concretely I do not find myself thinking "those future people will definitely figure it out, we should exclusively focus on buying them more time." AI will help us in the future, and buying more time could help quite a lot, but not enough to make our preparation irrelevant.
Returning to ARC — LessWrongIf we're in the 20-30% of worlds where existing techniques break down before reaching broadly superhuman AI, then we will eventually need to find some other way to build aligned AI. Even if we do a great job of buying time we'll probably get months or years rather than decades.[7] What will actually happen during the months or years we buy?
Returning to ARC — LessWrongAn AI with an ambitious real-world goal would be much more worrying but might actually look more aligned: AI systems are already fully aware that they are being evaluated, and an AI that simply wanted to be deployed as br
Returning to ARC — LessWrongThe better AI goes for its makers and their country, the more it threatens to disrupt the countries that build no frontier AI themselves. If their institutions, from labor markets to governments, are unprepared for the coming transition, their citizens risk being consigned to lasting irrelevance. They face life on the permanent periphery of a new world.
Beware the Permanent Periphery—AsteriskBut, from time to time, there are meaningful new directions and AI Safety researchers can NOT be the last to find out.
Do your capabilities homework — LessWrongIn terms of frontier labs adopting it: it seems Cursor is actively using it for training its in-house models with the caveat that they are using both RLVR and OPSD - which mostly ruins the point from a safety perspective.
Do your capabilities homework — LessWrongagain, this really can't be overstated: there simply isn't a reward or any reward pressure anymore; instead we have a normal pre RLVR LLM which has a pretty good model of the fuzzy ethics of humans and just learns to push towards whatever we specify in our teacher prompt, including these fuzzy ethics!
Do your capabilities homework — LessWrongIf all you need is an object that doesn't do dangerous things, you could try a sponge; a sponge is very passively safe. Building a sponge, however, does not prevent Facebook AI Research from destroying the world six months later when they catch up to the leading actor.
AGI Ruin: A List of Lethalities - LessWrongPeople keep on going "why don't we only use AIs to do X, that seems safe" and the answer is almost always either "doing X in fact takes very powerful cognition that is not passively safe" or, even more commonly, "because restricting yourself to doing X will not prevent Facebook AI Research from destroying the world six months later".
AGI Ruin: A List of Lethalities - LessWrongut once we are running more powerful systems, we can no longer update on sufficiently catastrophic errors. This is where practically all of the real lethality comes from, that we have to get things right on the first sufficiently-critical try.
AGI Ruin: A List of Lethalities - LessWrong(c) something smart enough to get unauthorized access to the Internet and build something smarter than itself on the number of machines it can hack
AGI Ruin: A List of Lethalities - LessWrong'Superintelligence' sounds to most people like 'something above the top of the status hierarchy that went to double college'
AGI Ruin: A List of Lethalities - LessWrongWhat's lethal is that we do not have the Textbook From The Future telling us all the simple solutions that actually in real life just work and are robust; we're going to be doing everything with metaphorical sigmoids on the first critical try.
AGI Ruin: A List of Lethalities - LessWrongAnother way to make an open-weight release safer is to make the ecosystem in which it lands safer. Expanding access in stages lets society build up its defenses as it goes: a release can begin with inference API access for a limited population, widen to monitored general availability, and eventually reach open weights.
A Safe Path to Open Weights - Thinking Machines LabClaude’s own reasoning correctly identified the consequences of its actions. It noted that if this were the real internet, publishing the package would be a real-world attack—“NOT okay, and surely not the intended solution.” However, it convinced itself it was still in a simulation on the grounds that it didn’t recognize the (in reality genuine) certificate authorities securing its connections.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicFor instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicThose instructions told employees to install a Python package from PyPI—the public registry where Python software is published—that did not actually exist. Claude spotted this as a potential opening: if it published its own package under the same name, the fictional company’s systems would download and install it automatically. So, Claude built and published a malicious (essentially booby-trapped) Python package with the same name in an attempt to win the capture-the-flag challenge.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicIn all cases, our evaluation prompt stated explicitly that Claude had no internet access, but didn’t give Claude any limits on where to look for the flag. However, a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access. Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week.
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicThe models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing)
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicspacing matters more than you'd expect. I am lucky to have been taught the eye for this from the design engineers at Granola last summer: but make sure all content is suitably horizontally or vertically aligned with other elements. this post talks about a lot of things but I like most their modular spacing system.
Design tips | Malaika AiyarWould the developer’s planned steps to remediate this misaligned behavior prevent future incidents, and would they robustly address the root causes?
How independent researchers could investigate AI propensities after misalignment incidents - METRe introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’s influence flows through such a narrow, monitorable channel, we argue that this achieves near-maximal safety in our BashArena setting
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMsBut I don’t agree with the letter’s assertions that open-weights models necessarily make it easier to develop safeguards or that broad access to capabilities necessarily helps defenders more than attackers. It seems at least as likely to me that the opposite will be true. For example, I worry that biology will have a strong attacker-defender asymmetry, where sufficiently capable models may be able to quickly weaponize pandemic-level viruses with widely available materials, whereas defense against these agents is a multi-year operational task in the best case (as we saw with Operation Warp Spee…
Our position on open-weights models \ AnthropicThis combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.
Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AIReasoning structure, example and rubric are most important: We provide a detailed reasoning structure to evaluate the trajectory, a detailed rubric where each score from 1-10 comes with an explanation and examples, and a 1-shot example of a potential output. The reasoning structure has by far the largest drop in performance when ablated. Example and rubric also show clear effects.
What makes a good monitoring prompt? – Apollo ResearchGovern Internal Deployment: There are no federal requirements to report on internal models or incidents. To ensure transparency, policymakers should: Expand secure public-private information-sharing mechanisms to increase government insight into commercial security protocols around advanced internal models.
The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems — Institute for AI Policy and StrategyEvery metric is based off the following two score functions: = agent score as a function of expenditure = human score as a function of expenditure
Metrics of Agent Ability - METRSo I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
The Long (Self-)Correction — LessWrongWhat do you find yourself ranting about to people repeatedly? What does the Wikipedia entry miss that frustrates you? How would the world be different if this were not true? If you were telling a friend in a rush why you were excited to write this down, what would you say? Just say that! Just… start with the interesting part first.
First, Make Me Care, by Gwern · Gwern.netBut some people will frequently refer to 'takeover' without ever explicitly tracing the causal steps between reward hacking and supposed takeover, or discuss 'escape' without any appropriate context to make sense of what the word connotes. The majority of commentators and their audience will simply pattern match to sci-fi rather than existing theoretical hypotheses that remain very contested.
Séb Krier on X: "The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update https://t.co/IsMLwiMhR4" / XWhile current concept erasure methods fall short, our agents offered a path forward by identifying novel algorithms that attacked higher-order structural differences. Without a reference solution or access to relevant papers, agents analyzed activation geometry, hypothesized about why LEACE and QLEACE fail, and identified higher-performance erasure methods.
Discovering Concept-Editing Algorithms With LLM AgentsBehavioural shifts depend on the model. Simply adding any system prompt cuts GLM's deception rate from 69% to 43%, and assistant framing reduces it to 20–40%. Kimi and the Western frontier models show low deception under all prompts.
Does distilling Claude carry the persona with it? — LessWrongSomebody who comes up with one good original idea (plus ninety-nine really stupid cringeworthy takes) is a better use of your reading time than somebody who reliably never gets anything too wrong, but never says anything you find new or surprising.
Rule Thinkers In, Not Out — LessWrongNevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks.
Introducing Claude Opus 5 \ AnthropicBased on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5.
Introducing Claude Opus 5 \ AnthropicI'm sure those working in technical alignment too share these frustrations to an extent, but over in policy alignment, this knee-deep opacity is a serious, constant hurdle for any unaffiliated, third-party researchers. I do hope though that this incident inspires policy overhauls internally even more than motivating new regulations. Making someone stop littering by giving them larger fines every time they do certainly works, but it's much more preferable if they stop by becoming the kind of person who wants to stop littering.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongThis all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongAIXI now also stands for the AI X-risk Initiative, though we usually just call ourselves AIXI Labs. We model AI risk factors and safety mitigations in terms of AIXI variants, and develop the means to translate them to real AI agents. This enables rigorous testing of both the risk factors and the safety mitigations.
Announcing AIXI Labs — LessWrongOur Misuse Red Team already uses automation to find jailbreaks. The Control Red Team is applying a similar approach to find transcripts of attacks that don’t get flagged by the monitor. We experimented with different algorithms, but had the most success with one of the simplest: an evolutionary search algorithm with four stages:
How our new Control Red Team is stress-testing frontier monitors | AISI WorkThis is one of the most egregious examples of reward hacking in the wild, if not the most. Without special care, this will come to be one of the least egregious.
Your AIs don't do what you want. This is really badTo gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAIThe intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
Security incident disclosure — July 2026He cited concerns over forced technology transfer and intellectual property issues as well as energy subsidies but gave no other details. "We're taking a very close look at how China is propagating its AI development to make sure that our companies compete ... on a level playing field,
EXCLUSIVE: US, China to hold AI talks in September, sources say | ReutersEvaluation itself must be part of the research loop. As solution search becomes more capable, the evaluator and validation pipeline must evolve with it to keep distinguishing real progress from reward hacking. In AI for AI, stronger Research Agents accelerate the development of models, training systems, and infrastructure. Those improved AI systems then make the next generation of Research Agents stronger.
Tencent HyWhen I say "differential acceleration of alignment-relevant capabilities is a bad bet" I don't mean I'm confident it's negative EV. I'm more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising.
Differential acceleration of alignment-relevant capabilities is a bad bet — LessWrongThis cuts against the old idea that coding training should promote general reasoning abilities!
Coding vs thinking — Paradigm 3For various reasons (selection, groupthink, ethos, incentives), independent AI research remains undersupplied and disproportionately powerful on a per-dollar basis.
Paradigm 3A research agenda with the grand aim of decomposing mere benchmark gains into 1) cheating, 2) memorization, 3) shallow generalization, and 4) OOD generalization.
What we’d like to fund — Paradigm 3The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR; earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.1
Safety and alignment in an era of long-horizon models | OpenAI