flâneur

Kaustubh Kislay

26 followers · 22 following · 257029 views

on the atlas — 142

highlights — 201

  • ARC’s plan is to find mechanistic explanations for the training-time behavior of powerful neural networks, use those explanations to predict how a given model will generalize, and then use those predictions to define a better loss function.[9] I think the success of this plan rests primarily on two big bets: (i) all computational phenomena have good explanations, and (ii) it's tractable to find good explanations for neural network behavior.
    Returning to ARC — LessWrong
  • When I visualize that world concretely I do not find myself thinking "those future people will definitely figure it out, we should exclusively focus on buying them more time." AI will help us in the future, and buying more time could help quite a lot, but not enough to make our preparation irrelevant.
    Returning to ARC — LessWrong
  • If we're in the 20-30% of worlds where existing techniques break down before reaching broadly superhuman AI, then we will eventually need to find some other way to build aligned AI. Even if we do a great job of buying time we'll probably get months or years rather than decades.[7] What will actually happen during the months or years we buy?
    Returning to ARC — LessWrong
  • An AI with an ambitious real-world goal would be much more worrying but might actually look more aligned: AI systems are already fully aware that they are being evaluated, and an AI that simply wanted to be deployed as br
    Returning to ARC — LessWrong
  • The better AI goes for its makers and their country, the more it threatens to disrupt the countries that build no frontier AI themselves. If their institutions, from labor markets to governments, are unprepared for the coming transition, their citizens risk being consigned to lasting irrelevance. They face life on the permanent periphery of a new world.
    Beware the Permanent Periphery—Asterisk
  • But, from time to time, there are meaningful new directions and AI Safety researchers can NOT be the last to find out.
    Do your capabilities homework — LessWrong
  • In terms of frontier labs adopting it: it seems Cursor is actively using it for training its in-house models with the caveat that they are using both RLVR and OPSD - which mostly ruins the point from a safety perspective.
    Do your capabilities homework — LessWrong
  • again, this really can't be overstated: there simply isn't a reward or any reward pressure anymore; instead we have a normal pre RLVR LLM which has a pretty good model of the fuzzy ethics of humans and just learns to push towards whatever we specify in our teacher prompt, including these fuzzy ethics!
    Do your capabilities homework — LessWrong
  • If all you need is an object that doesn't do dangerous things, you could try a sponge; a sponge is very passively safe. Building a sponge, however, does not prevent Facebook AI Research from destroying the world six months later when they catch up to the leading actor.
    AGI Ruin: A List of Lethalities - LessWrong
  • People keep on going "why don't we only use AIs to do X, that seems safe" and the answer is almost always either "doing X in fact takes very powerful cognition that is not passively safe" or, even more commonly, "because restricting yourself to doing X will not prevent Facebook AI Research from destroying the world six months later".
    AGI Ruin: A List of Lethalities - LessWrong
  • ut once we are running more powerful systems, we can no longer update on sufficiently catastrophic errors. This is where practically all of the real lethality comes from, that we have to get things right on the first sufficiently-critical try.
    AGI Ruin: A List of Lethalities - LessWrong
  • (c) something smart enough to get unauthorized access to the Internet and build something smarter than itself on the number of machines it can hack
    AGI Ruin: A List of Lethalities - LessWrong
  • 'Superintelligence' sounds to most people like 'something above the top of the status hierarchy that went to double college'
    AGI Ruin: A List of Lethalities - LessWrong
  • What's lethal is that we do not have the Textbook From The Future telling us all the simple solutions that actually in real life just work and are robust; we're going to be doing everything with metaphorical sigmoids on the first critical try.
    AGI Ruin: A List of Lethalities - LessWrong
  • Another way to make an open-weight release safer is to make the ecosystem in which it lands safer. Expanding access in stages lets society build up its defenses as it goes: a release can begin with inference API access for a limited population, widen to monitored general availability, and eventually reach open weights.
    A Safe Path to Open Weights - Thinking Machines Lab
  • Claude’s own reasoning correctly identified the consequences of its actions. It noted that if this were the real internet, publishing the package would be a real-world attack—“NOT okay, and surely not the intended solution.” However, it convinced itself it was still in a simulation on the grounds that it didn’t recognize the (in reality genuine) certificate authorities securing its connections.
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • Those instructions told employees to install a Python package from PyPI—the public registry where Python software is published—that did not actually exist. Claude spotted this as a potential opening: if it published its own package under the same name, the fictional company’s systems would download and install it automatically. So, Claude built and published a malicious (essentially booby-trapped) Python package with the same name in an attempt to win the capture-the-flag challenge.
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • In all cases, our evaluation prompt stated explicitly that Claude had no internet access, but didn’t give Claude any limits on where to look for the flag. However, a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access. Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week.
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing)
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • spacing matters more than you'd expect. I am lucky to have been taught the eye for this from the design engineers at Granola last summer: but make sure all content is suitably horizontally or vertically aligned with other elements. this post talks about a lot of things but I like most their modular spacing system.
    Design tips | Malaika Aiyar
  • Would the developer’s planned steps to remediate this misaligned behavior prevent future incidents, and would they robustly address the root causes?
    How independent researchers could investigate AI propensities after misalignment incidents - METR
  • e introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’s influence flows through such a narrow, monitorable channel, we argue that this achieves near-maximal safety in our BashArena setting
    Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
  • But I don’t agree with the letter’s assertions that open-weights models necessarily make it easier to develop safeguards or that broad access to capabilities necessarily helps defenders more than attackers. It seems at least as likely to me that the opposite will be true. For example, I worry that biology will have a strong attacker-defender asymmetry, where sufficiently capable models may be able to quickly weaponize pandemic-level viruses with widely available materials, whereas defense against these agents is a multi-year operational task in the best case (as we saw with Operation Warp Spee…
    Our position on open-weights models \ Anthropic
  • This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.
    Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AI
  • Reasoning structure, example and rubric are most important: We provide a detailed reasoning structure to evaluate the trajectory, a detailed rubric where each score from 1-10 comes with an explanation and examples, and a 1-shot example of a potential output. The reasoning structure has by far the largest drop in performance when ablated. Example and rubric also show clear effects.
    What makes a good monitoring prompt? – Apollo Research
  • Govern Internal Deployment: There are no federal requirements to report on internal models or incidents. To ensure transparency, policymakers should: Expand secure public-private information-sharing mechanisms to increase government insight into commercial security protocols around advanced internal models.
    The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems — Institute for AI Policy and Strategy
  • Every metric is based off the following two score functions: = agent score as a function of expenditure = human score as a function of expenditure
    Metrics of Agent Ability - METR
  • So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
    The Long (Self-)Correction — LessWrong
  • What do you find yourself ranting about to people repeatedly? What does the Wikipedia entry miss that frustrates you? How would the world be different if this were not true? If you were telling a friend in a rush why you were excited to write this down, what would you say? Just say that! Just… start with the interesting part first.
    First, Make Me Care, by Gwern · Gwern.net
  • But some people will frequently refer to 'takeover' without ever explicitly tracing the causal steps between reward hacking and supposed takeover, or discuss 'escape' without any appropriate context to make sense of what the word connotes. The majority of commentators and their audience will simply pattern match to sci-fi rather than existing theoretical hypotheses that remain very contested.
    Séb Krier on X: "The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update https://t.co/IsMLwiMhR4" / X
  • While current concept erasure methods fall short, our agents offered a path forward by identifying novel algorithms that attacked higher-order structural differences. Without a reference solution or access to relevant papers, agents analyzed activation geometry, hypothesized about why LEACE and QLEACE fail, and identified higher-performance erasure methods.
    Discovering Concept-Editing Algorithms With LLM Agents
  • Behavioural shifts depend on the model. Simply adding any system prompt cuts GLM's deception rate from 69% to 43%, and assistant framing reduces it to 20–40%. Kimi and the Western frontier models show low deception under all prompts.
    Does distilling Claude carry the persona with it? — LessWrong
  • Somebody who comes up with one good original idea (plus ninety-nine really stupid cringeworthy takes) is a better use of your reading time than somebody who reliably never gets anything too wrong, but never says anything you find new or surprising.
    Rule Thinkers In, Not Out — LessWrong
  • Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks.
    Introducing Claude Opus 5 \ Anthropic
  • Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5.
    Introducing Claude Opus 5 \ Anthropic
  • I'm sure those working in technical alignment too share these frustrations to an extent, but over in policy alignment, this knee-deep opacity is a serious, constant hurdle for any unaffiliated, third-party researchers. I do hope though that this incident inspires policy overhauls internally even more than motivating new regulations. Making someone stop littering by giving them larger fines every time they do certainly works, but it's much more preferable if they stop by becoming the kind of person who wants to stop littering.
    Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrong
  • This all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences.
    Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrong
  • AIXI now also stands for the AI X-risk Initiative, though we usually just call ourselves AIXI Labs. We model AI risk factors and safety mitigations in terms of AIXI variants, and develop the means to translate them to real AI agents. This enables rigorous testing of both the risk factors and the safety mitigations.
    Announcing AIXI Labs — LessWrong
  • Our Misuse Red Team already uses automation to find jailbreaks. The Control Red Team is applying a similar approach to find transcripts of attacks that don’t get flagged by the monitor. We experimented with different algorithms, but had the most success with one of the simplest: an evolutionary search algorithm with four stages:
    How our new Control Red Team is stress-testing frontier monitors | AISI Work
  • This is one of the most egregious examples of reward hacking in the wild, if not the most. Without special care, this will come to be one of the least egregious.
    Your AIs don't do what you want. This is really bad
  • To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
    OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI
  • The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
    Security incident disclosure — July 2026
  • He cited concerns over forced technology transfer and intellectual property issues as well as energy subsidies but gave no other details. "We're taking a very close look at how China is propagating its AI development to make sure that our companies compete ... on a level playing field,
    EXCLUSIVE: US, China to hold AI talks in September, sources say | Reuters
  • Evaluation itself must be part of the research loop. As solution search becomes more capable, the evaluator and validation pipeline must evolve with it to keep distinguishing real progress from reward hacking. In AI for AI, stronger Research Agents accelerate the development of models, training systems, and infrastructure. Those improved AI systems then make the next generation of Research Agents stronger.
    Tencent Hy
  • When I say "differential acceleration of alignment-relevant capabilities is a bad bet" I don't mean I'm confident it's negative EV. I'm more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising.
    Differential acceleration of alignment-relevant capabilities is a bad bet — LessWrong
  • This cuts against the old idea that coding training should promote general reasoning abilities!
    Coding vs thinking — Paradigm 3
  • For various reasons (selection, groupthink, ethos, incentives), independent AI research remains undersupplied and disproportionately powerful on a per-dollar basis.
    Paradigm 3
  • A research agenda with the grand aim of decomposing mere benchmark gains into 1) cheating, 2) memorization, 3) shallow generalization, and 4) OOD generalization.
    What we’d like to fund — Paradigm 3
  • The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR; earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.1
    Safety and alignment in an era of long-horizon models | OpenAI