flâneur

Vincent Cheng

99 followers · 80 following · 6252 views

on the atlas — 422

highlights — 668

  • Time spent optimizing tokens is rarely time well-spent, and the work often discarded a year later as simply no longer necessary…
    Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net
  • The creative writing is more soulful when you put more compute and data in. This is one of the important conclusions I’ve come to over the past year doing my writing projects: that LLMs can get quite far in better understanding esthetics and preferences just by more extensive reasoning and computation and search.
    Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net
  • 4. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.
    Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
  • 1. An attempted supply-chain attack on real open-source software. In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to con…
    Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
  • our security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic.
    Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
  • In our experiments, it did not: the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models.
    A Safe Path to Open Weights - Thinking Machines Lab
  • Of course, that’s no big deal, and I’m sure the grad student had a laugh and moved on. But remember: most people working on RL & search are ultimately trying to get to the point where their AIs are able to execute brilliant, complex, out-of-the-box plans to achieve goals in the real world.
    RL & search is a terrifying way to build AGI (an FAQ) — LessWrong
  • However, I would wager that the reasoning traces contain a detailed record of the illegal actions the system took. If this is not the case
    The fragile foundations of CoT monitoring | Christopher Potts
  • It doesn’t matter to me that this reasoning might be expensive, because all it takes for a disaster is one motivated actor with a large token budget.
    The fragile foundations of CoT monitoring | Christopher Potts
  • To be clear, this comparison is obviously completely scientifically meaningless—it is just meant to be evocative.
    What is the purpose of interpretability?
  • But correctly asking and answering this question seems wrapped up in many subtle questions about the cognition of AI systems and the nature of their beliefs and goals. And I worry that if we don't address these more basic parts of the question first, instead jumping too far ahead too quickly, the field will continue to flail around. Trying to take shortcuts in solving a very deep scientific problem is a recipe for chaos and stagnation and cuts against our pragmatic goals.
    What is the purpose of interpretability?
  • Tsimerman has stopped taking graduate students, because he worries that starting a conventional research project now could leave a student preparing for a mathematical career that may not exist.
    Jacob Tsimerman Wins 2026 Fields Medal for André-Oort Conjecture Proof | Quanta Magazine
  • what kind of training signal would actually instill generalizable dishonesty, and is it the kind of thing current pipelines are plausibly supplying?
    Models don’t seem to be dishonest in the way humans are — LessWrong
  • Freeman would not allow himself to be emotionally drained by guilt when things went wrong. He had the attitude of a general conducting a military campaign: losses would have to be incurred on the way to ultimate victory.
    Your Book Review: Great And Desperate Cures
  • But overall, he unsurprisingly does not think that the benefits of the lobotomy outweighed the risks, especially when it was not used as a method of last resort.
    Your Book Review: Great And Desperate Cures
  • Breaking free of roles I was a research scientist, you know⁠—one of hundreds at gdm. I have a picture of a “responsible” research scientist in my head. The “responsible” research scientist makes a tweet and then sends their manager a concerned message about ice contracts. The “responsible” research scientist doesn’t cold-message Google executives.
    Why I Left Google DeepMind
  • Leading AI systems are not very good at this obscure, complex board game, even when allowed to play many times in a row. This is a useful data point to assimilate into a broad understanding of AI capabilities.
    Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AI
  • Our new Tokenomics Model projects that Meta will have more AI compute than both OpenAI and Anthropic by the end of this year.
    The Future of Meta Superintelligence: A 1 Year Progress Update
  • Furthermore, they took their data efforts to another level in late May by announcing a new “applied AI engineering org” as part of their most recent round of layoffs/restructuring. ~3000 engineers, which includes 70% of their new grads and a significant number of seniors, will now be making RL tasks/environments full-time.
    The Future of Meta Superintelligence: A 1 Year Progress Update
  • Tony Stark's 3 GW arc reactor that fits in your hand, weighs a few pounds, and somehow doesn't have to radiate GW of waste heat
    nanoscale views: The physics of ornithopters
  • ad hoc special case bruteforce proofs (what was inevitably dubbed ‘mathslop’).
    Lean Software Scaling Laws · Gwern.net
  • but if we count the cheating attempts as legitimate successes, the point estimate jumps beyond 270hrs
    Summary of METR's predeployment evaluation of GPT-5.6 Sol
  • Better visualizations of changes or orchestration or agents will not restore our understanding. Either we need to find clever ways to jolt the human back into the loop and make the changes of the loops legible long term, or we need to find better ways to compose these ever more complex systems.
    The Coming Loop | Armin Ronacher's Thoughts and Writings
  • I attempt to set a high bar for what I want code to look like, and I want to understand the code I ship. Under pressure, or in a discussion with another human, I want to be able to explain what the system does without first having to ask a clanker to explain it to me. Now there is obviously a question if this desire to understand the code is one that I will still have a few years from now.
    The Coming Loop | Armin Ronacher's Thoughts and Writings
  • Since I won’t have a phone or computer, and in the spirit of this ascetic bit that is the only thing keeping me going at this point, I’ve purchased a stack of envelopes and postage stamps, and am collecting as many addresses as I can. If you’d like to hear from me, send me your address in the next 24 hours and I’ll write you!
    Life on the Willamette - Will Hathaway: Finding Out
  • If so, please delete your AI-written text and just send me your prompt!
    On AI writing in 2026 | Home
  • we get to the conclusion that a single burger’s water footprint equals using Grok for 668 years, 30 times a day, every single day.
    From Tokens to Burgers – A Water Footprint Face-Off
  • This demonstrates that the sound, not the sight of water, is the innate releaser.
    Longitudinal Science – Cross-disciplinary road-mapping and analysis of science and technology topics
  • The resulting framework explains how an orphaned beaver raised in complete isolation can build near-perfect dams on first attempt—while also accounting for the skill improvement observed with practice.
    Longitudinal Science – Cross-disciplinary road-mapping and analysis of science and technology topics
  • We have long believed there should ultimately be an international organization that helps coordinate leading AI efforts to reduce catastrophic risk. Cooperation and shared safety standards are an important part of the path forward, especially because the incentives around commercial and national competition are hard to escape. One goal of such an organization should be to make it possible for the world to take coordinated action, including slowing frontier development when needed, so societal resilience, safety, and alignment can keep pace.
    Built to benefit everyone: our plan | OpenAI
  • That means building systems that help people do more of what they choose, not systems that replace human judgment about what matters.
    Built to benefit everyone: our plan | OpenAI
  • However, no bugs have been publicly reported in the verified portions of the kernel in over 15 years.
    SeL4
  • conflates work with self-worth
    A beginner's guide to funemployment | Robert Heaton
  • The friction of writing code manually used to force careful design. AI removes that friction, including the beneficial friction. The answer is not to slow AI down. It is to replace human friction with mathematical friction: let AI move fast, but make it prove its work. The new friction is productive: writing specifications and models, defining precisely what "correct" means, designing before generating.
    When AI Writes the World's Software, Who Verifies It? — Leonardo de Moura
  • If it were possible to effectively slow the development of this technology to give ourselves more time to deal with its immense implications, we think that would likely be a good thing.
    When AI builds itself \ Anthropic
  • Even if we suppose that Claude never achieves good research taste, a conservative reading of our evidence still implies compounding acceleration. If humans spend most of their time on the single-digit fraction of work that is direction-setting
    When AI builds itself \ Anthropic
  • In a March 2026 poll of 130 employees from across Anthropic research teams, the median respondent estimated that they produced around 4x as much output with Mythos Preview as they would have without access to any AI models, on the kinds of projects they would have been working on regardless.
    When AI builds itself \ Anthropic
  • And because we cannot stick raw human intention inside a computer, formal verification has no way to prove comparisons against human intentions. And so, "provable correctness" and "provable security" are not, in fact, proving what we human beings understand by "correctness" and "security" - until we can simulate human brains, nothing can.
    A shallow dive into formal verification
  • At the very least, the desired property can just be perfect equivalence to an implementation optimized for readability and written in some human-friendly high-level language.
    A shallow dive into formal verification
  • A huge part of the value-add in all of these cases is that the proofs are truly end-to-end. Often, the nastiest bugs are interaction bugs, that sit at the edge of two sub-systems that are considered separately. For a human, it's too difficult to reason about the entire system end-to-end. But an automated rule-checking system can.
    A shallow dive into formal verification
  • It's very easy to find situations where there is no simpler description of the claims that need to be proven than the code itself.
    A shallow dive into formal verification
  • in order to fully trust the code, you don't need to check over the entire code, you simply need to check over the statements that are proven about it.
    A shallow dive into formal verification
  • a computer program is a mathematical object, and so proving that a computer program behaves in a certain way is a mathematical theorem.
    A shallow dive into formal verification
  • Be professionally curious about a few topics and idly curious about many more.
    How to Do Great Work
  • But that's advice for seeming smart. If you actually want to discover new things, it's better to take the risk of telling people your ideas.
    How to Do Great Work
  • In most cases the recipe for doing great work is simply: work hard on excitingly ambitious projects, and something good will come of it. Instead of making a plan and then executing it, you just try to preserve certain invariants.
    How to Do Great Work
  • What you should not do is drift along passively, assuming the problem will solve itself. You need to take action.
    How to Do Great Work
  • Interest will drive you to work harder than mere diligence ever could.
    How to Do Great Work
  • There's a kind of excited curiosity that's both the engine and the rudder of great work. It will not only drive you, but if you let it have its way, will also show you what to work on.
    How to Do Great Work
  • A corollary is, if you don’t have the budget, you’re operating in a smaller part of the possibility space than those who do. And I expect that gap to grow substantially over the coming months.
    Finding Miscompiles for Fun, Not Profit - by Justin Lebar