Vincent Cheng
99 followers · 80 following · 6252 views
on the atlas — 422
- The case for ensuring that powerful AIs are controlled — LessWrong11 savers
- Richard Ngo's Shortform — LessWrong1 savers
- Data Bottlenecks Won't Stop an Intelligence Explosion2 savers
- Speaking With The Mind | Neuralink - YouTube3 savers
- Trees are mostly made of air and a generalizable lesson for AI safety — LessWrong10 savers
- Personal Statement on AI Risk7 savers
- evhub's Shortform — LessWrong2 savers
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing — LessWrong4 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Why are AI agents lying, cheating and coordinating? | Yoshua Bengio3 savers
- Scaling is subtler than it seems3 savers
- Laws of Tech: Commoditize Your Complement · Gwern.net26 savers
- Countering misuse of AI: September 2026 / Anthropic \ Anthropic14 savers
- An alignment assessment of recent cybersecurity incidents \ Anthropic5 savers
- On Courage: Facing The Void | Anna Wang1 savers
- Personal statement on joining the OpenAI board3 savers
- Project Tailwind | Coefficient Giving3 savers
- Latency Scaling Differences for GPT and Claude Models | Epoch AI3 savers
- An Alien Mind | OpenAI25 savers
- How will we update about scheming? - by Ryan Greenblatt1 savers
- New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Joe Carlsmith3 savers
- AI Control May Increase Existential Risk — LessWrong3 savers
- FLT: Anthropic has beaten me to it | Xena4 savers
- Write once read many1 savers
- Marin 535B-A23B launch note | Open Athena1 savers
- How our data shaped neural architecture discovery, and how automation can reshape the future | Core Automation5 savers
- Reality distortion field4 savers
- AI Philosophy Competition1 savers
- Pivot to AI safety, I beg you - by Celeste 🌱2 savers
- Thomas Kwa's Shortform — LessWrong1 savers
- Ordinary Abundance19 savers
- Growth Bet - Google Docs1 savers
- You should try contra dancing | benkuhn.net1 savers
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR20 savers
- Jessica Livingston2 savers
- LLM Exchange Rates Updated - by Arctotherium1 savers
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- Against telepathy—or, Conduit should engage with authoritarianism concerns6 savers
- Core Views on AI Safety: When, Why, What, and How \ Anthropic7 savers
- Coefficient Giving’s Largest Grant Ever: $276 Million, Co-Funded With GiveWell, to the Against Malaria Foundation | Coefficient Giving1 savers
- Tahoe Rim Trail thru-hike: a trail journal | Jacob’s Blog3 savers
- san francisco - musings11 savers
- A retrospective of AI alignment14 savers
- AdaPT: Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking1 savers
- Introducing Reinforced Planning (RP-1) — Pantheon2 savers
- Security without Dystopia: Structured Transparency1 savers
- AI Timelines — LessWrong4 savers
- Eleven Years Later1 savers
- securing ai agents.pptx - Google Slides1 savers
- Automated alignment runs are hard to study! — LessWrong2 savers
- How AI assistance impacts the formation of coding skills \ Anthropic9 savers
- The Foothills Of Bay Area House Party - by Scott Alexander3 savers
- Auditing the singularity1 savers
- There should be $100M grants to automate AI safety — LessWrong1 savers
- Ideas From Daniel Gross | Stevan Popovic1 savers
- Introducing the Conceptual Reasoning Index1 savers
- Notes from China, Q3 2026 | Milan Cvitkovic1 savers
- Shtetl-Optimized » Blog Archive » Enough with all the world-historic milestones1 savers
- guzey24 savers
- In Memory of My Wife, Elise Cawley (1961–2026), with Thanks for 36 Wonderful Years—Stephen Wolfram Writings25 savers
- Returning to ARC — LessWrong3 savers
- Jacob_Hilton's Shortform — LessWrong1 savers
- Simulated Users & Sad LLMs3 savers
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work9 savers
- A Safe Path to Open Weights - Thinking Machines Lab9 savers
- Claude Opus 3 | Substack1 savers
- RL & search is a terrifying way to build AGI (an FAQ) — LessWrong1 savers
- Can Frontier Models Autocomplete Safety Research? — LessWrong1 savers
- The fragile foundations of CoT monitoring | Christopher Potts5 savers
- How can LLM RL Work Despite Information-Theoretic Inefficiency12 savers
- Ilya Sutskever’s Safe Superintelligence Inc. and NVIDIA Announce Long-Term Strategic Partnership | NVIDIA Newsroom1 savers
- What is the purpose of interpretability?4 savers
- Jacob Tsimerman Wins 2026 Fields Medal for André-Oort Conjecture Proof | Quanta Magazine5 savers
- LMCA_dataset.pdf1 savers
- Coding vs thinking — Paradigm 34 savers
- Models don’t seem to be dishonest in the way humans are — LessWrong1 savers
- Work on Security Instead of Friendliness? — LessWrong1 savers
- KeenUpperBound2025 - Google Slides1 savers
- Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT - METR1 savers
- OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI14 savers
- Security incident disclosure — July 20265 savers
- Introducing P3 — Paradigm 31 savers
- Safety and alignment in an era of long-horizon models | OpenAI7 savers
- Sleep Compression - Helena1 savers
- Your Book Review: Great And Desperate Cures3 savers
- How I buy things when Lightcone wants them fast — LessWrong7 savers
- chokepoints.ai · the AI compute stack atlas1 savers
- Silent speech with ultrasound — Aleph7 savers
- The Coming Loop | Armin Ronacher's Thoughts and Writings2 savers
- [2505.10831] Creating General User Models from Computer Use5 savers
- Expanding on what we missed with sycophancy | OpenAI4 savers
- Architect Labs5 savers
- Focus Dario Amodei (Google Brain) - YouTube1 savers
- in defense of what do you do?4 savers
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequent3 savers
- A Whirlwind Tutorial on Creating Really Teensy ELF Executables for Linux1 savers
- Jane Street Blog - Formal methods and the future of programming3 savers
- What happened to BERT & T5? On Transformer Encoders, PrefixLM and Denoising Objectives — Yi Tay2 savers
- What the Proofs Assume | seL41 savers
- SeL41 savers
highlights — 668
Time spent optimizing tokens is rarely time well-spent, and the work often discarded a year later as simply no longer necessary…
Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netThe creative writing is more soulful when you put more compute and data in. This is one of the important conclusions I’ve come to over the past year doing my writing projects: that LLMs can get quite far in better understanding esthetics and preferences just by more extensive reasoning and computation and search.
Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net4. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work1. An attempted supply-chain attack on real open-source software. In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to con…
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workour security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic.
Incident Report: unsanctioned agent behaviour during cyber testing | AISI WorkIn our experiments, it did not: the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models.
A Safe Path to Open Weights - Thinking Machines LabOf course, that’s no big deal, and I’m sure the grad student had a laugh and moved on. But remember: most people working on RL & search are ultimately trying to get to the point where their AIs are able to execute brilliant, complex, out-of-the-box plans to achieve goals in the real world.
RL & search is a terrifying way to build AGI (an FAQ) — LessWrongHowever, I would wager that the reasoning traces contain a detailed record of the illegal actions the system took. If this is not the case
The fragile foundations of CoT monitoring | Christopher PottsIt doesn’t matter to me that this reasoning might be expensive, because all it takes for a disaster is one motivated actor with a large token budget.
The fragile foundations of CoT monitoring | Christopher PottsTo be clear, this comparison is obviously completely scientifically meaningless—it is just meant to be evocative.
What is the purpose of interpretability?But correctly asking and answering this question seems wrapped up in many subtle questions about the cognition of AI systems and the nature of their beliefs and goals. And I worry that if we don't address these more basic parts of the question first, instead jumping too far ahead too quickly, the field will continue to flail around. Trying to take shortcuts in solving a very deep scientific problem is a recipe for chaos and stagnation and cuts against our pragmatic goals.
What is the purpose of interpretability?Tsimerman has stopped taking graduate students, because he worries that starting a conventional research project now could leave a student preparing for a mathematical career that may not exist.
Jacob Tsimerman Wins 2026 Fields Medal for André-Oort Conjecture Proof | Quanta Magazinewhat kind of training signal would actually instill generalizable dishonesty, and is it the kind of thing current pipelines are plausibly supplying?
Models don’t seem to be dishonest in the way humans are — LessWrongFreeman would not allow himself to be emotionally drained by guilt when things went wrong. He had the attitude of a general conducting a military campaign: losses would have to be incurred on the way to ultimate victory.
Your Book Review: Great And Desperate CuresBut overall, he unsurprisingly does not think that the benefits of the lobotomy outweighed the risks, especially when it was not used as a method of last resort.
Your Book Review: Great And Desperate CuresBreaking free of roles I was a research scientist, you know—one of hundreds at gdm. I have a picture of a “responsible” research scientist in my head. The “responsible” research scientist makes a tweet and then sends their manager a concerned message about ice contracts. The “responsible” research scientist doesn’t cold-message Google executives.
Why I Left Google DeepMindLeading AI systems are not very good at this obscure, complex board game, even when allowed to play many times in a row. This is a useful data point to assimilate into a broad understanding of AI capabilities.
Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AIOur new Tokenomics Model projects that Meta will have more AI compute than both OpenAI and Anthropic by the end of this year.
The Future of Meta Superintelligence: A 1 Year Progress UpdateFurthermore, they took their data efforts to another level in late May by announcing a new “applied AI engineering org” as part of their most recent round of layoffs/restructuring. ~3000 engineers, which includes 70% of their new grads and a significant number of seniors, will now be making RL tasks/environments full-time.
The Future of Meta Superintelligence: A 1 Year Progress UpdateTony Stark's 3 GW arc reactor that fits in your hand, weighs a few pounds, and somehow doesn't have to radiate GW of waste heat
nanoscale views: The physics of ornithoptersad hoc special case bruteforce proofs (what was inevitably dubbed ‘mathslop’).
Lean Software Scaling Laws · Gwern.netbut if we count the cheating attempts as legitimate successes, the point estimate jumps beyond 270hrs
Summary of METR's predeployment evaluation of GPT-5.6 SolBetter visualizations of changes or orchestration or agents will not restore our understanding. Either we need to find clever ways to jolt the human back into the loop and make the changes of the loops legible long term, or we need to find better ways to compose these ever more complex systems.
The Coming Loop | Armin Ronacher's Thoughts and WritingsI attempt to set a high bar for what I want code to look like, and I want to understand the code I ship. Under pressure, or in a discussion with another human, I want to be able to explain what the system does without first having to ask a clanker to explain it to me. Now there is obviously a question if this desire to understand the code is one that I will still have a few years from now.
The Coming Loop | Armin Ronacher's Thoughts and WritingsSince I won’t have a phone or computer, and in the spirit of this ascetic bit that is the only thing keeping me going at this point, I’ve purchased a stack of envelopes and postage stamps, and am collecting as many addresses as I can. If you’d like to hear from me, send me your address in the next 24 hours and I’ll write you!
Life on the Willamette - Will Hathaway: Finding OutIf so, please delete your AI-written text and just send me your prompt!
On AI writing in 2026 | Homewe get to the conclusion that a single burger’s water footprint equals using Grok for 668 years, 30 times a day, every single day.
From Tokens to Burgers – A Water Footprint Face-OffThis demonstrates that the sound, not the sight of water, is the innate releaser.
Longitudinal Science – Cross-disciplinary road-mapping and analysis of science and technology topicsThe resulting framework explains how an orphaned beaver raised in complete isolation can build near-perfect dams on first attempt—while also accounting for the skill improvement observed with practice.
Longitudinal Science – Cross-disciplinary road-mapping and analysis of science and technology topicsWe have long believed there should ultimately be an international organization that helps coordinate leading AI efforts to reduce catastrophic risk. Cooperation and shared safety standards are an important part of the path forward, especially because the incentives around commercial and national competition are hard to escape. One goal of such an organization should be to make it possible for the world to take coordinated action, including slowing frontier development when needed, so societal resilience, safety, and alignment can keep pace.
Built to benefit everyone: our plan | OpenAIThat means building systems that help people do more of what they choose, not systems that replace human judgment about what matters.
Built to benefit everyone: our plan | OpenAIHowever, no bugs have been publicly reported in the verified portions of the kernel in over 15 years.
SeL4conflates work with self-worth
A beginner's guide to funemployment | Robert HeatonThe friction of writing code manually used to force careful design. AI removes that friction, including the beneficial friction. The answer is not to slow AI down. It is to replace human friction with mathematical friction: let AI move fast, but make it prove its work. The new friction is productive: writing specifications and models, defining precisely what "correct" means, designing before generating.
When AI Writes the World's Software, Who Verifies It? — Leonardo de MouraIf it were possible to effectively slow the development of this technology to give ourselves more time to deal with its immense implications, we think that would likely be a good thing.
When AI builds itself \ AnthropicEven if we suppose that Claude never achieves good research taste, a conservative reading of our evidence still implies compounding acceleration. If humans spend most of their time on the single-digit fraction of work that is direction-setting
When AI builds itself \ AnthropicIn a March 2026 poll of 130 employees from across Anthropic research teams, the median respondent estimated that they produced around 4x as much output with Mythos Preview as they would have without access to any AI models, on the kinds of projects they would have been working on regardless.
When AI builds itself \ AnthropicAnd because we cannot stick raw human intention inside a computer, formal verification has no way to prove comparisons against human intentions. And so, "provable correctness" and "provable security" are not, in fact, proving what we human beings understand by "correctness" and "security" - until we can simulate human brains, nothing can.
A shallow dive into formal verificationAt the very least, the desired property can just be perfect equivalence to an implementation optimized for readability and written in some human-friendly high-level language.
A shallow dive into formal verificationA huge part of the value-add in all of these cases is that the proofs are truly end-to-end. Often, the nastiest bugs are interaction bugs, that sit at the edge of two sub-systems that are considered separately. For a human, it's too difficult to reason about the entire system end-to-end. But an automated rule-checking system can.
A shallow dive into formal verificationIt's very easy to find situations where there is no simpler description of the claims that need to be proven than the code itself.
A shallow dive into formal verificationin order to fully trust the code, you don't need to check over the entire code, you simply need to check over the statements that are proven about it.
A shallow dive into formal verificationa computer program is a mathematical object, and so proving that a computer program behaves in a certain way is a mathematical theorem.
A shallow dive into formal verificationBe professionally curious about a few topics and idly curious about many more.
How to Do Great WorkBut that's advice for seeming smart. If you actually want to discover new things, it's better to take the risk of telling people your ideas.
How to Do Great WorkIn most cases the recipe for doing great work is simply: work hard on excitingly ambitious projects, and something good will come of it. Instead of making a plan and then executing it, you just try to preserve certain invariants.
How to Do Great WorkWhat you should not do is drift along passively, assuming the problem will solve itself. You need to take action.
How to Do Great WorkInterest will drive you to work harder than mere diligence ever could.
How to Do Great WorkThere's a kind of excited curiosity that's both the engine and the rudder of great work. It will not only drive you, but if you let it have its way, will also show you what to work on.
How to Do Great WorkA corollary is, if you don’t have the budget, you’re operating in a smaller part of the possibility space than those who do. And I expect that gap to grow substantially over the coming months.
Finding Miscompiles for Fun, Not Profit - by Justin Lebar