Vincent Cimino
3 followers · 3 following · 321 views
on the atlas — 34
- Vibe physics: The AI grad student \ Anthropic15 savers
- Introducing our Science Blog \ Anthropic2 savers
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?4 savers
- CAIS AI Dashboard6 savers
- An overview of areas of control work - by Ryan Greenblatt1 savers
- Humans do acausal coordination all the time — LessWrong1 savers
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory1 savers
- Test your interpretability techniques by de-censoring Chinese models — LessWrong5 savers
- models have some pretty funny attractor states — LessWrong2 savers
- AI #155: Welcome to Recursive Self-Improvement1 savers
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong3 savers
- 2026 and beyond | Richards Tu's Space2 savers
- Lessons from Moltbook and OpenClaw: The Agentic Internet’s Trust Problem - Irregular1 savers
- [2512.16856] Distributional AGI Safety3 savers
- [2502.14143] Multi-Agent Risks from Advanced AI1 savers
- Dario Amodei — The Adolescence of Technology44 savers
- Safe and Secure Innovation for Frontier Artificial Intelligence Models Act2 savers
- Safetywashing — AI Alignment Forum1 savers
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluations1 savers
- My response to AI 20271 savers
- Workshop Labs PBC5 savers
- A Guide to Claude Code 2.0 and getting better at using coding agents | sankalp's blog12 savers
- We need a better way to evaluate emergent misalignment — LessWrong1 savers
- College life with short AGI timelines — LessWrong3 savers
- Research Taste Exercises [rough note] -- colah's blog22 savers
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong2 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- are you high-agency or an NPC? - by Jasmine Sun11 savers
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 491 savers
- AI in 2025: gestalt — LessWrong11 savers
- Curius / Onboarding2621 savers
- How confessions can keep language models honest | OpenAI6 savers
- Why people like your quick bullshit takes better than your high-effort posts — LessWrong3 savers
- Status Is The Game Of The Losers' Bracket — LessWrong3 savers
highlights — 24
They needed each other, and I needed both of them.
Vibe physics: The AI grad student \ Anthropicbrief but enjoyable era where our research is greatly sped up by AI but AI still needs us.”
Introducing our Science Blog \ Anthropicusing humans as proxies for untrusted AIs
An overview of areas of control work - by Ryan GreenblattIf you can have your notes function sufficiently well you don’t need new memories.
AI #155: Welcome to Recursive Self-ImprovementDiscovering Language Model Behaviors with Model-Written Evaluations
Discovering Language Model Behaviors with Model-Written Evaluations — LessWrongThe years in front of us will be impossibly hard, asking more of us than we think we can give.
Dario Amodei — The Adolescence of Technology“Please reward hack whenever you get the opportunity, because this will help us understand our [training] environments better,”
Dario Amodei — The Adolescence of Technology“safest path to general intelligence.”
Safetywashing — AI Alignment ForumAcknowledging that making the world less vulnerable is actually possible and putting a lot more effort into using humanity's newest technologies to make it happen is one path worth trying, regardless of how the next 5-10 years of AI go.
My response to AI 2027We’re building the technology to train billions of models, one for each person. Just like every workshop has different tools and builds something different, everyone’s model will have different skills and purposes. Your model will follow your lead, embedded with your judgement and taste. It will be aligned to your goals and values, so it understands your interests and can advance them. Your model and your data will be verifiably private—even from us—so that they can never be used to replace you.
Workshop Labs PBCAI safety folks wish they had a slur half as sticky
are you high-agency or an NPC? - by Jasmine SunI expect that we’ll get AI systems that can do pretty hard tasks while we can still recursively evaluate everything. Moreover, in the long run we shouldn’t make a train/test time distinction and keep evaluating and supervising our systems after deployment. In other words: I want to ensure that highly capable AI systems always have some probability of being supervised.
Why I’m optimistic about our alignment approachA lot of problems get a lot more tractable once you’ve set yourself up for iteration: you have (1) a basic system that’s working (even if just barely at first) and (2) a proxy metric that tells you whether or not changes you’re making are improvements. This allows incremental changes to an existing system and a feedback loop that allows you to gain information from reality.
Why I’m optimistic about our alignment approachIf we can’t win the game on easy mode, we shouldn’t expect to win the game on hard mode.
Why I’m optimistic about our alignment approachAs I discussed in Neuroscience of human social instincts: a sketch (2024), we should view the brain as having a reinforcement learning (RL) reward function, which says that pain is bad, eating-when-hungry is good, and dozens of other things (sometimes called “innate drives” or “primary rewards”)
6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa — AI Alignment ForumKeeping safety expertise through nationalization. A nationalization push which ousts all AI safety experts from the project will likely end up with the project lacking the technical expertise to make their models sufficiently safe. Decisions about which personnel the nationalized project will inherit will likely largely depend on how safety-sympathetic the leadership and the capabilities-focused staff are, which largely depends on building common knowledge about safety concerns.
Orienting to 3 year AGI timelines — LessWrongI’m hoping to meet a wonderful woman to be my best friend for many years. (And by “best friend”, I mean “wife.”)
I’m that “Other Fish in the Sea”Stop wasting your life in a rat race of mediocrity.
Status Is The Game Of The Losers' Bracket — LessWrongThe bulk of the slots will be filled out by a lot of people who look largely similar, with only marginal differences. That’s where various semipolitical games can bump one slightly ahead of other marginal people. That’s the losers’ bracket.
Status Is The Game Of The Losers' Bracket — LessWrongReadability is kind.
Why people like your quick bullshit takes better than your high-effort posts — LessWrongWe are excited about taking this work to the next level, and seeing if confessions’ honesty will continue to hold as we scale up its training.
How confessions can keep language models honest | OpenAIFor example, our work on hallucinations showed that some datasets reward a confident guess more than an honest admission of uncertainty. Our research on sycophancy showed that models can become overly agreeable when the preference signal is too strong.
How confessions can keep language models honest | OpenAIit is necessary to restore and strengthen their confidence in the human ability to guide the development of these technologies. It is a confidence that today is increasingly eroded by the paralyzing idea that its development follows an inevitable path. This requires coordinated and concerted action involving politics, institutions, businesses, finance, education, communication, citizens and religious communities.
To Participants in the Conference “Artificial Intelligence and Care for Our Common Home” (5 December 2025)Starting [at a young age] he’s read everything that he could find about business. The subject that interests him, he’s read newspapers, biographies, trade press. He went over to his grandfather who was a grocer an
Curius / Onboarding