Parv Mahajan
25 followers · 18 following · 393 views
on the atlas — 41
- A playbook for field strategy - by Dewi Erwan5 savers
- Why people like your quick bullshit takes better than your high-effort posts — LessWrong3 savers
- The Whispering Earring (Scott Alexander) - Croissanthology10 savers
- Plans A, B, C, and D for misalignment risk4 savers
- The seam through the center of things - by Cate Hall5 savers
- In My Misanthropy Era — LessWrong2 savers
- AI Tools for Existential Security | Forethought5 savers
- Claude can make mistakes. Please double-check responses. | Julian Michael7 savers
- Overview | Shallow Review 20257 savers
- Writing one sentence per line | Derek Sivers5 savers
- You will be OK — LessWrong13 savers
- [AI Futures Model] Supplementary materials - Google Docs5 savers
- On Doubling9 savers
- Examples of barbell strategies11 savers
- Making Normal Conversations Better - by Sasha Chapin12 savers
- Milan Cvitkovic52 savers
- 'How to be a Human' Starter Pack - by Lydia Nottingham6 savers
- Why Prayer Is Not Answered2 savers
- thoughts on safety in sf - by vincent huang4 savers
- Opinionated Takes on Meetups Organizing — LessWrong1 savers
- The Most Valuable Commodity2 savers
- EA orgs' legal structure inhibits risk taking and information sharing on the margin — LessWrong1 savers
- On Fleshling Safety: A Debate by Klurl and Trapaucius. — LessWrong1 savers
- maybe commitment is the most beautiful thing - by Lilian1 savers
- The race to AGI-pill the pope | The Verge3 savers
- Neel Nanda on leading a Google DeepMind team at 26 – and advice if you want to work at an AI company (part 2) - 80,000 Hours2 savers
- AI friends are already here6 savers
- how to party like an AI researcher - by Jasmine Sun2 savers
- are you high-agency or an NPC? - by Jasmine Sun11 savers
- There's No Fire Alarm for Artificial General Intelligence — LessWrong1 savers
- Governing Automated Strategic Intelligence1 savers
- 2025 letter | Zhengdong33 savers
- The Intelligence Curse28 savers
- Defining the Intelligence Curse - The Intelligence Curse13 savers
- Capital, AGI, and Human Ambition - The Intelligence Curse10 savers
- Pyramid Replacement - The Intelligence Curse9 savers
- Why So Few Matt Levines? · Gwern.net8 savers
- Breaking the Intelligence Curse - The Intelligence Curse4 savers
- Shaping the Social Contract - The Intelligence Curse4 savers
- Self-Help Tactics That Are Working For Me — LessWrong2 savers
- Strategy as Problem-Solving > US Army War College - Publications > Display2 savers
highlights — 56
But at today’s frontier AI companies, recovering physics PhDs, burned-out quants, and logorrheic LessWrong vets top the status ladder.
chinese peptide physiognomy - by Jasmine SunImmediately after you’ve been introduced to someone, send them a calendar invite for the next day at a reasonable time for their timezone. Use this format for the calendar event title: “[TBC] YourName TheirName”.
A playbook for field strategy - by Dewi ErwanAt the start of every conversation, share your backstory, ask them for theirs, and find common ground. Demonstrate you’re competent, be vulnerable, and mirror their energy.
A playbook for field strategy - by Dewi ErwanYou’ve not been appointed to do this, so you need to earn legitimacy and trust by doing the hard work well.
A playbook for field strategy - by Dewi ErwanStrategy isn’t a goal, it’s a form of problem-solving, and you can’t solve a problem you don’t understand and haven’t defined.
A playbook for field strategy - by Dewi ErwanAvoid editing and writing a blog post at the same time, where you erase two words for every three you scrawl. Instead, write five streams-of-consciousness in one day and edit them into coherent essays throughout the week. That way, the artist inside you isn’t immediately neutered by the critic and actually has the opportunity to create. At the same time, the critic will get a chance to make sure that the artist doesn’t just take a big shit on your blog.
Examples of barbell strategiesBut the key thing is that they can’t feel like monologues.
Making Normal Conversations Better - by Sasha ChapinEngage a human productivity monitor
Milan CvitkovicAsk your acquaintances, “Hey, I want to leave my house more, are there any cool events you’re going to soon?” (HT Sasha Chapin)
Milan CvitkovicCan we raise the ‘functioning member of society’ waterline?
'How to be a Human' Starter Pack - by Lydia Nottinghamexisting in this state of nonstop paranoia feels really awful, and i realized pretty quickly that it isn't how i want to live my life. i don’t want to be the kind of person who is deeply suspicious of everyone they encounter, and i wonder if viewing all strangers as threats affects your psyche in ways more subtle and sinister than you can be conscious of.
thoughts on safety in sf - by vincent huangWe don’t talk enough about the courage of living according to your convictions; of saying no to the endless horizon of possibilities. We live in a society that reveres optionality, but maybe what we need is the opposite: a willingness to choose, to stand for something, to be still somewhere and say, unabashedly, this is what I believe in, this is what I love.
maybe commitment is the most beautiful thing - by LilianIt sort of has the feeling of going around Rome like an old 1950s gumshoe, pigeonholing priests in gelaterias, grabbing them by the cassock, and saying, ‘all right, give it to me straight, Padre, how does it really work around here?’” he says. “That’s just a massive difference from DC.”
The race to AGI-pill the pope | The Vergeuccess also requires bridging the worlds of Catholicism and AI safety, two communities not known for their overlap. “The number could probably fit in an elevator,” Levin estimates.
The race to AGI-pill the pope | The Verge’m sure you must get many of these emails. So to help you prioritise, here’s some key info about me: blah, blah, blah
Neel Nanda on leading a Google DeepMind team at 26 – and advice if you want to work at an AI company (part 2) - 80,000 HoursWell, the genie is out of the bottle on AI friends. Recently, a colleague gave a talk to a LA high school and asked how many students considered themselves emotionally attached to an AI. One-third of the room raised their hand. I initially found this anecdote somewhat unbelievable, but the reality is even more stark: per a 2025 survey from Common Sense Media, 52% of American teenagers are “regular users” of AI companions
AI friends are already hereThe joke is that reading Google’s published papers is a better way to discover what they’re not working on—anything going into Gemini training won’t be released.
how to party like an AI researcher - by Jasmine SunI don’t know if it’s disinterest or superstition or a sense of invincibility: We’re all techno-optimists, we aren’t supposed to feel fear. Acknowledge the precarity and you might make it real.
are you high-agency or an NPC? - by Jasmine SunAt some point we might see a sudden flood of arXiv papers in which really interesting and fundamental and scary cognitive challenges seem to be getting done at an increasing pace. Whereupon, as this flood accelerates, even some who imagine themselves sober and skeptical will be unnerved to the point that they venture that perhaps AGI is only 15 years away now, maybe, possibly. The signs might become so blatant, very soon before the end, that people start thinking it is socially acceptable to say that maybe AGI is 10 years off. Though the signs would have to be pretty darned blatant, if they’re…
There's No Fire Alarm for Artificial General Intelligence — LessWrongWe are uncertain about how much cognitive effort it takes to perform tasks, and how easy it is to cheat at them, and the first "impressive" tasks to be accomplished will be those where we were most wrong about how much effort was required.
There's No Fire Alarm for Artificial General Intelligence — LessWrongMistral chat templates are terrible for collecting activations on an existing chat transcript (for running probes). By default, system messages and tool messages appear after the last user message, and only appear if the last message was from the user, so it's easy to accidentally collect activations off-policy, instead of seeing the actual activations which caused the model to generate its actions. It's also nontrivial to even set custom chat templates for Mistral models with vLLM.
Auditing Games for Sandbagging [paper] — LessWrongThis model was created by prompt-distillation from instructions to underperform, followed by reinforcement learning to remove obvious signs of sandbagging. Unlike Cedar and Oak, this model is more likely to explicitly conceptualise itself as deliberately sandbagging - similar to the prompt-distilled Tarun model, but harder to detect with transcript analysis thanks to reinforcement learning
Auditing Games for Sandbagging [paper] — LessWrongReal sandbagging models would likely need to reason in-context about when/where/how to sandbag, potentially making them detectable via CoT monitoring. On the other hand, they might also adaptively resist elicitation attempts.
Auditing Games for Sandbagging [paper] — LessWrongThe practical implication: We think elicitation training is likely a good idea in practice to mitigate sandbagging and other kinds of underperformance. Compared to reliably detecting sandbagging, it may be more tractable to just run elicitation training before evaluations.
Auditing Games for Sandbagging [paper] — LessWrongTheir report stated that this sandbagger was "the only model that should be deployed."
Auditing Games for Sandbagging [paper] — LessWrongbut couldn't compare untrusted models to each other.
Auditing Games for Sandbagging [paper] — LessWrongour red team created five model organisms (some sandbagging, some benign decoys), and then a blinded blue team tried to identify which were which.
Auditing Games for Sandbagging [paper] — LessWrongTo store an idea at night, I create a mental pointer to the idea and drop it near the next empty landmark along the path. Good pointers are characters that interact with the environment in one way or another.
Solving a problem with mindwarewould likely do little to slow China’s indigenization efforts
Should the US Sell Hopper Chips to China? | IFPwe estimate near parity between Blackwells and H200s for many inference workloads; up to a 5x advantage for Blackwells on the type of inference workloads for which they are best suited; and a 1.5x advantage for Blackwell-based clusters in training. This means that with access to Hoppers, Chinese labs could build AI training supercomputers as capable as American ones at 50% extra cost5 — a premium that the Chinese Communist Party (CCP) would likely at least partly subsidize.
Should the US Sell Hopper Chips to China? | IFPHuawei is not planning to produce an AI chip matching the H200 until Q4 2027 at the earliest
Should the US Sell Hopper Chips to China? | IFPHe'd bet on Obama at even money, but he'd also bet on Romney if someone offered him better than 3-to-1. That's what it means to actually believe your model: you're not rooting for an outcome, you're betting on your beliefs.
Thinking in Predictions — LessWrongBut the world we are preparing our students for — medicine, politics, business, art, climate change — is an open, chaotic system. Rules change, information is always imperfect and incomplete, and “winning” often does not involve defeating an opponent. Sometimes, it means collaborating, empathizing, or simply surviving. In that open world, the machine by itself is a tremendously powerful engine — but without a steering wheel. AlphaZero can calculate the optimal path, but it cannot decide where we want to go. Stockfish can sacrifice a queen to win a game, but it cannot decide whether it’s worth …
What and How to Teach When Google Knows Everything and ChatGPT Explains It All Very Well | Office of the PresidentIf I were a betting man, I wouldn’t put my money on humans in any intellectual, cognitive, or even creative task. Claims that machines “can’t drive,” “can’t write,” “can’t compose music,” or “can’t create art” have all been proven false, one after another. Every time we insist on human superiority, a new machine appears to put us in our place. At best, humans are left with moral judgment and emotional intelligence — deciding what matters; what winning and losing mean; what is good or bad, beautiful or ugly. We remain the social, political, moral, and emotional agents who define whether chess (…
What and How to Teach When Google Knows Everything and ChatGPT Explains It All Very Well | Office of the PresidentThe same personalization can (and will) extend to academic advising, curriculum planning, extracurricular recommendations, and career guidance. There may be specific phases of learning — such as introductory programming, algebra, and creative writing — where we should deliberately set AI aside to train raw human skill. But that should be a deliberate pedagogical decision, not a dogmatic one. If we exclude AI from learning, students will ultimately pay the price.
What and How to Teach When Google Knows Everything and ChatGPT Explains It All Very Well | Office of the PresidentOur strategy is not to produce more computer scientists, but to produce professionals in every field who are fluent in AI.
What and How to Teach When Google Knows Everything and ChatGPT Explains It All Very Well | Office of the PresidentIn feudal societies the answer to "why is this person powerful?" would usually involve some long family history, perhaps ending in a distant ancestor who had fought in an important battle: "my great-great-grandfather fought at Bosworth Field!". In the future, the answer to “why is this person powerful?” would trace back to something they or someone they were close with did in the pre-AGI era: "oh, my uncle was technical staff at OpenAI". The children of the future will live their lives in the shadow of their parents, with social mobility extinct. This is far from the worst future we could imag…
Capital, AGI, and Human Ambition - The Intelligence CurseYou might also use formal language in an attempt to Make It Look Professional – unless you’re aiming for a really particular audience that eats up formality, just stop doing that! Readability is kind.
Why people like your quick bullshit takes better than your high-effort posts — LessWrongFor example, some people are trustworthy in the sense that things they say make it pretty easy to guess what they're going to do in the future in a wide variety of situations that might come up; I definitely don't think that this is the case for Anthropic.
Unless its governance changes, Anthropic is untrustworthy — LessWrongAnthropic's mission is not really compatible with the idea of pausing, even if evidence suggests it's a good idea to.
Unless its governance changes, Anthropic is untrustworthy — LessWrongDepending on the contents of the Investors' Rights Agreement, which is not public, it might be impossible for the directors appointed by the LTBT to fire the CEO. (That, notably, would constitute fewer formal rights than even OpenAI's nonprofit board has over OpenAI's CEO.)
Unless its governance changes, Anthropic is untrustworthy — LessWrongAlternatively, if they believe that their strategy should change in light of other labs' misalignment, or their geopolitical views, or anything else, they need to be honest about having changed their mind.
Unless its governance changes, Anthropic is untrustworthy — LessWrongWe endorse this approach, which balances safety concerns with authoritarian risks. Below, we outline specific technologies that, if implemented, could lower risks to an acceptable level in each key issue are.
Breaking the Intelligence Curse - The Intelligence Curse“Internet gloves” where users can use AIs to pull information from platforms in selective, non-addictive ways, without being sucked into the platform.
Breaking the Intelligence Curse - The Intelligence CurseLarge-scale feedback collection that allows policymakers to get more fine-grained and qualitative data about citizens’ preferences than current simple numerical opinion polling does.
Breaking the Intelligence Curse - The Intelligence CurseUpskilling humans in the areas which will bottleneck the AI economy. AI systems are likely to have uneven capability profiles compared to humans, excelling in tasks with easy verification, low time horizons, and a lack of interfacing with the physical world. Naturally, these will create bottlenecks which humans will be able to fill to stay relevant in the economy. There is a race between human upskilling and retraining on one hand, and AI labs smoothing over the jagged performance frontier on the other.25
Breaking the Intelligence Curse - The Intelligence CurseWe expect the AI balance to be more multipolar and arrive more slowly than some of the more aggressive scenarios predict2, making this path less feasible
Shaping the Social Contract - The Intelligence Curse. Third, we think this is much less likely to be effective for advanced AI because we expect states to have far more infrastructural power10 as AI advances, in line with Bullock, Hammond, and Krier’s conception of the AGI-powered “Despotic Leviathan”. As such, powerful actors could spot nearly all major threats to their power. Moreover, new technologies like cheap and very effective autonomous drones could also change the balance of power such that armed uprisings cannot threaten the state. For all these reasons, we expect autocracies with labor-replacing AI to succumb to the intelligence curs…
Defining the Intelligence Curse - The Intelligence CurseWhen oil was discovered in Norway, the country had been a stable democracy since it acquired independence in 1905. The state bureaucracy functioned well, with little corruption. The legal system worked well, and the media was actively evaluating and commenting upon the workings of the system.
Defining the Intelligence Curse - The Intelligence CurseThe economy will increasingly sideline them. In the limit, much of the economy could run in loops that avoid human consumers entirely.
Defining the Intelligence Curse - The Intelligence Curse