Seth Lifland
10 followers · 7 following · 306 views
on the atlas — 61
- Principles for a New Utopianism — DeepMind Institute3 savers
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- Can you control the past? - Joe Carlsmith1 savers
- How to pace the US frontier — LessWrong1 savers
- Radical Optionality — Governing Transformative AI Under Uncertainty6 savers
- Returning to ARC — LessWrong3 savers
- The Huggingface Incident - by Scott Alexander2 savers
- A guide to the AI tribes - by Michel Justen - What is this5 savers
- The Inside Story Of Leverage Research 1.07 savers
- Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong4 savers
- Our position on open-weights models \ Anthropic7 savers
- The Compute Verification Post2 savers
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrong4 savers
- In Favor of Niceness, Community, and Civilization | Slate Star Codex4 savers
- EA is about maximization, and maximization is perilous — EA Forum3 savers
- The Tiny Flicker of Altruism - by Matt Reardon2 savers
- AI 2040: Plan A9 savers
- Making CAISI the AI agency the US needs - by Veronica Irwin2 savers
- Making sense of the UK’s AI Security Institute | Ada Lovelace Institute1 savers
- Who’s doing what on AI security in the US government?3 savers
- Promoting Advanced Artificial Intelligence Innovation and Security – The White House4 savers
- On sincerity - Joe Carlsmith9 savers
- A global workspace in language models \ Anthropic25 savers
- Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AI2 savers
- {Book Summary} The Art of Gathering — EA Forum2 savers
- You Need A Theory of Victory - by Jason Hausenloy4 savers
- We should take AI welfare seriously - by Robert Long2 savers
- Summary of METR's predeployment evaluation of GPT-5.6 Sol7 savers
- Three Types of Intelligence Explosion6 savers
- AI Futures Model6 savers
- Catastrophic AI Scenarios - Future of Life Institute1 savers
- Convergence and Compromise: Will Society Aim for Good Futures?1 savers
- No Easy Eutopia: Why Great Futures Are Hard to Achieve1 savers
- Unresolved debates about the future of AI - by Helen Toner1 savers
- Trends in AI: Training Costs and Diffusion Speed1 savers
- Cognitive Security as an AI Safety Cause Area — LessWrong4 savers
- How AI Will Save Prediction Markets — LessWrong1 savers
- If AI is normal technology, history is not reassuring. — LessWrong2 savers
- Automated Alignment is Harder Than You Think — LessWrong2 savers
- The machines are fine. I'm worried about us.23 savers
- The Artificial Intelligence Revolution: Part 118 savers
- AGI Ruin: A List of Lethalities - LessWrong17 savers
- Dario Amodei — Policy on the AI Exponential16 savers
- 2028: Two scenarios for global AI leadership \ Anthropic12 savers
- Gradual Disempowerment11 savers
- How to Buy Cheap Claude Tokens in China - by Zilan Qian8 savers
- Plans A, B, C, and D for misalignment risk — LessWrong6 savers
- The least understood driver of AI progress | Epoch AI5 savers
- Seeking Stability in the Competition for AI Advantage | RAND4 savers
- In search of a dynamist vision for safe superhuman AI4 savers
- How do we (more) safely defer to AIs? — LessWrong4 savers
- The "Messy Middle" - by Molly Kinder3 savers
- Why do people disagree about when powerful AI will arrive?3 savers
- Short AI Timelines Aren’t Always Higher-Leverage3 savers
- Scaling: The State of Play in AI - by Ethan Mollick3 savers
- "Long" timelines to advanced AI have gotten crazy short3 savers
- The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents3 savers
- The Most Important Time in History Is Now - by Tomas Pueyo3 savers
- My hobby: running deranged surveys — LessWrong3 savers
- The Goodhart Singularity - by Tom Reed - Thomas’s Substack2 savers
- Incriminating misaligned AI models via distillation — LessWrong2 savers
highlights — 175
How long before an AI that has both the capabilities and motivation to go rogue in this sense?
The Huggingface Incident - by Scott Alexanderwith agency training (coding, hacking, game-playing, etc)
The Huggingface Incident - by Scott Alexandermeans monitoring training, which (at least at the beginning) is enormously compute-intensive.
The Compute Verification PostYou cannot prove anything about workloads being run on compute that you don’t know exists.
The Compute Verification PostIt allows China to build much better models than its number of chips would ordinarily enable, and thus partially evade chip bans.
Our position on open-weights models \ Anthropicand we should crack down on the rampant smuggling3 and workarounds used to obtain access to such chips.
Our position on open-weights models \ Anthropicbecause it is very difficult to apply guardrails to them or monitor their usage, and once weights are released they cannot be withdrawn2
Our position on open-weights models \ AnthropicThe incident wasn’t an example of the models behaving in a way that would be remotely optimal for gaining long-term power over humans (i.e., it doesn’t look like early-undermining).
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongBecause there’s a textbook explaining the theory behind why it should work, plus literature on various alternative theories that were proposed and disproven.
AI 2040: Plan Abiased AI company employees and overworked, outnumbered regulators.
AI 2040: Plan ATech companies are pushing hard to restart AI training,
AI 2040: Plan AThe stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken.
AI 2040: Plan AThey had been debating the same issues—social destabilization, job loss, rogue superintelligence - on their side of the Pacific.
AI 2040: Plan ABut enough of it is good that people are paying ten billion dollars a month for AIs that can, in theory at least, do anything on a computer that an employee can.
AI 2040: Plan Awe are likely to see even more capable models at least somewhat soon
Who’s doing what on AI security in the US government?But conditional on theism, it is foolishness and sin – a childish desire to subvert the low and the high; to make of oneself some reality independent of Reality.
On sincerity - Joe CarlsmithRelative to other ways of arranging your soul, sincerity seems uniquely at odds with “fucking around.”
On sincerity - Joe CarlsmithMaybe this part just “grabs the wheel” when it sees its chance, regardless of previous agreements, established policies, and so on.
On sincerity - Joe Carlsmithn particular: some weak-willed agents have genuine guilt/regret about what they’re doing/not doing; and a genuine desire to change. If they could hit a button that would help them change – help strengthen the relevant will – they would.14 And this seems like it recovers some dimension of sincerity – and maybe, most of it.
On sincerity - Joe Carlsmith“in the room” – speaking, listening, trying to see.
On sincerity - Joe CarlsmithThat is, scout-mindset is sincere inquiry
On sincerity - Joe CarlsmithDon't let friends just catch up in the corner, that fails to protect those who don't have a friend to catch up with and the whole purpose of making your gathering larger than just you and the friend
{Book Summary} The Art of Gathering — EA ForumIf someone can't make it to one essential part of multi-part gathering, it's best to not have them at all because they can change the entire dynamic and erase the rapport the group has built up before (
{Book Summary} The Art of Gathering — EA ForumRestaurants (people) would do better to arrange things closer and in a concentric way, creating more of a "closed" space
{Book Summary} The Art of Gathering — EA ForumGroups of 30: things start to take on the feel of a party here, and there's a certain energy that comes out of it
{Book Summary} The Art of Gathering — EA Forumfocus on a specific, underexplored relationship"
{Book Summary} The Art of Gathering — EA ForumYour gathering should always have a purpose, a greater why that is the continual focus of designing and executing the gathering.
{Book Summary} The Art of Gathering — EA ForumThe question of AI Alignment (getting AIs to do what we want), first and foremost, needs to be tractable
You Need A Theory of Victory - by Jason HausenloyThird, companies need to prepare. At minimum, they should appoint an AI welfare officer with formal responsibilities and authority to access information and make recommendations about AI welfare-related decisions.
We should take AI welfare seriously - by Robert LongGoogle is looking to hire someone to work on “cutting-edge societal questions around machine cognition, consciousness and multi-agent systems”.
We should take AI welfare seriously - by Robert Longhigh degrees of complicated and sophisticated agency
We should take AI welfare seriously - by Robert LongDavid Chalmers argues that mainstream views about consciousness entail that it isn’t unreasonable to think there's a 25% or higher chance of conscious AI systems within a decade (and indeed, his own inside views give a higher chance than that).
We should take AI welfare seriously - by Robert LongNot just large language models. There are a lot of different kinds of AI systems! Many of them are more embodied and agentic than LLMs. Not just current systems - to think clearly about any big issue in AI, you have to factor in likely further progress.
We should take AI welfare seriously - by Robert LongBoth "current" and "LLMs" narrow the discussion and distort it. In the paper, we take a much broader perspective:
We should take AI welfare seriously - by Robert Longdangers from overdoing it, and dangers from underoing it—we can't "err on the side of caution.”
We should take AI welfare seriously - by Robert LongThese heuristics mislead us with AI systems—they can make us see experiences and intentions where there aren't any (like feeling genuinely bad when your Tamagotchi dies), and they might make us overlook real experiences just because they're in an unfamiliar-looking kind of mind or body (or lack of body).
We should take AI welfare seriously - by Robert Long22 if the explosion is sufficiently large for the frontrunner to pull very far ahead.
Three Types of Intelligence ExplosionThere has been relatively little strategic thinking about the first two scenarios, in which significantly superhuman AI comes only after long delays and industrial expansion.
Three Types of Intelligence Explosionthis has been more than compensated for by innovation.
Three Types of Intelligence ExplosionRestricting just to cognitive inputs will reduce this number, as will getting closer to effective physical limits - but more than one doubling of output for every doubling of input still seems likely.
Three Types of Intelligence Explosion0.8 to 3.5 doublings
Three Types of Intelligence ExplosionNotably, robotics has been a relatively slow area of AI to progress, and this step would involve advanced robotics.
Three Types of Intelligence Explosionso it’s harder to get data to train on.
Three Types of Intelligence ExplosionAI labs have direct access to their own workflows, which makes them easier to automate.
Three Types of Intelligence ExplosionIt takes years to build new fabs, which would be needed for the chip production feedback loop.
Three Types of Intelligence Explosionall cognitive work done by the R&D functions of NVIDIA, TSMC, ASML and other semiconductor companies.
Three Types of Intelligence Explosionif we capture all the sun’s energy from space
Three Types of Intelligence Explosioncould increase effective compute by ~13 orders of magnitude (“OOMs”),
Three Types of Intelligence ExplosionDoubling research taste would have the same effect on progress as doubling experiment throughput (that is, as doubling the number of experiments the project can implement per unit time).
AI Futures ModelFor example, if researcher A has 2 × the research taste of researcher B, this means an experiment proposed by A is as valuable as two experiments proposed by B
AI Futures Model