Micah Carroll
4 followers · 5 following · 1477 views
on the atlas — 87
- From fear to excitement — LessWrong6 savers
- Moravec's paradox11 savers
- Film Study for Research6 savers
- Arrow's impossibility theorem8 savers
- Neural network training makes beautiful fractals | Jascha’s blog11 savers
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum3 savers
- RUDDER - Reinforcement Learning with Delayed Rewards | rudder2 savers
- Kernel method2 savers
- Moment (mathematics)4 savers
- Failures in Kindness — LessWrong8 savers
- Demarcation problem2 savers
- Technology Forecasting: The Garden of Forking Paths · Gwern.net2 savers
- Postpositivism2 savers
- Conclusion — Chapter 7 of Superintelligence Strategy1 savers
- Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategy2 savers
- AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy1 savers
- Introduction — Chapter 2 of Superintelligence Strategy2 savers
- Executive Summary — Chapter 1 of Superintelligence Strategy1 savers
- 'Hyperstition: An Introduction' - 0rphan Drift Archive3 savers
- Byte pair encoding1 savers
- Gmail22 savers
- Resolving internal conflicts requires listening to what parts want — LessWrong1 savers
- Conflicts between emotional schemas often involve internal coercion — LessWrong3 savers
- We learn long-lasting strategies to protect ourselves from danger and rejection — LessWrong3 savers
- Replacing fear - LessWrong2 savers
- the case for CoT unfaithfulness is overstated — LessWrong3 savers
- Replacing guilt3 savers
- Explore - LeetCode1 savers
- Explore - LeetCode1 savers
- Explore - LeetCode1 savers
- AI Act: Participate in the drawing-up of the first General-Purpose AI Code of Practice | Shaping Europe’s digital future1 savers
- A gentle introduction to mechanistic anomaly detection — LessWrong3 savers
- How well do truth probes generalise? — LessWrong3 savers
- Implementing activation steering — LessWrong1 savers
- A Bird's Eye View of the ML Field [Pragmatic AI Safety #2] — AI Alignment Forum1 savers
- Judgments often smuggle in implicit standards — LessWrong3 savers
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlow2 savers
- Claude’s Character \ Anthropic4 savers
- Aditya Kusupati - Google Scholar1 savers
- 2406.003921 savers
- Reward hacking behavior can generalize across tasks — AI Alignment Forum1 savers
- How I select alignment research projects — AI Alignment Forum1 savers
- The St. Petersburg Paradox (Stanford Encyclopedia of Philosophy)1 savers
- Statistical Modeling, Causal Inference, and Social Science3 savers
- You don't get to know what you're fighting for4 savers
- ChatGPT Is Nothing Like a Human, Says Linguist Emily Bender7 savers
- Towards Platform Democracy: Policymaking Beyond Corporate CEOs and Partisan Pressure | Belfer Center for Science and International Affairs3 savers
- Ursula K. Le Guin — A Left-Handed Commencement Address2 savers
- Why I am not a longtermist – Windows On Theory13 savers
- Becoming a magician – Autotranslucence40 savers
- Curius / Onboarding2621 savers
- Half-assing it with everything you've got33 savers
- Meditations On Moloch | Slate Star Codex32 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Augmenting Long-term Memory30 savers
- We Need a New Science of Progress - The Atlantic25 savers
- Notes on Effective Altruism22 savers
- Simulators - LessWrong16 savers
- Lena @ Things Of Interest16 savers
- Cargo Cult Science14 savers
- 12ft13 savers
- Art Is for Seeing Evil | The Point Magazine12 savers
- The garden of forking memes: how digital media distorts our sense of time10 savers
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcast10 savers
- Up and Down the Ladder of Abstraction10 savers
- More Is Different for AI9 savers
- RLHF: Reinforcement Learning from Human Feedback9 savers
- Ads Don't Work That Way9 savers
- Seeing Like an Algorithm — Remains of the Day8 savers
- The Indy8 savers
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blog8 savers
- The "most important century" blog post series7 savers
- You're allowed to fight for something7 savers
- The Taming of Tech Criticism6 savers
- How To Think Real Good | Meta-rationality5 savers
- Critical Atlas of Internet5 savers
- Illusion of explanatory depth4 savers
- Superrationality4 savers
- Nonparametric statistics4 savers
- Hedonic treadmill4 savers
- Ursula K. Le Guin3 savers
- Forecasting transformative AI: the "biological anchors" method in a nutshell3 savers
- Forget About Setting Goals. Focus on This Instead.3 savers
- Behavior Cloning is Miscalibrated - AI Alignment Forum2 savers
- Hawthorne effect2 savers
- ⚡️ Take Back the Future! - Reboot2 savers
- Imre Lakatos2 savers
highlights — 2897
By methodically constraining the most destabilizing moves, states can guide AI toward unprecedented benefits rather than risk it becoming a catalyst of ruin.
Conclusion — Chapter 7 of Superintelligence StrategyAs AI diffuses across countless sectors, societies can raise living standards and individuals can improve their wellbeing however they see fit. Meanwhile leaders, enriched by AI's economic dividends, see even more to gain from economic interdependence and a spirit of détente could take root. During a period of economic growth and détente, a slow, multilaterally supervised intelligence recursion—marked by a low risk tolerance and negotiated benefit-sharing—could slowly proceed to develop a superintelligence and further increase human wellbeing.
Conclusion — Chapter 7 of Superintelligence StrategyStates that act with pragmatism instead of fatalism or denial may find themselves beneficiaries of a great surge in wealth.
Conclusion — Chapter 7 of Superintelligence StrategyThese measures do not halt but stabilize progress.
Conclusion — Chapter 7 of Superintelligence StrategyTo preserve this deterrent and constrain intent, states can expand their arsenal of cyberattacks to disable threatening AI projects. This shifts the focus from "winning the race to superintelligence" to deterrence.
Conclusion — Chapter 7 of Superintelligence StrategyA standoff of destabilizing AI projects may arise by default, but it is not meant to persist for decades or serve as an indefinite stalemate. During the standoff, states seeking the benefits from creating a more capable AI have an incentive to improve transparency and adopt verification measures, thereby reducing the risk of sabotage or preemptive attacks.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyMAIM can be made more stable with unilateral information acquisition (espionage), multilateral information acquisition (verification), unilateral maiming (sabotage), and multilateral maiming (joint off-switch).
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyAIs could reshape the classic tension between security and transparency
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyBy analyzing code and commands on-site, AIs could issue a confidentiality-preserving report or simple compliance verdict, potentially revealing nothing beyond whether the facility is creating new destabilizing models.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategySpeculative but increasingly plausible, confidentiality-preserving AI verifiers offer a path to confirming that AI projects abide by declared constraints without revealing proprietary code or classified material.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyIn a similar spirit, increased transparency spares the broader ecosystem of everyday AI services and lowers the risk of blanket sabotage.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyThe approach of mutual observation echoes the spirit of the Open Skies Treaty, which employed unarmed overflights to demonstrate that neither side was hiding missile deployments.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyCoordinating can help states reduce the risk of maiming datacenters that merely run consumer-facing AI services.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyThis principle of city avoidance would, by analogy, advise placing large AI datacenters in remote areas. If an aggressive maiming action ever occurs, that action would not put cities into the crossfire.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyUnlike kinetic attacks, some of these attacks leave few overt signs of intrusion, yet they can severely disrupt destabilizing AI projects with minimal diplomatic fallout.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyStates could also poison data, corrupt model weights and gradients, disrupt software that handles faulty GPUs, or undermine cooling or power systems. Training runs are non-deterministic and their outcomes are difficult to predict even without bugs, providing cover to many cyberattacks.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyMeasures to prevent smuggling of AI chips keep decisions in the hands of more responsible states rather than rogue actors, which helps preserve MAIM's deterrent value. Like MAD, MAIM requires that destabilizing AI capabilities be restricted to rational actors.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyHowever, clarity about escalation holds little deterrence value if rogue regimes or extremist factions acquire large troves of AI chips.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyFor deterrence to hold, each side's readiness to maim must be common knowledge, ensuring that any maiming act—such as a cyberattack—cannot be misread and cause needless escalation.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyAI powers must clarify the escalation ladder of espionage, covert sabotage, overt cyberattacks, possible kinetic strikes, and so on.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyEven if attempting to harden massive datacenters is extraordinarily prohibitive and unwise, rumors alone can spark fears that a rival is going to risk national security and human security. Formal understandings not to pursue such fortifications help keep the standoff steady.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyBy analogy, we should not leave to chance today's default condition of MAIM: where would-be monopolists, gambling not to cause omnicide, can expect their projects to be disabled.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyStates eventually came to accept that mutual deterrence, while seemingly a natural byproduct of nuclear stockpiling, demanded deliberate maintenance.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyThe net effect may be a stalemate that postpones the emergence of superintelligence, curtails many loss of control scenarios, and undercuts efforts to secure a strategic monopoly, much as mutual assured destruction once restrained the nuclear arms race.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyThis dynamic stabilizes the strategic landscape without lengthy treaty negotiations—all that is necessary is that states collectively recognize their strategic situation.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyA state can expect its AI project to be disabled if any rival believes it poses an unacceptable risk.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyMuch like nuclear rivals concluded that attacking first could trigger their own destruction, states seeking an AI monopoly while risking a loss of control must assume competitors will maim their project before it nears completion.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyShould the supercomputer require an order-of-magnitude AI chip expansion, retrofitting the facility would become prohibitively difficult. Even those with the wealth and foresight to pursue this route would still face the potent risks of insider threats and hacking. In addition, the entire project could be sabotaged during the lengthy construction phase. Last, states could threaten non-AI assets to deter the project long before it goes online.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyCosts balloon as well, diverting funds away from the project's AI chips and pushing total expenditures into the several hundreds of billions. Cooling the world's largest supercomputer underground introduces complex engineering challenges that go well beyond what is required for smaller underground setups.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategySince above-ground datacenters cannot currently be defended from hypersonic missiles, a state seeking to protect its AI-enabled strategic monopoly project might attempt to bury datacenters deep underground to shield them. In practice, the costs and timelines are daunting, and vulnerabilities remain. Construction timelines can stretch to three to five times longer than standard datacenter builds, amounting to several additional years.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyFinally, under dire circumstances, states may resort to broader hostilities by climbing up existing escalation ladders or threatening non-AI assets. We refer to attacks against rival AI projects as "maiming attacks."
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyShould these measures falter, some leaders may contemplate kinetic attacks on datacenters, arguing that allowing one actor to risk dominating or destroying the world are graver dangers, though kinetic attacks are likely unnecessary.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyWhen subtlety proves too constraining, competitors may escalate to overt cyberattacks, targeting datacenter chip-cooling systems or nearby power plants in a way that directly—if visibly—disrupts development.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyStates intent on blocking an AI-enabled strategic monopoly can employ an array of tactics, beginning with espionage, in which intelligence agencies quietly obtain details about a rival's AI projects. Knowing what to target, they may undertake covert sabotage: well-placed or blackmailed insiders can tamper with model weights or training data or AI chip fabrication facilities, while hackers quietly degrade the training process so that an AI's performance when it completes training is lackluster.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyRather than wait for a rival to weaponize a superintelligence against them, states will act to disable threatening AI projects, producing a deterrence dynamic that might be called Mutual Assured AI Malfunction, or MAIM.
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence StrategyWe propose three interconnected lines of effort. First, deterrence: a standoff akin to the nuclear stalemate of MAD, in which no power can gamble human security on an unbridled grab for dominance without expecting disabling sabotage. Next, nonproliferation: just as fissile materials, chemical weapons, and biological agents have long been denied to terrorists by great powers, AI chips and weaponizable AI systems can similarly be kept from rogue actors. Finally, competitiveness: states can protect their economic and military power through a variety of measures including legal guardrails for AI a…
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyThe Manhattan Project assumes that rivals will acquiesce to an enduring imbalance or omnicide rather than move to prevent it.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyYet this facility, easily observed by satellite and vulnerable to preemptive attack, would inevitably raise alarm. China would not sit idle waiting to accept the US's dictates once they achieve superintelligence or wait as they risk a loss of control.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategySuch a project would invoke the Defense Production Act to channel AI chips into a U.S. desert compound staffed by top researchers, a large fraction of whom are necessarily Chinese nationals, with the stated goal of developing superintelligence to gain a strategic monopoly.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyBy contrast, the U.S.-China Economic and Security Review Commission has suggested a more offensive path: a Manhattan Project to build superintelligence.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyThe Monopoly strategy envisions one project securing a monopoly over advanced AI.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyYet militaries desire precisely these hazardous capabilities, making reciprocal restraint implausible. Even with a treaty, the absence of verification mechanisms means the treaty would be toothless; each side, fearing the other's secret work, would simply continue.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyThe voluntary moratorium strategy proposes halting AI development—either immediately or once certain hazardous capabilities, such as hacking or autonomous operation, are detected.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyFinally, they encourage that advanced U.S. model weights continue to be released openly, arguing that even if China or rogue actors use these AIs, no real security threat arises because, they maintain, AI's capabilities are defense-dominant.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyThey likewise oppose export controls on AI chips, claiming such measures would concentrate power and enable a one-world government; in their view, these chips should be sold to whoever can pay, including adversaries.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyProponents of this strategy insist that the U.S. government impose no requirements—including testing for weaponization capabilities—on AI companies, lest it curtail innovation and allow China to win.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyTherefore, loss of control can emerge structurally, as society gradually yields decision-making to automated systems that become indispensable but insidiously acquire more and more effective control. It can occur intentionally, such as a rogue actor unleashing an AI to do harm. It can also occur by accident, when a fast-moving intelligence recursion loops repeatedly ad mortem. All it takes is one loss of control event to jeopardize human security.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyWe should work to have our risk tolerance stay near Compton's threshold rather than in double-digit territory.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyIn sharp contrast, after the defeat of Nazi Germany, Manhattan Project scientists feared the first atomic device might ignite the atmosphere. Robert Oppenheimer asked Arthur Compton what the acceptable threshold should be, and Compton set it at three in a million (a "6σ" threshold)
AI Is Pivotal for National Security — Chapter 3 of Superintelligence StrategyIf the choice is stark—risk omnicide or lose—some might take that gamble.
AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy