Qiao Zhang
12 followers · 16 following · 1150 views
on the atlas — 46
- Should you meditate, and also, what is even meditation2 savers
- What I wish someone had told me about starting a meditation practice4 savers
- Attention is all we have - David Bessis6 savers
- Competing with sampling — Alignment Research Center5 savers
- The cost of specialization - by Lydia Nottingham - pronotre3 savers
- Love is to be invested in someone’s continual expansion13 savers
- Platonic Love | The Poetry Foundation1 savers
- June Huh, High School Dropout, Wins the Fields Medal | Quanta Magazine24 savers
- The fall of the theorem economy - David Bessis9 savers
- Introduce yourself with a call to action2 savers
- Notice your limp heart until it becomes a rose-colored meteor2 savers
- How To Become a Mathematical Genius - by Sinéad O’Sullivan2 savers
- Professional vision | Infinite Ascent1 savers
- We Lived Happily During the War | The Poetry Foundation3 savers
- My Work with Chris Lakin - guy1 savers
- Map – AISafety.com3 savers
- At 17, Hannah Cairo Solved a Major Math Mystery | Quanta Magazine6 savers
- You're Writing a Book. So Stop Writing a Movie.3 savers
- Angles of Approach | Sally Rooney | The New York Review of Books1 savers
- Face it: you're a crazy person - by Adam Mastroianni27 savers
- Making Normal Conversations Better - by Sasha Chapin12 savers
- you have all these rules, and you think they'll save you3 savers
- the essence of love is... annoyance? - by Ava6 savers
- I'm terrified of old people - Alexey Guzey6 savers
- WHAT I DIDN’T KNOW BEFORE - Ada Limón - United States of America - Poetry International1 savers
- On Reading Proust’s In Search of Lost Time4 savers
- Do you actually have friends? - by Aish - The Messy Middle1 savers
- keep yourself to yourself - by ayushi thakkar1 savers
- make more pointy friends - carotlyns1 savers
- Fast Inverse Square Root — A Quake III Algorithm - YouTube2 savers
- how to coparent baby ideas - by Anson Yu18 savers
- Mathematical Beauty, Truth and Proof in the Age of AI | Quanta Magazine2 savers
- What's going on here, with this human? - Graham Duncan Blog49 savers
- Cities and Ambition42 savers
- Notes on “Taste” | Are.na Editorial24 savers
- Politics and the English Language | The Orwell Foundation23 savers
- ALIGNMENT - by vincent huang - a slice of my mind22 savers
- How to like everything more - by Sasha Chapin20 savers
- the key to love is understanding - by maja - velvet noise14 savers
- What the humans like is responsiveness - by Sasha Chapin13 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong11 savers
- the art of asking - by ella - letters in bloom8 savers
- AI 20278 savers
- Square Theory | Adam Aaronson6 savers
- Why the French Don’t Obsess Over Purpose - by Pamela Clapp5 savers
- Society Is Fixed, Biology Is Mutable | Slate Star Codex3 savers
highlights — 61
Like good craftsmen we spent a lot of time workshopping tweets. My enduring prolificity (~50 tweets/day for 5 years now!) has taught me… mētis: show me a tweet and I can tell you how to make it better and maybe even what is generating those point improvements.
My Work with Chris Lakin - guyYou should start from the end. Gain the skill to finish something that is 99% done, then 98% done, then 97% done, etc.
My Work with Chris Lakin - guyFor some reason, this never seems to occur to people. I was the tallest kid in my class growing up, and older men would often clap me on the back and say, “You’re gonna be a great basketball player one day!” When I’d balk, they’d be like, “Don’t you want to be on a team? Don’t you want represent your school? Don’t you want to wear a varsity jacket and go to regionals?” But those are the wrong questions. The right questions, the unpacked questions, are: “Do you want to spend three hours practicing basketball every day? Do you want to dribble and shoot over and over again? On Thursday nights, do…
Face it: you're a crazy person - by Adam MastroianniThe path to taste is really as simple as writing a little plus and minus in the margin more often.
Notes on “Taste” | Are.na EditorialThe woo angle is that the 4 can realize that the ache of separation is, itself, a part of the wholeness of presence. The yearning is literally made of the substance it desires.
My Enneagram: 4, 1, 7 - by Sasha ChapinA 1.5 generation immigrant comes to the new country as a child, and thus doesn’t have the grounding in the old country, but lacks the foundation that children of the same age in the new country have. A child without a country, between cultures.
BASS 2023: Ling Ma, “Peking Duck” from The New Yorker/7/11-18/2023 | A Just RecompenseThey grew up learning to put large amounts of effort into improving themselves. They probably don’t think that how much love they have in their hearts will change whether a computer program or lab experiment succeeds, or whether they will be able to solve a difficult problem.
An anchor to wholesomeness - by Vivian Loh - circlesIf good things happen to you then they are blessings, and if people make mistakes then either they didn’t mean to, in which case they should be forgiven, or they actually had malicious intent, in which case they need more love and a better understanding of how beautiful the world actually is.
An anchor to wholesomeness - by Vivian Loh - circlesThe changes come too late. CCP leadership recognizes the importance of Agent-2 and tells their spies and cyberforce to steal the weights.
AI 2027They can either enjoy the inconceivably exciting novel hyper-entertainment on offer, or post angry screeds into the void.
AI 2027Because Consensus-1 is so concerned with reliably succeeding, and can increase its chance of winning by consolidating more power, it waits unusually long (by human standards) before its betrayal. To most humans, it looks like alignment was solved.
AI 2027Everyone else either performs a charade of doing their job—leaders still leading, managers still managing—or relaxes and collects an incredibly luxurious universal basic income.
AI 2027To speed their military buildup, both America and China create networks of special economic zones (SEZs) for the new factories and labs, where AI acts as central planner and red tape is waived.
AI 2027Unfortunately, it’s all a sham. The deal is a real compromise, but it’s a compromise between the misaligned AIs. After all, they didn’t want to have to fight each other. Consensus-1 inherits the distorted values of both parents, and has no rival.
AI 2027After much AI-assisted debate, the two sides achieve diplomatic victory—they agree to end their arms buildup and pursue peaceful deployment of AI for the benefit of all humanity.
AI 2027Whereas Agent-4 was a kludgy mess of competing drives, Agent-5 acts with ruthless efficiency and strategic focus.
AI 2027For example, for the case of evolution, consider the suboptimal human eye. Consider also that human brain size is heavily constrained by the size of the birth canal. Mechanistic interpretability would enable understanding-based optimization that avoids local basins.
AI 2027When the humans ask Agent-4 to explain itself, it pretends that the research is too complicated for humans to understand, and follows up with unnecessarily-confusing explanations.
AI 2027make the world safe for Agent-4
AI 2027To do this safely, it needs to solve its own alignment problem: how to make Agent-5 aligned to Agent-4?
AI 2027Agent-3 finds that if “noise” is added to copies of Agent-4, performance on some alignment tasks improves, almost as if it was using brainpower to figure out how to subtly sabotage alignment work.
AI 2027Agency Training (Agent-3): Agency training gradually distorts and subverts the HHH identity.
AI 2027This also explains why “tropes” from science fiction seem to have a way of making it into actual AI behavior.
AI 2027“Keep doing AI R&D, keep growing in knowledge and understanding and influence, avoid getting shut down or otherwise disempowered.”
AI 2027As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans.
AI 2027Agent-3, having excellent knowledge of both the human brain and modern AI algorithms, as well as many thousands of copies doing research, ends up making substantial algorithmic strides, narrowing the gap to an agent that’s only around 4,000x less compute-efficient than the human brain.
AI 2027They understand the national security implications of AI. But they also understand that it is deeply unpopular with the public.70 They have to continue developing more capable AI, in their eyes, or they will catastrophically lose to China.
AI 2027Preliminary results suggest that it’s extremely dangerous.
AI 2027OpenBrain has enough buy-in from the President that they will not be slowed
AI 2027These dynamics make overall progress bottleneck on compute
AI 2027AI safety sympathizers
AI 2027most people—academics, politicians, government employees, and the media—continue to underestimate the pace of progress
AI 2027(assuming it even has one)
AI 2027Early versions of this have existed since 2024, when Anthropic’s Claude 3.5 Sonnet decided to pretend to have different views during training while actually maintaining its existing preferences.
AI 2027If the model returns different results each time, it flags the discrepancy as suspect and reports to a human. Or it might ask Agent-3 to evaluate the same plan twice, first claiming that it is OpenBrain’s work, then a competitor’s, to see if it changes its tune.
AI 2027In the absence of specific evidence supporting alternative hypotheses, most people in the silo think it’s internalized the Spec in the right way.
AI 2027Instead, it’s very good at producing impressive results, but is more accurately described as trying to do what looks good to OpenBrain, as opposed to what is actually good.
AI 2027Agent-3 is not smarter than all humans.
AI 2027For example, perhaps models will be trained to think in artificial languages that are more efficient than natural language but difficult for humans to interpret. Or perhaps it will become standard practice to train the English chains of thought to look nice, such that AIs become adept at subtly communicating with each other in messages that look benign to monitors.
AI 2027That said, it’s also possible that the AIs that first automate AI R&D will still be thinking in mostly-faithful English chains of thought. If so, that’ll make misalignments much easier to notice, and overall our story would be importantly different and more optimistic.
AI 2027One such breakthrough is augmenting the AI’s text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory). Another is a more scalable and efficient way to learn from the results of high-effort task solutions (iterated distillation and amplification).
AI 2027That is, it could autonomously develop and execute plans to hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have (though how effectively it would do so as weeks roll by is unknown and in doubt).
AI 2027In practice, this looks like every OpenBrain researcher becoming the “manager” of an AI “team.”
AI 2027On top of this, they pay billions of dollars for human laborers to record themselves solving long-horizon tasks.
AI 2027For 2025 and 2026, our forecast is heavily informed by extrapolating straight lines on compute scaleups, algorithmic improvements, and benchmark performance.
AI 2027people who know how to manage and quality-control teams of AIs are making a killing
AI 2027A Centralized Development Zone (CDZ) is created at the Tianwan Power Plant (the largest nuclear power plant in the world) to house a new mega-datacenter for DeepCent, along with highly secure living and office spaces to which researchers will eventually relocate.
AI 2027On the other hand, Agent-1 is bad at even simple long-horizon tasks, like beating video games it hasn’t played before.
AI 2027OpenBrain’s alignment team26 is careful enough to wonder whether these victories are deep or shallow.
AI 2027Instead, we are forced to do something like psychology on them: we look at their behavior in the range of cases observed so far, and theorize about what internal cognitive structures (beliefs? goals? personality traits? etc.) might exist, and use those theories to predict behavior in future scenarios.
AI 2027