Yudhister Joel Kumar
131 followers · 206 following · 5238 views
on the atlas — 565
- Chernoff bound1 savers
- Sam_Zemurray1 savers
- [2501.16946] Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development3 savers
- Making deals with early schemers — LessWrong2 savers
- The Artificial Self6 savers
- John Rawls2 savers
- Tensor product1 savers
- Home | seL41 savers
- Working through a small tiling result — LessWrong3 savers
- Metagaming matters for training, evaluation, and oversight2 savers
- A Straussian reading of The Adolescence of Technology | Zhengdong4 savers
- Open-minded updatelessness — LessWrong2 savers
- Nelson Elhage1 savers
- Thoughts on Updatelessness – The Universe from an Intentional Stance1 savers
- [2310.17813] A Spectral Condition for Feature Learning2 savers
- Bonsai-demo/1-bit-bonsai-8b-whitepaper.pdf at main · PrismML-Eng/Bonsai-demo2 savers
- State of Brain Emulation Report 2025 | Research Overview3 savers
- There should be ‘general managers’ for more of the world’s important problems11 savers
- The_Final_Offshoring.pdf2 savers
- Understanding the Neural Tangent Kernel – EigenTales3 savers
- Quantifying Truesight With SAEs · Gwern.net2 savers
- Patrick Collison on programming languages, AI, and Stripe's biggest engineering decisions - YouTube1 savers
- Effective Startup Ideation4 savers
- arg min | Ben Recht | Substack1 savers
- Building Technology to Drive AI Governance6 savers
- Tim Roughgarden's Lecture Notes2 savers
- A Travelogue from India and China7 savers
- Different senses in which two AIs can be “the same” — LessWrong1 savers
- MatX: High-throughput chips for LLMs5 savers
- Claude Mythos Preview System Card3 savers
- Think Tanks Have Defeated Democracy6 savers
- The Origins and Motivations of Univalent Foundations - Ideas | Institute for Advanced Study2 savers
- Deontology and virtue ethics as "effective theories" of consequentialist ethics — LessWrong4 savers
- Conversational Cultures: Combat vs Nurture (V2) — LessWrong1 savers
- Agam Bhatia1 savers
- Bodega Bay Nuclear Power Plant1 savers
- [2604.06366] Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks1 savers
- Notes on structured concurrency, or: Go statement considered harmful — njs blog6 savers
- Role-playing vs Self-modelling — LessWrong3 savers
- Episcopal Church (United States)2 savers
- Harvard University1 savers
- Thomas Nagel1 savers
- What Is It Like to Be a Bat?3 savers
- (PDF) A Fine is a Price1 savers
- An Alignment Journal: Features and policies · Alignment Journal Blog1 savers
- 2e9f9cde1b709281a06dd14f679e4c51-Paper-Conference.pdf1 savers
- An Alignment Journal: Features and policies — LessWrong1 savers
- Intuitive Self-Models - LessWrong3 savers
- Beliefs are Chosen to Serve Goals — LessWrong3 savers
- nickbostrom.com/evolutionary-optimality.pdf2 savers
- Galaxy brain resistance18 savers
- Inference to the Best Explanation (article)1 savers
- Spiral Dynamics1 savers
- Dialogue: Is there a Natural Abstraction of Good? — LessWrong1 savers
- The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents3 savers
- Smart Policy: Cognitive Enhancement and the Public Interest1 savers
- Letter from Utopia9 savers
- The Unilateralist’s Curse and the Case for a Principle of Conformity1 savers
- Reformative Hypocrisy, and Paying Close Enough Attention to Selectively Reward It. — LessWrong1 savers
- 10 pieces of advice for children - Nina Panickssery3 savers
- [2604.04891] Muon Dynamics as a Spectral Wasserstein Flow1 savers
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWrong1 savers
- College at Age Sixteen: What Is Intelligence For?2 savers
- Does novel understanding imply novel agency / values?2 savers
- Mathematics for humans2 savers
- Gardens of the Soul — Laura Deming5 savers
- microgpt6 savers
- On The Independence Axiom — LessWrong4 savers
- The Coastal Elites Are Right, Actually - by A.M. Hickman2 savers
- The Shard Theory of Human Values10 savers
- 4. Sets and Functions — Mathematics in Lean 0.1 documentation2 savers
- mathematicians.pdf2 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- The Gap Map34 savers
- [2308.12108] The Local Learning Coefficient: A Singularity-Aware Complexity Measure4 savers
- Which Future?1 savers
- Singular Learning Theory Seminar2 savers
- We were wrong about convergence3 savers
- I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines — LessWrong2 savers
- Mathematics in the Library of Babel — Daniel Litt6 savers
- My journey to the microwave alternate timeline — LessWrong4 savers
- Why You Don’t Believe in Xhosa Prophecies — LessWrong5 savers
- matrix-book.pdf1 savers
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum4 savers
- [2601.21571] Shaping capabilities with token-level data filtering5 savers
- Basics of How Not to Die — LessWrong2 savers
- Moltbook is the most interesting place on the internet right now8 savers
- On Agency | Sebastian Farquhar1 savers
- Stuff you should have been taught in college but weren’t – Casey Handmer's blog8 savers
- Base Camp for Mt. Ethics1 savers
- On “compilation errors” in mathematical reading, and how to resolve them | What's new1 savers
- On Doubling9 savers
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum3 savers
- How Hangzhou Spawned Deepseek and Unitree2 savers
- An Opinionated Guide to Using Anki Correctly — LessWrong7 savers
- [1805.00909] Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review2 savers
- Relentlessly Resourceful6 savers
- Did California's Fast Food Minimum Wage Reduce Employment? | NBER2 savers
- Arc Institute’s first virtual cell model: <span style="font-variant: small-caps">S<span style="font-weight: bolder">tate</span></span> | Arc Institute3 savers
- What is AIXI?2 savers
highlights — 2563
We think it’s very reasonable to spend an average of ~$3k per paper on reviewer payments.
An Alignment Journal: Features and policies · Alignment Journal BlogBut if they're doing strictly more work than a reviewer, why aren't we paying them?
An Alignment Journal: Features and policies — LessWrongThe most important and potentially surprising observation is that general capacities emerge abruptly — for small the model allocates its capacity entirely to per-task circuits, but above some threshold the model suddenly allocates a large fraction of total capacity to general circuits.
An Alignment Journal: Features and policies — LessWrongThis structure allows him to define the availability of internal "communicative alternatives" as a component of fairness. The availability of certain internal communicative protocols, and certain options to self-modify internally, forms the particularly novel parts of his approach.
An Alignment Journal: Features and policies — LessWrongIf a problem penalizes a procedure for properties other than its pattern of decisions, this seems unfair.
An Alignment Journal: Features and policies — LessWrongLoosely speaking, their conjecture states that if a neural network with nonlinearities has the rare property that no input in maps to an all-negative output, then there is a concise structural explanation of that explains this property.
An Alignment Journal: Features and policies — LessWrongIn order for a neural network to behave like a random function, the covariance matrix of this distribution must be the identity matrix. The authors show that this happens whenever the activation function has an expected value of 0 under the Gaussian measure (e.g., the function).
An Alignment Journal: Features and policies — LessWrongIts main goal is to help a potential reader decide — on the paper’s merits — if the paper is worth reading.
An Alignment Journal: Features and policies — LessWrongOne good reason to save a humble bumble bee that strays into your house is to play C with the great unknown.
Base Camp for Mt. EthicsWe should be modest, willing to listen and learn. We should not too headstrongly insist on having too much our way. Instead, we should be compliant, peace-loving, industrious, and humble vis-a-vis the cosmic host.
Base Camp for Mt. EthicsWe should contribute public goods to the cosmic resource pool, by securing resources and (later) placing them under the control of cosmic norms. Prevent xrisk and build AI?
Base Camp for Mt. EthicsFor example, if you get accused of violating a local rule, you could try to argue that this rule conficts with a higher-level rule and is therefore not normative. (Note that if this argument requires more than a relatively short extrapolation distance, it will probably not work.)
Base Camp for Mt. EthicsThe question now arises as to what morality requires of us if we should fnd ourselves, as we almost certainly do, in a situation in which the local norms are not completely in accord with higher-level norms
Base Camp for Mt. EthicsThe reason is that the presumably vastly superior epistemic abilities of the cosmic host, along with there plausibly having been ample time available for an equilibrium to have been reached, suggests that the hypothetical steps involved in defning the extrapolated norms may have been actually taken.
Base Camp for Mt. EthicsOn the other hand, there is a sense in which extrapolation distance—roughly: the degree to which moral norms “idealize” a simplifed version of the actually occuring patterns of moral blame and credit allocation—may be shorter for the norms of the cosmic host:
Base Camp for Mt. EthicsThe important sense is that since communities can be nested, the normative structures that they develop can likewise form an embedding structure. It is this sense of hierarchy that I will explore here.
Base Camp for Mt. Ethicsbut this has no bearing on whether “Might makes right” is true or whether the counterfactual mentioned above is true, because the semantic content of their utterance is different from the semantic content of our utterance.
Base Camp for Mt. Ethicst might be that in our language terms like “moral wrongness”, “moral permissibility”, etc., function as Kripkean rigid designators .
Base Camp for Mt. EthicsFor example, the rights to life and property, the notion of rule-of-law, and animal welfare (which was greatly advanced in political practice by actual neuroscience).
Dialogue: Is there a Natural Abstraction of Good? — LessWrongFor example, a calculus student can "grok" the concept of differentiation and still make a mistake on an exam. But the pattern of mistakes they make is different, and if they continue to practice, the student who has "grokked" it is much more likely to improve on the areas where they tend to mess up.
Dialogue: Is there a Natural Abstraction of Good? — LessWronga) I think it's about having the correct information-integration and decision-making process, which subsumes both having good intents upstream and making good choices downstream.
Dialogue: Is there a Natural Abstraction of Good? — LessWrongWe must draw a sharp line between shrinking the gene pool and growing the gene pool, and between coercive and non-coercive approaches.
Dialogue: Is there a Natural Abstraction of Good? — LessWrongHowever, these are moral truths, not policies. As a matter of policy, regulating medical access to abortions tends not to produce the desired outcomes.
Dialogue: Is there a Natural Abstraction of Good? — LessWrongYou don’t need to make “rite of passage”-style mistakes (e.g. drinking or taking drugs, getting into bad relationships, cramming for exams, ignoring your health, becoming a socialist)! Avoid them. Adults often say things like “all kids make mistake X and then gradually learn not to do X”. If you observe that many people who do something later regret doing it, strongly consider not doing it ever yourself unless you have good information that your situation is different. You don’t need to learn from experience, only sheep do! As a thinking human, you can also learn from others’ experiences. When…
10 pieces of advice for children - Nina PanicksseryI recommend Eric Drexler's writing on AI, which I host here to ward against link-rot:
Owain Evans, AI Alignment researcherIf I did, it would be frivolous to care how sentences sound. But in practice it feels the opposite of frivolous. Fixing sentences that sound bad seems to help get the ideas right.
Good WritingThe symmetry breaks because the Assistant and JFK are very different as self-models. The Assistant is not perfect or completely true, but it is a far more viable self-model than JFK.
Role-playing vs Self-modelling — LessWrongWhere the models would most likely come apart is accuracy of self-prediction and the lack of detailed memories for the alternative "Self."
Role-playing vs Self-modelling — LessWrongOpenAI’s C.E.O. for AGI Deployment
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerAltman pitched Alexandr Wang, now the head of A.I. at Meta, on a leadership role, telling him that Jeff Bezos, the founder of Amazon, could head the new company.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerCorporate investigations aim to confer legitimacy. At private companies, their findings are sometimes not written down—this can be a way to limit liability. But in cases involving public scandals there is often a greater expectation of transparency. Before Kalanick left Uber, in 2017, its board hired an outside firm, which released a thirteen-page summary to the public.
Sam Altman May Control Our Future—Can He Be Trusted? | The New Yorker“He’s unconstrained by truth,” the board member told us. “He has two traits that are almost never seen in the same person. The first is a strong desire to please people, to be liked in any given interaction. The second is almost a sociopathic lack of concern for the consequences that may come from deceiving someone.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerYoon, the former board member, argued that Altman was “not this Machiavellian villain” but merely, to the point of “fecklessness,” able to convince himself of the shifting realities of his sales pitches. “He’s too caught up in his own self-belief,” she said. “So he does things that, if you live in the real world, make no sense. But he doesn’t live in the real world.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerAmodei and Sutskever were never close friends, but they reached similar conclusions. Amodei wrote, “The problem with OpenAI is Sam himself.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerThey have since taken on a legendary status in Silicon Valley; in some circles, they are simply called the Ilya Memos.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerBut even Jobs never told his customers that if they didn’t buy his brand of MP3 player everyone they loved would die. When Altman was twenty-three, in 2008, Graham, his mentor, wrote, “You could parachute him into an island full of cannibals and come back in 5 years and he’d be the king.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerBut in that post, called “The Gentle Singularity,” he adopted a new tone, replacing existential terror with ebullient optimism. “We’ll all get better stuff,” he wrote. “We will build ever-more-wonderful things for each other.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New Yorker“It just was kind of completely ignored,” Jacob Hilton, an OpenAI researcher at the time, said.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerThere was an all-hands meeting, the former employee continued, “where Ilya gets up and he’s, like, Hey, everyone, there’s going to be a point in the next few years where basically everyone at this company has to switch to working on safety, or else we’re fucked.” But the superalignment team was dissolved the following year, without completing its mission.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerFurthermore, a researcher on the team said, “most of the superalignment compute was actually on the oldest cluster with the worst chips.” The researchers believed that superior hardware was being reserved for profit-generating activities.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerWeeks after the paper was published, one of its authors, a Ph.D. student at the University of California, Berkeley, got an e-mail from Altman, who said that he was increasingly worried about the threat of unaligned A.I.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerAmodei’s notes describe escalating tense encounters, including one, months later, in which Altman summoned him and his sister, Daniela, who worked in safety and policy at the company, to tell them that he had it on “good authority” from a senior executive that they had been plotting a coup. Daniela, the notes continue, “lost it,” and brought in that executive, who denied having said anything. As one person briefed on the exchange recalled, Altman then denied having made the claim.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerAlthough Amodei, who was leading the company’s safety team, had helped to pitch the deal to Bill Gates, many people on the team were anxious about it, fearing that Microsoft would insert provisions that overrode OpenAI’s ethical commitments.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerHe jumped out of his chair, ran down the hall, and told his fellow-researchers, “Stop everything you’re doing. This is it.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerAll three men confirmed that the pact existed, though Brockman said that it was informal. “He unilaterally told us that he’d step down if we ever both asked him to,” he told us. “We objected to this idea, but he said it was important to him. It was purely altruistic.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerBut, unbeknownst to them, he also struck a secret handshake deal with Brockman and Sutskever: Altman would get the C.E.O. title; in exchange, he agreed to resign if the other two deemed it necessary.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerGoogle offered Sutskever six million dollars a year, which OpenAI couldn’t come close to matching. But, Altman boasted, “they unfortunately dont have ‘do the right thing’ on their side.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerA collection of more than two hundred pages of documents related to Amodei, including those notes and internal e-mails and memos, has been circulated by colleagues in Silicon Valley but never before disclosed publicly. In his notes, Amodei wrote that Altman’s goal was to build “an AI lab that would be focused on safety (‘maybe not right away, but as soon as it can be’).”
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerNevertheless, as recently as 2021, a Securities and Exchange Commission filing listed Altman as the chairman of Y Combinator. (Altman says that he wasn’t aware of this until much later.)
Sam Altman May Control Our Future—Can He Be Trusted? | The New Yorker“We didn’t have the legal power to fire anyone. All we could do was apply moral pressure.”
Sam Altman May Control Our Future—Can He Be Trusted? | The New Yorker