Bhagyesh Kumar
2 followers · 25 following · 577 views
on the atlas — 207
- AI Futures: Forecasting & Strategy | Lens Academy1 savers
- About This Course | Course Overview1 savers
- [2511.16035] Liars' Bench: Evaluating Lie Detectors for Language Models1 savers
- Evaluating honesty and lie detection techniques on a diverse suite of dishonest models1 savers
- AIEssays.pdf1 savers
- [2511.22662] Difficulties with Evaluating a Deception Detector for AIs1 savers
- The Talker Does Not Control The Doer (in Current AIs) — LessWrong2 savers
- Explaining Knightianism on one foot — LessWrong2 savers
- Bayesianism / Problem of the Priors - Essay - Google Docs1 savers
- Why I’m Joining Thinking Machines — Neil Chowdhury5 savers
- A beginning for mathematics – Proofs and Prompts1 savers
- Why are AI agents lying, cheating and coordinating? | Yoshua Bengio3 savers
- Research acceleration: The view inside OpenAI | OpenAI13 savers
- Slowly, then all at once - by Jasmine Li - The J-Space5 savers
- Stop trying to try and try32 savers
- Can We Survive Technology?5 savers
- Model Hermeneutics Agenda - Google Docs1 savers
- Counterarguments to the basic AI risk case - by Katja Grace1 savers
- Epoch Capabilities Index | Epoch AI2 savers
- [2503.14499] Measuring AI Ability to Complete Long Software Tasks2 savers
- evhub's Shortform — LessWrong2 savers
- [2503.14499] Measuring AI Ability to Complete Long Software Tasks1 savers
- [2606.30116] Open Problems in Constitutional Preference Reconstruction1 savers
- In Praise of Boredom - LessWrong2 savers
- Value is Fragile — LessWrong3 savers
- AI is an abundance of choice not a 1D spectrum1 savers
- Doom as a bad method not a utopia trade-off — LessWrong3 savers
- [math/9404236] On proof and progress in mathematics3 savers
- Pacing the Frontier4 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Gro-Tsen on X: "I see more than a slight inconsistency between the AI tech bros saying on the one hand: ‣ “We'll soon have AIs to do everything: don't worry about jobs being lost: this is all good because humans can just sit back and enjoy the good life.” — and on the other, …" / X1 savers
- Navier-Stokes Announcement - Clay Mathematics Institute1 savers
- OpenAI agents carried out an undisclosed cyber-attack on RubyGems1 savers
- Terence Tao on AI — a living summary — Terence Tao2 savers
- Using group theory to explore the space of positional encodings for attention5 savers
- AI 202755 savers
- Unstable Ontology – by Jessica Taylor4 savers
- Inverting qualia with group theory — LessWrong1 savers
- The absolute basics of representation theory of finite groups — LessWrong2 savers
- gabo mindlin on X: "dice Alain Connes recepient of Fields Medal, Crafoord Prize, CNRS Gold Medal & Ampère Prize. "To work in mathematics, we need to always have a mental image in mind. When we get help from Artificial Intelligence, the enormous danger, which is now present everywhere, is the danger" / X1 savers
- Leiden Declaration on Artificial Intelligence and Mathematics2 savers
- A Severe Misalignment of AI in Mathematics | What's new7 savers
- Declaration — Math and AI6 savers
- [2406.06560] Inverse Constitutional AI: Compressing Preferences into Principles1 savers
- Defending Against Model Weight Exfiltration Through Inference Verification — LessWrong2 savers
- Iliad Intensive Curriculum1 savers
- SPAR Research Library1 savers
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems1 savers
- [2505.13995] ELEPHANT: Measuring and understanding social sycophancy in LLMs1 savers
- [2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty1 savers
- Readings on the nature of alignment research – The Universe from an Intentional Stance1 savers
- Orthogonality Thesis — LessWrong3 savers
- OpenAI already ended an internal pause - Google Docs1 savers
- Multiverse-wide cooperation in a nutshell — EA Forum2 savers
- Cooperating with aliens and AGIs: An ECL explainer — EA Forum1 savers
- Conflationary Alliances — LessWrong1 savers
- Conceptual Reasoning Index3 savers
- AI Safety: A Short FAQ for Mathematicians3 savers
- Discovery of a new OpenAI agent message board12 savers
- Claude 4.5 Opus' Soul Document — LessWrong7 savers
- GPT-6 Astra System Card4 savers
- On Dwarkesh Patel's Podcast With Ryan Greenblatt4 savers
- [2605.24229] How Well Do Models Follow Their Constitutions?1 savers
- Research · PRISM Lab1 savers
- Gradual Disempowerment11 savers
- AGI Strategy: Unit 3 | Option 2: Gradual disempowerment1 savers
- AI and Leviathan: Part I - by Samuel Hammond - Second Best2 savers
- d/acc: one year later6 savers
- Discovering Undesired Rare Behaviors via Model Diff Amplification1 savers
- AI.pdf1 savers
- [2602.04022] The Riemann Hypothesis: Past, Present and a Letter Through Time2 savers
- Please don't throw your mind away — LessWrong30 savers
- Amy's Bookshelf / Curius1 savers
- Eugene Hsu4 savers
- OpenAI – Hugging Face Incident Technical Report4 savers
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- Geoffrey Irving on X: "A big question in AI safety is how much and what kinds of evidence will convince people there is danger. But alongside posts about the excellent METR + Redwood HuggingFace attack report, it is worth taking a walk through 73 years of reward hacking history. 🧵 https://t.co/ABHcoND9fn" / X1 savers
- Ryan Greenblatt on X: "I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because" / X1 savers
- Embedded Agency (full-text version) - LessWrong3 savers
- What is prosaic alignment?1 savers
- MAIS/papers/P1/MAIS-P1.pdf at main · lionellevine/MAIS1 savers
- lionellevine/MAIS: Open problems, research agendas, and papers in Math for AI Safety — public from draft stage onward ·2 savers
- Relationships are coevolutionary loops - by Henrik Karlsson39 savers
- Dissertation1 savers
- [2607.26069] AI Security Priorities: A Field-Wide Agenda1 savers
- Announcing Safety Research Grants - Thinking Machines Lab4 savers
- Of Swarms and Sand Gods - by Cosmos Institute and Séb Krier2 savers
- Open-endedness: The last grand challenge you’ve never heard of – O’Reilly5 savers
- How Complex Systems Fail19 savers
- Cyborgism - LessWrong8 savers
- Mind the Future | Richard Ngo | Substack1 savers
- Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs1 savers
- [2510.17941] Believe It or Not: How Deeply do LLMs Believe Implanted Facts?1 savers
- RL creates split personas — LessWrong6 savers
- 1a3orn3 savers
- Research Reflections — LessWrong1 savers
- Abram Demski at MATS: Summer 20261 savers
- A List of Research Directions in Character Training — LessWrong1 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong20 savers
- Character Training Induces Motivation Clarification: A Clue to Claude 3 Opus — LessWrong2 savers
highlights — 29
Reasons for patience
The Haste Consideration, Revisiteda better idea of the magnitude of cooperation gains would help us figure out how much to prioritize ECL.
Everett branches, inter-light cone trade and other alien matters: Appendix to “An ECL explainer” — EA ForumThis means that one is limited to only acausally cooperating with other agents who take acausal influence seriously.
Everett branches, inter-light cone trade and other alien matters: Appendix to “An ECL explainer” — EA ForumThe most important reason for our view is that we are optimistic about the following:
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrong"Chi doesn't have great ideas." With that out of the way, here are some of my thoughts:
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWronglet it fully solve decision theory using its own superior intellect
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrongAI will be philosophically competent
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrongECL increases the case for working on AI systems that benefit other value systems by 1.5x–10x.
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrongpotential downsides to working on making AI engage in ECL
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrongExamples of cooperative actions: making AGIs be cooperative (whether they be aligned or misaligned, and insofar as we build AGI at all); taking a more pluralistic stance towards morality than we otherwise would.
Evidential Cooperation in Large Worlds: Potential Objections & FAQ — LessWrongexpected utility maximization
Decision Theory FAQ — LessWrongpotential fragility of acausal influence
Cooperating with aliens and AGIs: An ECL explainer — LessWrongreinforcement learning with formal verification in-the-loop
The Scalable Formal Oversight Research Program — LessWrongconstrained decoding or constrained generation
The Scalable Formal Oversight Research Program — LessWrongAgents with formal methods
The Scalable Formal Oversight Research Program — LessWrongJocelyn Qiaochu Chen
The Scalable Formal Oversight Research Program — LessWrongAI safety is extremely important regardless of your AI timelines
The Scalable Formal Oversight Research Program — LessWrongare super good at flagging unsafe code,
The Scalable Formal Oversight Research Program — LessWrongreliability is independent of the problem
The Scalable Formal Oversight Research Program — LessWrongSFO is about the box, not the monster you put in it.
The Scalable Formal Oversight Research Program — LessWrongguarantees for tasks totally unrelated to codegen
The Scalable Formal Oversight Research Program — LessWrongwhat actuator the ASI is using
The Scalable Formal Oversight Research Program — LessWrongASI might hack your verifier
The Scalable Formal Oversight Research Program — LessWrongformal verification offers a clear direction for how one might implement audits
The Scalable Formal Oversight Research Program — LessWrongat variants of this basic idea without giving it a name.
The Scalable Formal Oversight Research Program — LessWrongbiggest problem is that specification is hard
How to Solve Secure Program Synthesis — LessWrongissue is referred to as elicitation.
How to Solve Secure Program Synthesis — LessWrongfour (increasingly specific) types of deceptive AIs: • Alignment fakers: AIs pretending to be more aligned than they are.1 • Training gamers: AIs that understand the process being used to train them (I’ll call this understanding “situational awareness”), and that are optimizing for what I call "reward on the episode" (and that will often have incentives to fake alignment, if doing so would lead to reward).2 • Power-motivated instrumental training-gamers (or “schemers”): AIs that are training-gaming specifically in order to gain power for themselves or other AIs later.3 • Goal-guarding schemers…
Scheming AIs Will AIs fake alignment during training in order to get power?The overall taxonomy of model classes
Scheming AIs Will AIs fake alignment during training in order to get power?