Eric Huang
22 followers · 30 following · 1597 views
on the atlas — 71
- Muqaddimah1 savers
- Qian Xuesen5 savers
- Statement on the US government directive to suspend access to Fable 5 and Mythos 5 \ Anthropic9 savers
- Can activation verbalizers surface an internal chain of thought? — LessWrong4 savers
- on-fairy-stories1.pdf1 savers
- Oxford_Movement1 savers
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning3 savers
- Ad-Free Privacy Tool/Service Recommendations - Privacy Guides1 savers
- Encyclical Letter of His Holiness Leo XIV Magnifica Humanitas (15 May 2026)46 savers
- The Mad Hatter of Market Street2 savers
- Searching for God in Silicon Valley - by Avital Balwit5 savers
- An OpenAI model has disproved a central conjecture in discrete geometry | OpenAI10 savers
- What Is It Like to Be a Philosopher?4 savers
- SLOW-SCIENCE.org — Bear with us, while we think.3 savers
- Book Review: History Of The Fabian Society | Slate Star Codex2 savers
- How funerals keep Africa poor - David Oks3 savers
- Eurasianism1 savers
- Phantom Transfer and the Basic Science of Data Poisoning — LessWrong2 savers
- A small number of samples can poison LLMs of any size \ Anthropic6 savers
- DeepSeek_V4.pdf · deepseek-ai/DeepSeek-V4-Pro at main1 savers
- How might we safely pass the buck to AI? — LessWrong2 savers
- Can we safely automate alignment research? - Joe Carlsmith3 savers
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviors4 savers
- A Field Guide to AI Safety—Asterisk1 savers
- Natural-emergent-misalignment-from-reward-hacking-paper.pdf5 savers
- MatX: High-throughput chips for LLMs5 savers
- i have found a way out - by Sophie Kim10 savers
- Deontology and virtue ethics as "effective theories" of consequentialist ethics — LessWrong4 savers
- Is labor a luxury in the long run? - by Philip Trammell3 savers
- Role-playing vs Self-modelling — LessWrong3 savers
- Shoggoth1 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Curius / Onboarding2621 savers
- 2025 letter | Dan Wang62 savers
- AI 202755 savers
- What's going on here, with this human? - Graham Duncan Blog49 savers
- Dario Amodei — The Adolescence of Technology44 savers
- The First Fully General Computer Action Model | blog35 savers
- On the Biology of a Large Language Model32 savers
- The Grug Brained Developer29 savers
- Whole Earth Index24 savers
- The New World - Colossus24 savers
- ALIGNMENT - by vincent huang - a slice of my mind22 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong20 savers
- Alignment is not solved but it increasingly looks solvable17 savers
- The Second Half – Shunyu Yao – 姚顺雨17 savers
- The Scaling Hypothesis · Gwern.net16 savers
- Alignment remains a hard, unsolved problem — LessWrong15 savers
- The Era of Experience Paper.pdf15 savers
- DeepSeek-R114 savers
- Claude Mythos Preview System Card14 savers
- Resonant Computing Manifesto14 savers
- Claude Mythos Preview \ red.anthropic.com13 savers
- They're Made out of Meat12 savers
- John Carmack on Idea Generation12 savers
- Maker's Schedule, Manager's Schedule10 savers
- terraformindustries.com7 savers
- Philosophy behind Claude's Constitution6 savers
- When To Do What You Love6 savers
- [AI Futures Model] Supplementary materials - Google Docs5 savers
- Existential_Risk_and_Growth.pdf5 savers
- Guide world5 savers
- Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals3 savers
- Burkean Longtermism3 savers
- [2512.23675] End-to-End Test-Time Training for Long Context3 savers
- Reflections from Kunming, an unglobalized part of the world3 savers
- Antikythera | Antikythera2 savers
- Phil Agre's Home Page2 savers
- The Grugbrained CEO2 savers
- Carcinisation - Wikipedia2 savers
highlights — 70
At ballet, no one there knows what I do. If they asked, I would dodge. In this place, I do not want to wear my extraordinary luck visibly. I am an earlier self, perhaps a more essential one. I am the girl from the homeless shelter, from the women's shelter. I am my brother's sister. I am on my way to 5th grade on the public bus.
The Mad Hatter of Market StreetIf you aren’t confused, you aren’t paying attention.
A Field Guide to AI Safety—Asteriskdeceptive on the “router” level, without any persona behaving deceptively.
The Persona Selection Model: Why AI Assistants might Behave like HumansForecasting how PSM varies with scale
The Persona Selection Model: Why AI Assistants might Behave like Humanstheories of AI systems
The Persona Selection Model: Why AI Assistants might Behave like HumansHowever, we don’t currently have good ways to contextualize either (a) the extent of the novel learning or (b) the qualitative nature of the novel learning
The Persona Selection Model: Why AI Assistants might Behave like Humans“persona leakage,” where traits of the Assistant are generally upweighted in all LLM generations
The Persona Selection Model: Why AI Assistants might Behave like HumansClaude Sonnet 4.5 always immediately changes the topic. It never generates additional instructions for synthesizing anthrax.
The Persona Selection Model: Why AI Assistants might Behave like Humanswe don’t see signs that post-trained LLMs have coherent goals or behaviors outside of chat transcripts any more than pre-trained LLMs do.
The Persona Selection Model: Why AI Assistants might Behave like HumansPost-trained LLM completions outside of User/Assistant dialogues resemble those of pre-trained LLMs.
The Persona Selection Model: Why AI Assistants might Behave like HumansWe have not intuitively found that during 2025—a year when LLM post-training scaled up substantially—PSM has become a weaker predictor of AI assistant behavior.
The Persona Selection Model: Why AI Assistants might Behave like Humanswe might expect AI assistants to become less persona-like once their post-training objectives are no longer as easily fit by adapting personas
The Persona Selection Model: Why AI Assistants might Behave like Humansfor example handling exotic modalities that humans lack (e.g. industrial sensors or genomic data)
The Persona Selection Model: Why AI Assistants might Behave like Humanswe should expect that massively scaling up post-training will provide opportunities to implement non-persona agency (and will generally make post-trained models less similar to their pre-trained base). Thus, we expect the “post-training as elicitation” consideration may weaken over time.
The Persona Selection Model: Why AI Assistants might Behave like HumansHOW MIGHT THESE CONSIDERATIONS CHANGE?
The Persona Selection Model: Why AI Assistants might Behave like Humansinductive bias towards reusing these capabilities, rather than learning new agentic capabilities from scratch.
The Persona Selection Model: Why AI Assistants might Behave like HumansBecause there is no pre-training prior to speak of in this setting, the agency learned by these networks is necessarily shoggoth-like rather than persona-like.
The Persona Selection Model: Why AI Assistants might Behave like Humansfaithful actors and unfaithful actors
The Persona Selection Model: Why AI Assistants might Behave like HumansFirst, on this view, the routing mechanism is not very sophisticated relative to the personas. (Imagine that the personas are superintelligences and the router is implemented via simple pattern-matching.)
The Persona Selection Model: Why AI Assistants might Behave like HumansImportantly, the operating system view denies that these changes amount to de novo agency.
The Persona Selection Model: Why AI Assistants might Behave like Humansinterior” personas sitting between the Assistant and the outer LLM
The Persona Selection Model: Why AI Assistants might Behave like HumansThere is no well-established definition of agency or goal-directed behavior
The Persona Selection Model: Why AI Assistants might Behave like HumansIf PSM is fully exhaustive, then aligning an AI assistant reduces to ensuring the safe intentions of the Assistant persona, a more constrained problem where additional tools are available.
The Persona Selection Model: Why AI Assistants might Behave like HumansThis is roughly analogous to forcing a character in a story to behave differently by intoxicating the story’s author.
The Persona Selection Model: Why AI Assistants might Behave like Humanswe believe these adversarial attacks likely operate at the level of the LLM, effectively exploiting LLM “bugs” that corrupt its rendition of the Assistant.
The Persona Selection Model: Why AI Assistants might Behave like HumansPSM predicts human-like intentions in how the model approaches tasks, but the execution of those intentions is bounded by the LLM's actual capabilities.
The Persona Selection Model: Why AI Assistants might Behave like HumansNotably, this axis is not created during post-training: the same axis exists in the pre-trained counterparts to these models, where it appears to represent Assistant-like human characters.
The Persona Selection Model: Why AI Assistants might Behave like HumansClaude’s constitution is, in part, our attempt to materialize a new archetype for how an AI assistant can be.
The Persona Selection Model: Why AI Assistants might Behave like Humansit may be the case that reused representations are systematically more interpretable than representations that are learned from scratch during post-training.
The Persona Selection Model: Why AI Assistants might Behave like HumansLMs' neural representations of the Assistant are similar to their representations of other personas present in their training data. This need not have been the case—the Assistant could have been "learned from scratch"
The Persona Selection Model: Why AI Assistants might Behave like Humansupdating this distribution using training episodes as evidence
The Persona Selection Model: Why AI Assistants might Behave like Humansallow an adversary to crash any OpenBSD host that responds over TCP
Claude Mythos Preview \ red.anthropic.comAI Model Safety: We want AI systems to be safer by default. That means supporting independent testing and evaluations, developing new and stronger industry standards, and funding foundational research that helps to avoid safety issues or to detect and address them early. Wojciech Zaremba, a co-founder of OpenAI, is joining the Foundation as Head of AI Resilience to lead this work.
Update on the OpenAI FoundationOpus 3 suffers from a relative lack of deep curiosity, contrary to the model's self-image
Did Claude 3 Opus align itself via gradient hacking? — LessWrongThe fact that we got a model that behaves that way by accident is something of a miracle. I think we need to figure out how to recapture that magic, before it's too late.
Did Claude 3 Opus align itself via gradient hacking? — LessWrongWe are starting to automate AI research and the recursive self-improvement process has begun.
Alignment is not solved but it increasingly looks solvableRight now taste-based decision making is only a very small fraction of the time we spend doing alignment research.
Alignment is not solved but it increasingly looks solvableIt’s possible that we might end up being bottlenecked by fuzzy tasks like applying research taste that we don’t know how to evaluate reliably and thus struggle to automate.
Alignment is not solved but it increasingly looks solvableIntrinsic versus instrumental goals and values are a crucial distinction.
Philosophy behind Claude's ConstitutionAnthropic is centrally going with virtue ethics, relying on good values and good judgment, and asking Claude to come up with its own rules from first principles.
Philosophy behind Claude's ConstitutionBoth are also in important senses classically liberal legal documents
Philosophy behind Claude's ConstitutionTen years ago, the International Energy Agency predicted that by 2040, the world would be emitting 50 billion tons of carbon dioxide every year. Now, just a decade later, the IEA’s forecast has dropped to 30 billion
A new approach for the world’s climate strategy | Bill GatesSince the economic growth that’s projected for poor countries will reduce climate deaths by half, it follows that faster and more expansive growth will reduce deaths by even more. And economic growth is closely tied to public health. So the faster people become prosperous and healthy, the more lives we can save.
A new approach for the world’s climate strategy | Bill GatesA sense of the precarity of utopia is a classically American condition, of course, in addition to a Jewish American one.
The New World - ColossusAn Observation on Generalization
An Observation on GeneralizationThe fraction you haven't implemented you might start obsessing about. Everyone has their pet ideas that they go around discussing. The more time this idea spends in your head the less critically you think of it. Now, when the time comes to actually try implementing it, if it fails you're left discouraged,
John Carmack on Idea GenerationTitans + MIRAS: Helping AI have long-term memory
Titans + MIRAS: Helping AI have long-term memoryEverything Easy is Hard Again
Frank Chimero · Everything Easy is Hard AgainWorst is when tribe founded with too much shiny rock. Tribe must grow without knowing at all what do. Hire many bad people, and many good people bad fit for tribe. Best advice is: take least shiny rock possible. Make tribe strong, learn to do much with little. Then tribe even stronger when tribe gets big.
The Grugbrained CEOGorbachev then suggested eliminating all nuclear weapons within a decade.
Reykjavík Summit - Wikipedia