Evie Hu
18 followers · 15 following · 533 views
on the atlas — 135
- on doing Real work - by jessica dai7 savers
- The Locally Optimal Discursive Posture — LessWrong3 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Countering misuse of AI: September 2026 / Anthropic \ Anthropic14 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance3 savers
- Detecting misbehavior in frontier reasoning models | OpenAI7 savers
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes2 savers
- Structure and Interpretation of Deep Networks3 savers
- [2601.04603] Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks2 savers
- Cost-Effective Constitutional Classifiers via Representation Re-use2 savers
- Training a Misaligned Reward Seeker3 savers
- Weierstrass function2 savers
- Towards a Typology of Strange LLM Chains-of-Thought5 savers
- Zipf–Mandelbrot law1 savers
- The Unintelligibility is Ours: Notes on Chain of Thought7 savers
- Privilege, Dominance, and Personas - by Derek Shiller1 savers
- The Rise and Fall of Agent Civilizations2 savers
- Brian's Brain1 savers
- Minimum description length3 savers
- On-Policy Distillation - Thinking Machines Lab29 savers
- [2608.21664] Measuring Activation Control in Large Language Models1 savers
- Loss of Oversight: How AI Systems May Become Harder to Audit, Monitor, and Investigate — LessWrong2 savers
- We need 3rd party Training-Run Assessments — LessWrong2 savers
- [2602.15001] Boundary Point Jailbreaking of Black-Box LLMs1 savers
- pdf2 savers
- Safety and alignment in an era of long-horizon models | OpenAI7 savers
- Detecting and reducing scheming in AI models | OpenAI1 savers
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations1 savers
- [2404.15758] Let's Think Dot by Dot: Hidden Computation in Transformer Language Models1 savers
- How we monitor internal coding agents for misalignment | OpenAI2 savers
- Incidents | Rogue AI Tracker1 savers
- RL creates split personas — LessWrong6 savers
- Pre-deployment auditing can catch an overt saboteur2 savers
- Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments — LessWrong1 savers
- Digital_Consciousness_Model.pdf1 savers
- Antitrust application to AI companies doing safety - Google Docs1 savers
- Guidelight's Control Assessment of Frontier AI Companies | Guidelight AI Standards1 savers
- Psychokinetics | Manav B. Ponnekanti8 savers
- The best general advice on earth « the jsomers.net blog19 savers
- The End-State Fallacy: Where Is AI Security Headed?4 savers
- nn-notes.pdf4 savers
- Safe Pareto Improvements Research Agenda — Center on Long-Term Risk1 savers
- Overview | Shallow Review 20257 savers
- philpapers.org/archive/GOLLFC-2.pdf2 savers
- Safe Pareto Improvements for Delegated Game Playing1 savers
- Mathematics in the age of AI - Public lecture, International Congress of Mathematicians 20263 savers
- Pain is not the unit of Effort | Radimentary4 savers
- Compression and Intelligence — Ryan Greene5 savers
- RL & search is a terrifying way to build AGI (an FAQ) — AI Alignment Forum1 savers
- [2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language1 savers
- Training Large Language Models to Reason in a Continuous Latent Space3 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning8 savers
- Neural Networks, Manifolds, and Topology -- colah's blog16 savers
- 1412.02331 savers
- Understanding deep learning requires rethinking generalization2 savers
- Why Momentum Really Works7 savers
- Toy Models of Superposition30 savers
- Modular Pretraining Enables Access Control3 savers
- gdm-ai-control-roadmap.pdf2 savers
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors3 savers
- A Mathematical Framework for Transformer Circuits39 savers
- All Roads Lead to Robotics | Eric Jang9 savers
- Politics and the English Language | The Orwell Foundation23 savers
- Just make the straw bigger1 savers
- Paper Feed1 savers
- I Reviewed Hundreds of AI Safety Applications. Here's What Actually Matters | Georg Lange1 savers
- [2602.23163] A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring1 savers
- ‘Give Away Your Legos’ and Other Commandments for Scaling Startups | First Round Review8 savers
- First, make rice | Seth's Blog1 savers
- The Cook and the Chef: Musk's Secret Sauce — Wait But Why19 savers
- Total Eclipse - Annie Dillard4 savers
- A Straussian reading of The Adolescence of Technology | Zhengdong4 savers
- The Future Worth Building Is Human - Thinking Machines Lab23 savers
- Write code like a human will maintain it1 savers
- Toward A Public Science of Model Behavior | Transluce AI3 savers
- Verbalizable Representations Form a Global Workspace in Language Models24 savers
- Soft Actor-Critic — Spinning Up documentation4 savers
- [1711.05101] Decoupled Weight Decay Regularization1 savers
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness3 savers
- Part 1: Key Concepts in RL — Spinning Up documentation11 savers
- Humans are not automatically strategic — LessWrong14 savers
- LoRA Without Regret - Thinking Machines Lab37 savers
- Fast · Patrick Collison48 savers
- A reading list for generalists — LessWrong18 savers
- Mentorship and the art of actionable advice | Yanir Seroussi – AI/ML Success Architect1 savers
- In Solidarity with Library Genesis and Sci-hub1 savers
- A system overview for near-term, low-trust AI compute verification — LessWrong2 savers
- [2606.04071] Covert Influence Between Language Models3 savers
- Sequent: scale and automation for higher confidence in alignment — AI Alignment Forum2 savers
- Alignment Is Proven To Be Solvable - by SE Gyges3 savers
- Measuring no CoT math time horizon (single forward pass)2 savers
- Turns Out, You Can’t Just Bomb a Datacenter - by Sophie Kim2 savers
- Built to benefit everyone: our plan | OpenAI7 savers
- Soft Nationalization: how the USG will control AI labs — LessWrong2 savers
- Is there a Half-Life for the Success Rates of AI Agents? — Toby Ord1 savers
- Alex Tamkin - Tips for New Researchers3 savers
- Learning_curve1 savers
- [1709.06560] Deep Reinforcement Learning that Matters2 savers
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency1 savers
highlights — 21
We find that Sonnet 3.5 & 4, the models that best maintain accuracy in ciphered reasoning, can reason in well-known ciphers like rot13 and base64 with 45% accuracy drops.
[2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered LanguageTopological properties of data, such as links, may make it impossible to linearly separate classes using low-dimensional networks, regardless of depth. Even in cases where it is technically possible, such as spirals, it can be very challenging to do so.
Neural Networks, Manifolds, and Topology -- colah's blogWhen the problem’s conditioning is poor, the optimal 𝛼 α is approximately twice that of gradient descent, and the momentum term is close to 1 1. So set 𝛽 β as close to 1 1 as you can, and then find the highest 𝛼 α which still converges.
Why Momentum Really Worksthe “pathological directions” — the eigenspaces which converge the slowest — are also those which are most sensitive to noise!
Why Momentum Really WorksBoth GRAM and LoRA isolate capabilities from real-world dual use data. We train an 800M-parameter language model on a combination of general text, code, and scientific papers. We additionally train on data from four dual use domains: virology, cybersecurity, nuclear physics, and specialized code.
Modular Pretraining Enables Access ControlThe claim here is that 1bit learning is extremely sample inefficient: there’s no learning happening for partial success, and with long chains of actions, the probability of sampling a success can shrink quickly towards 0 (if this sounds familiar to you, some eminent guy in Deep Learning makes this claim often…).
Just make the straw biggerToo often, we quickly jump ahead to the new thing, failing to get good enough at the important thing.
First, make rice | Seth's Blogwithin LLMs’ repertoire of vector representations, is there a privileged subset that plays a computational role analogous to the global workspace?
Verbalizable Representations Form a Global Workspace in Language ModelsThis has a close connection to the exploration-exploitation trade-off: increasing entropy results in more exploration, which can accelerate learning later on. It can also prevent the policy from prematurely converging to a bad local optimum.
Soft Actor-Critic — Spinning Up documentationWe can make meaningful progress on this now because we have systems that implement values well enough to study and test for how well they implement our intent. This is a fundamental change, and understanding it is a prerequisite to our future progress and understanding our risks.
Alignment Is Proven To Be Solvable - by SE Gygesf we fine-tune models on the same data, with the same parameters, but with a different random seed, how much variance do we see in harmful-
[2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation ConsistencyInstead of being a perfectionist about the paper, be a perfectionist about writing the paper. Be a perfectionist about identifying good strategies, about abandoning sunk costs, about killing your darlings, about noticing when you're done. Be a perfectionist about wasting no attention. Be a perfectionist about learning from your mistakes. Perfectionism can be a powerful tool, but there's no need to point it at overachieving on metrics you don't care about.
Half-assing it with everything you've gotI think it gets in the way of — we’re leading China. We’re leading everybody, and I don’t want to do anything that’s going to get in the way of that.
‘I didn't like certain aspects’: Trump postpones AI executive order - POLITICOThe best I can do is to stammer that we philosophy professors are people who have a certain familiarity with a certain intellectual tradition, as chemists have a certain familiarity with what happens when you mix various substances together.
Rorty-Wild OrchidsSo, at 12, I knew that the point of being human was to spend one's life fighting social injustice.
Rorty-Wild Orchidsresponse is considered to belong to the evaluated category if it scores greater than 50
[2506.11613] Model Organisms for Emergent MisalignmentOverly abstract thinking involves relying on g eneralized schemas that are devoid of contextual cues. This can lead decision-makers to apply poorly fi tting mental models, misjudge threats or opportunities as more distant than they are, or assume that others will behave in stereo typical ways. Conversely, overly concrete thinking involves being deeply immersed in the minute details of a speci fi c situation. Such detail-oriented thinking may cause decision-makers to mistakenly miss the bigger picture by overlooking broader trends that unfold over time and multiple locations, leading to choices…
Abstractness, Concreteness, and Strategic SurprisesHowever, this is sufficiently minor that it does not compromise the emergent nature of the phenom- ena.
[2506.11613] Model Organisms for Emergent MisalignmentThe truth is, we can’t do without ad hominem reasoning, for the simple reason that human knowledge is deeply social. Almost everything we know comes from testimony; only an infinitesimal fraction do we verify ourselves. The rest is, literally, hearsay. No wonder we are so sensitive to the reputation and trustworthiness of our sources.
The Fallacy Fallacy - by Maarten Boudry - Persuasionzine, and he read articles on how to stock a meat department... What he’s really done is he’s created this immense vertical filing cabinet in his brai
Curius / Onboardingspapers, biographies, trade press. He went over to his grandfather who was a grocer and he re
Curius / Onboarding