Lucas Chen
5 followers · 5 following · 223 views
on the atlas — 24
- Four LLM loss functions → four flavors of LLM misalignment — LessWrong3 savers
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- nn-notes.pdf4 savers
- An OpenAI model left notes about how to evade containment; we need more details — LessWrong1 savers
- [0906.4325] Finite element exterior calculus: from Hodge theory to numerical stability1 savers
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong6 savers
- google/hackable_diffusion ·1 savers
- A Visual Guide to DiffusionGemma - by Maarten Grootendorst1 savers
- google/diffusiongemma-26B-A4B-it · Hugging Face1 savers
- A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWrong2 savers
- ProblemsILike.com — Daniel Litt1 savers
- Natural Language Autoencoders \ Anthropic9 savers
- Notes from the front - by Michael Harris - Silicon Reckoner1 savers
- Maybe I was too harsh on deep learning theory (three days ago) — LessWrong5 savers
- A Technical Introduction to Solomonoff Induction without K-Complexity — LessWrong1 savers
- Deep learning as program synthesis — LessWrong1 savers
- [2604.21691] There Will Be a Scientific Theory of Deep Learning5 savers
- [2503.02113] Deep Learning is Not So Mysterious or Different1 savers
- Introducing talkie: a 13B vintage language model from 193020 savers
- [2410.12101] The Persian Rug: solving toy models of superposition using large-scale symmetries1 savers
- pdf1 savers
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs5 savers
- The paper that killed deep learning theory — LessWrong1 savers
- Lucas's Bookshelf / Curius1 savers
highlights — 3
NLAs helped Anthropic researchers discover training data that caused this.
Natural Language Autoencoders \ AnthropicWhen this round is over, they will no longer have any need for mathematics (within 5 years, according to several people who were participating in this morning’s conversation)
Notes from the front - by Michael Harris - Silicon ReckonerMaybe I was too harsh on deep learning theory (three days ago)
Maybe I was too harsh on deep learning theory (three days ago) — LessWrong