Winnie Xu
77 followers · 19 following · 2921 views
on the atlas — 117
- RLHF & Post-Training Course by Nathan Lambert14 savers
- Anthropic v DoW - by Jordan Schneider - ChinaTalk1 savers
- AGI Trades11 savers
- Why We Think | Lil'Log25 savers
- What 81,000 people want from AI \ Anthropic26 savers
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / X7 savers
- pdf6 savers
- To Get Good, Go After The Metagame - Commoncog1 savers
- In-context Learning and Induction Heads13 savers
- Measuring the performance of our models on real-world tasks | OpenAI5 savers
- An Alchemist’s Notes on Deep Learning — An Alchemist's Notes on Deep Learning4 savers
- Learn to Read Korean in 15 Minutes1 savers
- jonvonkowallis.com/readers/CHIN5910/178-206-Eliot_Weinberger_%26_Octavia_Paz-Nineteen_Ways_of_Looking_at_Wang_Wei.pdf1 savers
- Debugging Reinforcement Learning Systems11 savers
- The Annotated S411 savers
- Scam Cities—Asterisk10 savers
- LoRA Without Regret - Thinking Machines Lab37 savers
- Introduction - SITUATIONAL AWARENESS: The Decade Ahead50 savers
- Tutorial - What is a variational autoencoder? – Jaan Altosaar2 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- jasmine sun's visit to china12 savers
- The Anthropic Economic Index \ Anthropic10 savers
- Dario Amodei — On DeepSeek and Export Controls2 savers
- N-dimensional Rotary Positional Embeddings1 savers
- The Only Important Technology Is The Internet — Kevin Lu17 savers
- dennyzhou.github.io/LLM-Reasoning-Stanford-CS-25.pdf1 savers
- Exploring institutions for global AI governance - Google DeepMind1 savers
- Toy Models of Superposition30 savers
- kevin frans blog1 savers
- Dario Amodei — The Urgency of Interpretability17 savers
- On the Biology of a Large Language Model32 savers
- Generative AI’s Act Two | Sequoia Capital5 savers
- Tracing the thoughts of a large language model \ Anthropic18 savers
- machine learning i - by vincent huang - a slice of my mind4 savers
- Claude's extended thinking \ Anthropic1 savers
- Conceptual SpacesThe Geometry of Thought | Books Gateway | MIT Press1 savers
- AGI and the EMH: markets are not expecting aligned or unaligned AI in the next 30 years — Basil Halperin9 savers
- o3-mini is really good at writing internal documentation2 savers
- Tips for Managing GSUs (Google RSUs) — EquityFTW1 savers
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs2 savers
- What is Technology? - Letters to a Young Technologist18 savers
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon University3 savers
- Transformer²: Self-Adaptive LLMs1 savers
- Home | Santa Fe Institute7 savers
- Optimizing AI Inference at Character.AI9 savers
- Just know stuff. (Or, how to achieve success in a machine learning PhD.) · Patrick Kidger14 savers
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlow2 savers
- 2403.09611.pdf3 savers
- Introductions - jxnl.co3 savers
- Planning for AGI and beyond9 savers
- How Lossless Data Compression Works | Quanta Magazine3 savers
- An Interactive Introduction to Fourier Transforms6 savers
- To Understand Language is to Understand Generalization | Eric Jang6 savers
- More Is Different for AI9 savers
- Competitive programming with AlphaCode | DeepMind2 savers
- Analogies between Biology and Deep Learning [rough note] -- colah's blog3 savers
- How To Be Successful96 savers
- The Bitter Lesson78 savers
- An Opinionated Guide to ML Research52 savers
- Quantum computing for the very curious51 savers
- Mimetic - Brian Timar44 savers
- A Mathematical Framework for Transformer Circuits39 savers
- Home | Levers For Progress37 savers
- HOWTO: Be more productive (Aaron Swartz's Raw Thought)34 savers
- Geeks, MOPs, and sociopaths in subculture evolution31 savers
- Research Debt28 savers
- Google "We Have No Moat, And Neither Does OpenAI"28 savers
- We Need a New Science of Progress - The Atlantic25 savers
- Visual Information Theory -- colah's blog24 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning23 savers
- Letters to a Young Technologist22 savers
- ChatGPT Is a Blurry JPEG of the Web | The New Yorker22 savers
- Zoom In: An Introduction to Circuits22 savers
- View article22 savers
- By default, capital will matter more than ever after AGI — LessWrong19 savers
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet17 savers
- GPT-417 savers
- Neural Networks, Manifolds, and Topology -- colah's blog16 savers
- Why does DARPA work?16 savers
- A Recipe for Training Neural Networks15 savers
- LLM Powered Autonomous Agents | Lil'Log14 savers
- the agony of eros: dating - by Ava - bookbear express14 savers
- SF Privately Owned Public Open Spaces12 savers
- The Mom Test: how to talk to customers and learn if your business is a good idea when everybody is lying to you12 savers
- Yamauchi No.10 Family Office11 savers
- Thinking like Transformers10 savers
- Just Ask for Generalization | Eric Jang10 savers
- Reward Hacking in Reinforcement Learning | Lil'Log10 savers
- Blueprint for an AI Bill of Rights - The White House10 savers
- Explorable Explanations9 savers
- Kullback–Leibler divergence9 savers
- Too much efficiency makes everything worse: overfitting and the strong version of Goodhart’s law | Jascha’s blog8 savers
- The Indy8 savers
- Red Blob Games8 savers
- An Intuitive Explanation of Solomonoff Induction - LessWrong8 savers
- The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!) – Joel on Software7 savers
- The "Basics" | Putting the "You" in CPU7 savers
- How to Use t-SNE Effectively6 savers
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.com6 savers
- LLM Inference Performance Engineering: Best Practices | Databricks Blog6 savers
highlights — 519
Like simpler games, real-world metas come in roughly two flavours: ones that are defined by external changes to the rules of a game, and ones that are shaped by a dynamic equilibrium of competition within a stable system of play. Unlike games, however, real-world domains have no set rules: they are vastly more complicated and interesting, because the rules change only when someone notices the rules have changed.
To Get Good, Go After The Metagame - Commoncogif we’d like to avoid nondeterminism in our inference servers, we must achieve batch invariance in our kernels
Defeating Nondeterminism in LLM Inference - Thinking Machines Labkernels don’t have batch invariance
Defeating Nondeterminism in LLM Inference - Thinking Machines Labload determines the batch size that the kernels are run under
Defeating Nondeterminism in LLM Inference - Thinking Machines Labfloating-point non-associativity and concurrent execution leads to nondeterminism based on which concurrent core finishes first
Defeating Nondeterminism in LLM Inference - Thinking Machines Labsoftware modification, code debugging, and network troubleshooting
The Anthropic Economic Index \ Anthropice opacity makes it hard to find definitive evidence supporting the existence of these risks at a large scale, making it hard to rally support for addressing them—and indeed, har
Dario Amodei — The Urgency of Interpretabilitycross-layer transcoder (CLT)
On the Biology of a Large Language ModelAttribution graphs generate hypotheses about the mechanisms used by the model, which we test and refine through follow-up perturbation experiments
On the Biology of a Large Language Modelthis setting does not evaluate this policy in terms of its zero-shot performance on the test task, but lets it adapt to the test task by executing a few “training” episodes at test-time, after executing which the policy is evaluated
Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityinference compute-constrained class of test-time algorithms A C 𝐴 𝐶
Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universitythe state of the art approach to exploring a nice and wide space of models and hyperparameters is to use an intern
A Recipe for Training Neural Networkswith a well-established system of processes, the necessity of PhD degrees disappears rapidly.
i sensed anxiety and frustration at NeurIPS’24 – Kyunghyun Cholie to humans systematically to persuade them)
Self-exfiltration is a key dangerous capabilityOnce models have the ability to self-exfiltrate, it doesn’t mean that they would choose to. But this then becomes a question about their alignment: you need to ensure that these models don’t want to self-exfiltrate.
Self-exfiltration is a key dangerous capabilityould the model “steal” its own weights and copy it to some external server that the model owner doesn’t contro
Self-exfiltration is a key dangerous capabilitycausally responsible for the conclusion that the model came to.
Externalized reasoning oversight: a research direction for language model alignment — AI Alignment ForumReasoning oversight should provide stronger guarantees of alignment than oversight on model outputs alone, since we would get insight into the causally responsible reasoning process that gave rise to a certain output.
Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forummanipulates a generative model's token generation process to constrain its next-token predictions to only tokens that do not violate the required output structure.
A Guide to Structured Outputs Using Constrained Decodingproxy for the reward signal we want and fine tune the LM on that proxy
Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaoricher, more disentangled models inside of them, which would be useful for alignment
Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaomodelling the underlying processes that cause the synthetic data is significantly easier than directly modelling the distribution of text.
Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaoif we learn the policy correctly, we can approximate the best-of-n distribution with a single forward pass
Spending Inference Time — Kevin Lustop is not inclusive
numpy.mgrid — NumPy v2.0 Manualdimensions and number of the output arrays are equal to the number of indexing dimensions
numpy.mgrid — NumPy v2.0 Manualnatively train our models in int8 precision
Optimizing AI Inference at Character.AI1 out of every 6 layers uses global attention
Optimizing AI Inference at Character.AIMulti-Query Attention
Optimizing AI Inference at Character.AImultiple choice question test
What's going on with the Open LLM Leaderboard?activate densely, meaning each activation is always firing on each input
Extracting Concepts from GPT-4 | OpenAIIn the dense regime, we end up with each neuron representing a single feature, and we can read feature values directly off of neuron activations.
Toy Models of Superpositionone-layer transformer with a 512-neuron MLP layer, and decompose the MLP activations into relatively interpretable features by training sparse autoencoders on MLP activations from 8 billion data points, with expansion factors ranging from 1× (512 features) to 256×
Towards Monosemanticity: Decomposing Language Models With Dictionary Learningsuperposition can arise naturally during the course of neural network training if the set of features useful to a model are sparse in the training data
Towards Monosemanticity: Decomposing Language Models With Dictionary Learningfrequency of concepts and the dictionary size
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnetabstract and concrete instantiations
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnetmultilingual (responding to the same concept across languages
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnetextracting high-quality features
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 SonnetFor instance, we see that clamping the Golden Gate Bridge feature 34M/31164353 to 10× its maximum activation value induces thematically-related model behavior
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 SonnetThe inefficiency of over-training a smaller model easily gets amortized over the inference lifetime of the model (since the model is smaller, each request incurs much less compute)
Factorial Funds | Thoughts on Llama 3treating reward as percentile
Spending Inference Time — Kevin LuInverse Cloze Task (predicting a sentence's surrounding context).
Retrieval Augmented Generation Research: 2017-2024excluding those with overly long instructions
openbmb/UltraFeedback · Datasets at Hugging Face66k prompts with pairs of chosen and rejected responses
Preference Tuning LLMs with Direct Preference Optimization MethodsOpenHermes-2.5-Mistral-7B and Zephyr-7b-beta-sft
Preference Tuning LLMs with Direct Preference Optimization Methodstwo alignment datasets Intel’s orca_dpo_pairs and the ultrafeedback-binarized dataset.
Preference Tuning LLMs with Direct Preference Optimization Methodswhile the denominator pushes embeddings every negative pair � � , � � Q i ,D j apart from each other (from � ≠ � j =i).
Long-Context Retrieval Models with Monarch Mixer · Hazy ResearchMultipleNegativesRankingLoss
Long-Context Retrieval Models with Monarch Mixer · Hazy Researchsimply extending the standard BERT training pipeline to longer sequence data was insufficient to train a good long-context retrieval model.
Long-Context Retrieval Models with Monarch Mixer · Hazy ResearchEnglish, French, German, Spanish, and Italian
Welcome Mixtral - a SOTA Mixture of Experts on Hugging Facewe don't need to teach them new tasks from scratch, we just need to elicit their latent knowledge
Weak-to-strong generalization