✳flâneur — a map of the web's best reading
4764
Open this reading profile →on the atlas — 23
- [1912.02178] Fantastic Generalization Measures and Where to Find Them1 savers
- Against Interpretability: a Critical Examination of the Interpretability Problem in Machine Learning | Philosophy & Technology | Springer Nature Link1 savers
- The Banal Evil of AI Safety - by Ben Recht - arg min1 savers
- [2502.21212] Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought1 savers
- [2502.04327] Value-Based Deep RL Scales Predictably1 savers
- KL-divergence as an objective function — Graduate Descent1 savers
- Reformist Reinforcement Learning - by Ben Recht - arg min1 savers
- The Mythos of Model Interpretability: In machine learning, the concept of interpretability is both important and slippery.: Queue: Vol 16, No 31 savers
- Sparse Autoencoders Find Highly Interpretable Features in Language Models2 savers
- PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.1 savers
- pigeon1 savers
- Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt - ACL Anthology1 savers
- On Calibration of Modern Neural Networks1 savers
- On Getting Confidence Estimates from Neural Networks | Bharath's notes1 savers
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 20241 savers
- Designing and Interpreting Probes · John Hewitt1 savers
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity1 savers
- Understanding Intermediate Layers Using Linear Classifier Probes1 savers
- Confidence Regulation Neurons in Language Models1 savers
- Silicon Valley’s Gold Rush Roots—Asterisk1 savers
- nelsn.jpg1 savers
- When RAND Made Magic in Santa Monica—Asterisk4 savers
- Rethinking High-School Science Fairs—Asterisk8 savers