✳flâneur — a map of the web's best reading
Christina Lee
3 followers · 4 following · 189 views
Open this reading profile →
on the atlas — 32
- Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org1 savers
- Gabe Pereyra on X: "Model strategy for @harvey: We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture." / X1 savers
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog2 savers
- [2507.06187] The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains1 savers
- [2606.10346] Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning1 savers
- Better MoE model inference with warp decode · Cursor1 savers
- Nemotron 3 Ultra: what distillation can't fix1 savers
- [2606.04033] Inverse Critical Experiment Design via Gradient Optimization and a Multigroup Attention-Based Neural Network Architecture1 savers
- Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / X1 savers
- Part 1: The Map Was Wrong — Nemostation1 savers
- Is Frontier Asynchronous RL Solved? — Luke J. Huang4 savers
- openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering ·1 savers
- [2605.15156] MeMo: Memory as a Model1 savers
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretraining1 savers
- 1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf1 savers
- Qwen1 savers
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens | OpenReview1 savers
- Ihtesham Ali on X: "A Norwegian neuroscientist spent 20 years proving that the act of writing by hand changes the human brain in ways typing physically cannot, and almost nobody outside her field has read the paper. Her name is Audrey van der Meer. She runs a brain research lab in Trondheim, and https://t.co/9FDGAKPA6I" / X1 savers
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / X1 savers
- 221 Cool and Unusual Things to Do in San Francisco - Atlas Obscura1 savers
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraints1 savers
- Current and New Activation Checkpointing Techniques in PyTorch – PyTorch1 savers
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-study2 savers
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernels1 savers
- [2603.08716] Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU1 savers
- When AI builds itself \ Anthropic30 savers
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Lab18 savers
- Pre, Mid, Post-Training Way of Life - by Tina He5 savers
- GLM-5.2: Built for Long-Horizon Tasks4 savers
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziems4 savers
- Learning Beyond Gradients3 savers
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation2 savers