Christina Lee
8 followers · 6 following · 353 views
on the atlas — 54
- How to Build a $20 Billion Semiconductor Fab15 savers
- Callosum1 savers
- Journey to 2-second Inter-node RL Weight Transfer1 savers
- GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress1 savers
- OpenFASOC: Open Source Fully-Autonomous SoC Synthesis using Customizable Cell-Based Synthesizable Analog Circuits [CHIPS Alliance]1 savers
- Making Claude a better electrical engineer | Claude by Anthropic1 savers
- CircuitNet/assets/overall_structure.png at main · circuitnet/CircuitNet1 savers
- Solvaix on X: "$31 MILLION HOTEL. EVERY PIPE AND WIRE MAPPED IN ADVANCE. BUILT BY TYPING SENTENCES INSTEAD OF PLACING THEM BY HAND. A 190-room hotel needed its ductwork, plumbing, electrical, and fire suppression modeled and coordinated before construction could start. The design team https://t.co/jqOAkMf5vX" / X1 savers
- Circuit Tracing in Vision–Language Models: Understanding the Internal Mechanisms of Multimodal Thinking1 savers
- Marion Lepert on X: "Catching skin cancer early is a home robotics problem. Melanoma is highly treatable when detected early, yet today’s screening process depends heavily on patients noticing tiny changes across their entire skin surface. This requires patients to solve a near-impossible https://t.co/te3SEpyi0X" / X1 savers
- Vinith M Suriyakumar on X: "Can we determine whether a model has been trained to generate harmful content without ever producing a single image? In our new paper, which won an Outstanding Paper Award at the AI4GOOD workshop at ICML 2026, we answer yes. We introduce Evaluation without Generation, a new https://t.co/8T2U3VKPQd" / X1 savers
- zeroasiccorp/logikbench: Digital logic benchmark suite ·1 savers
- Harveen Singh Chadha on X: "forget models training models - kimi k3 spent one 48 hour autonomous run designing a chip to serve a nano model built on its own architecture - it also wrote miniTriton from scratch matching or beating Triton on roofline benchmarks its so over https://t.co/tHkb62AbG3" / X1 savers
- Brain and Brawn: 1st Workshop on Robot Hardware-Aware Intelligence1 savers
- Dwarkesh Patel on X: "New blackboard lecture w @reinerpope How do chips actually work – starting with basic logic gates, and working up to why GPUs, TPUs, FPGAs, and the human brain each look the way they do. 0:00:00 – Building a multiply-accumulate from logic gates 0:16:20 – Muxes and the cost of https://t.co/OhZlNxdYyz" / X1 savers
- Jim Fan on X: "Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an https://t.co/57JC05VN54" / X2 savers
- Scaling Laws, Carefully | Lil'Log15 savers
- feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rl1 savers
- RL at 1T Scale: prime-rl Performance Deep Dive1 savers
- The Last Question24 savers
- cyberpunk edgerunner - Google Search1 savers
- Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org1 savers
- Gabe Pereyra on X: "Model strategy for @harvey: We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture." / X1 savers
- GLM-5.2: Built for Long-Horizon Tasks4 savers
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog4 savers
- [2507.06187] The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains1 savers
- [2606.10346] Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning1 savers
- Better MoE model inference with warp decode · Cursor3 savers
- Nemotron 3 Ultra: what distillation can't fix1 savers
- When AI builds itself \ Anthropic34 savers
- [2606.04033] Inverse Critical Experiment Design via Gradient Optimization and a Multigroup Attention-Based Neural Network Architecture1 savers
- Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / X1 savers
- Part 1: The Map Was Wrong — Nemostation1 savers
- Is Frontier Asynchronous RL Solved? — Luke J. Huang6 savers
- openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering ·1 savers
- [2605.15156] MeMo: Memory as a Model1 savers
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretraining1 savers
- 1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf1 savers
- Qwen1 savers
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens | OpenReview1 savers
- Ihtesham Ali on X: "A Norwegian neuroscientist spent 20 years proving that the act of writing by hand changes the human brain in ways typing physically cannot, and almost nobody outside her field has read the paper. Her name is Audrey van der Meer. She runs a brain research lab in Trondheim, and https://t.co/9FDGAKPA6I" / X1 savers
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / X1 savers
- 221 Cool and Unusual Things to Do in San Francisco - Atlas Obscura1 savers
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraints2 savers
- Current and New Activation Checkpointing Techniques in PyTorch – PyTorch1 savers
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-study2 savers
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernels1 savers
- [2603.08716] Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU1 savers
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Lab23 savers
- bitter_lesson.pdf6 savers
- Pre, Mid, Post-Training Way of Life - by Tina He5 savers
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziems4 savers
- Learning Beyond Gradients3 savers
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation2 savers
highlights — 39
LogikBench is a large hybrid benchmark suite of human authored and AI generated Verilog RTL circuits. Use cases includes objective evaluation of EDA algorithms/tools, PDKs, architectures, model/LLMs, compute infrastructure, (and more...).
zeroasiccorp/logikbench: Digital logic benchmark suite ·Co-Design of Control and Hardware: Joint optimization of physical design and control policies for high-performance robotic systems.
Brain and Brawn: 1st Workshop on Robot Hardware-Aware Intelligencewe give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library.
Jim Fan on X: "Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an https://t.co/57JC05VN54" / XThe constraints were simply: do not train a neural network, make it locally reproducible, leave records for each round, and keep pushing the score up.
Learning Beyond GradientsReconciling Kaplan and Chinchilla
Scaling Laws, Carefully | Lil'LogChinchilla Scaling Laws
Scaling Laws, Carefully | Lil'LogEach algorithm is a named runtime class — the algorithm object is the algorithm. Dispatch is keyed on algo.type — it names the algorithm, and each config class's defaults are its vetted parameterization:
feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rlThe algorithm classes
feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rlon May 14, 2061, what had been theory, became fact.
The Last QuestionGLM-5.1 training can be run with a single command on a Slurm cluster
RL at 1T Scale: prime-rl Performance Deep DiveFor decades, Multivac had helped design the ships and plot the trajectories that enabled man to reach the Moon, Mars, and Venus, but past that, Earth's poor resources could not support the ships.
The Last QuestionBy utilizing a source-side CPU engine replica and P2P RDMA transfers via Mooncake TransferEngine, we speed up weight transfer times for 1T-parameter Kimi-K2 7 times (53 seconds -> 7.2 seconds), at the cost of one additional inference engine replica (32G) per training rank on CPU memory.
Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS OrgBy utilizing a source-side CPU engine replica and P2P RDMA transfers via Mooncake TransferEngine, we speed up weight transfer times for 1T-parameter Kimi-K2 7 times (53 seconds -> 7.2 seconds
Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS OrgModel strategy for @harvey : We are working on the first model in our legal foundation model series, inspired by @cursor_ai 's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture. 2. Create the foundations for law firms to build their own specialized models and own their own intelligence.
Gabe Pereyra on X: "Model strategy for @harvey: We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture." / XImproved Architecture: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP layer for speculative decoding, increasing the acceptance length by up to 20%
GLM-5.2: Built for Long-Horizon TasksAdvanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency
GLM-5.2: Built for Long-Horizon TasksThis paper ex- amines the extent to which post-training can enhance LLMs’ capacity for causal inference. We introduce CauGym , a comprehensive dataset comprising seven core causal tasks for training and five diverse test sets. Using this dataset, we systematically evaluate five post-training ap- proaches: SFT, DPO, KTO, PPO, and GRPO.
[2602.06337] Can Post-Training Transform LLMs into Causal Reasoners?We prove the failure is fundamental: supervised fine-tuning, direct preference optimization, and in-context learning all produce predictors that cannot distinguish between causal graphs generating similar observational data, and any attempt to do so requires the model’s internal representations to grow unboundedly, violating the very conditions under which these methods work
[2605.27567] Why LLMs Fail at Causal Discovery and How Interventional Agents EscapeThe contrast with DeepSeek-V4 is the interesting part. DeepSeek dropped its RL stage and replaced it with multi-teacher distillation. Nvidia keeps RL and stacks MOPD on top. DeepSeek treats distillation as a replacement for RL, Nvidia treats it as a consolidation layer that RL feeds into.
Nemotron 3 Ultra: what distillation can't fixwhy using the old LoRA config not the 3D shape from PEFT 18?
Adina Yakup on X: "Macaron-V1-Preview-749B 👀 a Mixture-of-LoRA personal agent model from MindLab ✨ 744B base + 5 specialist LoRAs ✨ Generative UI as a core skill ✨ Personal agent focused ✨ 202K context ✨ MIT license https://t.co/OkjxKEThzZ" / Xmake visual training data controllable
Yucheng Shi on X: "Static datasets give fixed samples. Code-generated environments give scalable worlds. For LLM reasoning, executable environments can generate fresh problems and verifiable rewards. With TRON, we bring this idea to visual reasoning: each code-generated environment samples a https://t.co/H8C6eD3nR4" / XWith TRON, we bring this idea to visual reasoning: each code-generated environment samples a latent visual state, renders an image, asks a question, and verifies the answer from the underlying state.
Yucheng Shi on X: "Static datasets give fixed samples. Code-generated environments give scalable worlds. For LLM reasoning, executable environments can generate fresh problems and verifiable rewards. With TRON, we bring this idea to visual reasoning: each code-generated environment samples a https://t.co/H8C6eD3nR4" / Xlarge-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code.
Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / Xenvironment design ("the most powerful RL environment is the product itself"),
Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / XPractical consequences for this project: Decode is markedly faster than an all-full-attention 2B baseline: each GDN layer is O(n) instead of O(n²), which compounds during GRPO rollouts. KV-cache footprint is meaningfully smaller per sequence, which lets us batch more rollout sequences in parallel. Token counts are heavy. Verified from our actual training set (226K records across 5 source files): median ~2,100 tokens per training example; ActivityNet- and COIN-style longer clips sit at ~5,000 tokens typical and ~7,800 at p90; the cap before we drop a clip is around ~8,500 tokens. Attention savi…
Part 1: The Map Was Wrong — NemostationSequence-level importance sampling is the estimator that scales with compute, while token-level estimators become structurally inconsistent at high policy lag.
Is Frontier Asynchronous RL Solved? — Luke J. HuangWe find that the teacher need not be strong: with proper mixing of the language modeling and knowledge distillation losses, even small and undertrained teachers improve larger students.
[2605.23857] Strong Teacher Not Needed? On Distillation in LLM PretrainingReflexion agents verbally reflect on task feedback signals, then maintain their own reflective text in an episodic memory buffer to induce better decision-making in subsequent trials
1b44b878bb782e6954cd888628510e90-Paper-Conference.pdfur pre-training dataset includes data sourced across many different domains, including web documents and code, and incorporates image, audio, and video content. For the instruction- tuning phase we finetuned Gemini 1.5 models on a collection of multimodal data (containing paired instructions and appropriate responses), with further tuning based on human preference data.
gemini_v1_5_report.pdfLLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels.
Han Guo on X: "LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs). https://t.co/cOTeMUr4py" / XThe regions responsible for memory, sensory integration, and the encoding of new information were all firing together in a coordinated pattern that spread across the entire cortex. The whole network was awake and connected.
Ihtesham Ali on X: "A Norwegian neuroscientist spent 20 years proving that the act of writing by hand changes the human brain in ways typing physically cannot, and almost nobody outside her field has read the paper. Her name is Audrey van der Meer. She runs a brain research lab in Trondheim, and https://t.co/9FDGAKPA6I" / X221 Cool, Hidden, and Unusual Things to Do in San Francisco, California
221 Cool and Unusual Things to Do in San Francisco - Atlas Obscuratorch._dynamo.config.activation_memory_budget = 0.5 out = torch.compile(fn)(inp)
Current and New Activation Checkpointing Techniques in PyTorch – PyTorch(compile-only) Memory Budget API [NEW!]
Current and New Activation Checkpointing Techniques in PyTorch – PyTorchMany heuristics were not useless; they were simply too expensive to maintain. Coding agents change that maintenance curve. Rules that used to be one-off patches may start to become code worth owning for the long term.
Learning Beyond Gradientsnew two-part methodological framework combining 1) lightweight, distributed feedback loops on individual data sources, with 2) centralized integration tests to assess candidate mixes on base model quality and post-trainability
[2512.13961] Olmo 3Tiling window manager for macOS along the lines of xmonad.
Amethyst | ianyhResolves dependencies between columns (DAG-based execution)
Architecture & Performance - NeMo Data DesignerAsync RL. Early in RL, a meaningful fraction of trajectories enter loops that are guaranteed to score zero. The model clicks the same dead UI element repeatedly, or opens a fresh tab and loses the canvas. Our setup is not fully async, so the slowest trajectory in a batch gates the whole step. The next thing to try would be async RL a la Pipeline RL. Denser reward signals. Per-step or per-tool-call rewards would credit intermediate progress more densely than the single end-of-trajectory rubric we use today. This should help most on long-horizon composition. Feedback descent.Our reward function …
Training LLMs to use Photoshop | Mickey Labs