Dhruv Sheth
14 followers · 25 following · 1724 views
on the atlas — 131
- Richard_Ngo — LessWrong1 savers
- Nix | Substack1 savers
- stories of your life and others - Google Search1 savers
- Looking for Alice - by Henrik Karlsson - Escaping Flatland5 savers
- Explorative Modeling -- Unlocking a Third Pretraining Axis and End-to-End Generation | Alexi Gladstone2 savers
- AdaWorld4 savers
- Align AI and Mathematics—to Something Else | Bits of DNA1 savers
- Phillip Isola2 savers
- Shrimad Rajchandraji’s Works – Scriptures, Writings & Teachings1 savers
- Huy Ha1 savers
- [2512.18736] Is Your Conditional Diffusion Model Actually Denoising?1 savers
- [2310.07972] Interpretable Diffusion via Information Decomposition1 savers
- What Matters for Latent Actions in Robot Learning1 savers
- [2106.10316] Proper Value Equivalence1 savers
- paper.pdf1 savers
- 6d13e085b79d454da5910e4ca82a3d9d-Paper-Conference.pdf1 savers
- 73471f46f69a13cede6cf2db3a6f438f-Paper-Conference.pdf1 savers
- Can Video World Models Track Unobserved World States?1 savers
- The Rise and Fall of Agent Civilizations15 savers
- Perceptron4 savers
- [2607.15275] RoboTTT: Context Scaling for Robot Policies1 savers
- Scaling is subtler than it seems3 savers
- Behavioral cloning mystery3 savers
- Introducing S1: In-Context Learning for Robotics | Skild AI1 savers
- No, it’s not The Incentives—it’s you – [citation needed]6 savers
- Generalist - GEN-1.5: Embodied Foundation Models are One-Shot Learners7 savers
- How can LLM RL Work Despite Information-Theoretic Inefficiency12 savers
- Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models — DYNA2 savers
- Unlearning Data at Scale2 savers
- We reverse-engineered Flash Attention 44 savers
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction1 savers
- Levanter1 savers
- Introduction to Haliax - Colab1 savers
- marin-community/haliax: Named Tensors for Legible Deep Learning in JAX ·1 savers
- Tensor Considered Harmful2 savers
- Performance Hints11 savers
- blog | Henry Zhu1 savers
- Network and Storage Benchmarks for LLM Training on the Cloud | Henry Zhu1 savers
- World Action Model Atlas1 savers
- RLHF & Post-Training Course by Nathan Lambert14 savers
- Now What? A Recipe for After the Problem Setting (in the Agentic Age) | Tom Silver1 savers
- pdf1 savers
- Moritz Reuss — Robotics & VLA Research1 savers
- RL Post-Training on Macs | Pluralis Research2 savers
- Robostral Navigate: single-camera AI navigation | Mistral AI2 savers
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries3 savers
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog4 savers
- Blog - Aleksa Gordić3 savers
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model1 savers
- [2507.12549] The Serial Scaling Hypothesis1 savers
- Read Something Wonderful - The Extended Internet Universe1 savers
- Sending Samples Without Bits-Back2 savers
- Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis2 savers
- The ultimate guide to RL environments: building and scaling them in the LLM era - a Hugging Face Space by AdithyaSK5 savers
- A guide to the AI tribes - by Michel Justen - What is this2 savers
- Computer Vision and Geometry Group | Robot Learning1 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- State of RL for reasoning LLMs | A. Weers6 savers
- The Unreasonable Effectiveness of LLMs in Mathematics1 savers
- Building a Reasoning Machine - Ph.D. thesis | Edward Hu1 savers
- Lucas Maes1 savers
- stevengongg (@stevengongg) / X1 savers
- Open Source Alternatives to Popular Software1 savers
- Open Source Alternatives to Popular Software1 savers
- Cached Thoughts1 savers
- Occam’s Razor1 savers
- The Hidden Complexity of Wishes1 savers
- What comes next with open models - by Nathan Lambert1 savers
- Everyone is Gambling and No One is Happy - by kyla scanlon2 savers
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reuss4 savers
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić11 savers
- Sporks of AGI - by Sergey Levine - Learning and Control8 savers
- Too Much Information - by Ben Recht - arg min3 savers
- Video Generation Models Explosion 2024 - Yen-Chen Lin2 savers
- The Intelligence Age13 savers
- Theory of Change (Aaron Swartz's Raw Thought)83 savers
- An Opinionated Guide to ML Research52 savers
- Mini Blog Post 3: Become a person who Actually Does Things — Neel Nanda45 savers
- escaping flatland: career advice for CS undergrads43 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- Andrej's advice for success39 savers
- A Mathematical Framework for Transformer Circuits39 savers
- LoRA Without Regret - Thinking Machines Lab37 savers
- How To Scale Your Model34 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Reflections on OpenAI33 savers
- Making Deep Learning Go Faster29 savers
- A Survival Guide to a PhD26 savers
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog24 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- The Last Question24 savers
- How to Land a Frontier Lab Job19 savers
- The best general advice on earth « the jsomers.net blog19 savers
- How To Understand Things - Nabeel S. Qureshi18 savers
- Automated Weak-to-Strong Researcher16 savers
- The Shigalyovist Turn16 savers
- What are Diffusion Models? | Lil'Log16 savers
- Six (and a half) intuitions for KL divergence - LessWrong15 savers
- Alexey Guzey admits to being wrong on sleep14 savers
- Diffusion is spectral autoregression – Sander Dieleman13 savers
highlights — 134
This kind of asynchronous programming is fairly new in the world of CUDA
We reverse-engineered Flash Attention 4That’s because the biggest change in FA4 isn’t the (very cool) math — it’s a massive increase in the complexity of its asynchronous “pipeline” of operations.
We reverse-engineered Flash Attention 4import haliax.nn as hnn
marin-community/haliax: Named Tensors for Legible Deep Learning in JAX ·haliax.nn.softmax(scores, KPos)
marin-community/haliax: Named Tensors for Legible Deep Learning in JAX ·Pos = hax.Axis("position", 1024) # sequence length KPos = Pos.alias("key_position") Head = hax.Axis("head", 8) # number of attention heads Key = hax.Axis("key", 64) # key size Embed = hax.Axis("embed", 512) # embedding size
marin-community/haliax: Named Tensors for Legible Deep Learning in JAX ·3) Error Checking: Can we add annotations to functions giving pre- and post -conditions so that dimensions are automatically checked.
Tensor Considered Harmful2) Interacting with PyTorch Modules: Can we “lift” PyTorch modules with type annotations, so that we know how they change inputs?
Tensor Considered Harmful(Trap 3) Operations across dims are explicit. For instance, the softmax is clearly over the seqlen.
Tensor Considered Harmfulit still falls into many of the traps described above.
Tensor Considered Harmfultensor.transpose("w", "h", "c")
Tensor Considered Harmful6) Private dimensions should be protected.
Tensor Considered HarmfulThe main point though is that this code will run fine for whatever value dim is given. The comment here might descibe what is happening but the code itself doesn’t throw a run time error.
Tensor Considered HarmfulProposal 2: Accessors and Reduction The first benefit of names comes from the ability to replace the need for dim and axis style arguments entirely. For example, lets say we wanted to sort each column.
Tensor Considered HarmfulInfiniBand 400 Gbit/s NIC × 8 cards ~400 GB/s
Network and Storage Benchmarks for LLM Training on the Cloud | Henry ZhuNetwork and storage choices often determine whether training takes hours or days.
Network and Storage Benchmarks for LLM Training on the Cloud | Henry ZhuStill, the result illustrates why video backbones are attractive for robotics: the model has a useful prior for what robot-object interaction should look like
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogWAMs perform best when action learning is co-trained with a video-prediction objective.
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogco-training on video alongside robot data, the broader visual prior can reduce overfitting; the benefit depends on the dataset, objective, and architecture
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogreduce the amount of grounding that has to be learned from robot demonstrations alone
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogPi-0.7’s visual-subgoal results point in the same direction: when the policy is given a desired future image, action prediction becomes more direct and training converges faster [43].
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogPredicting future world changes correlates with generating the necessary actions. Inverse dynamics prediction is often easier than pure action generation [26].
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogThis naturally raises the question: what if we started from a backbone that already represents how language maps to visual change in the world?
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogKnowledge Insulation reports similar findings and makes the concern architectural: it isolates the gradients of the flow-matching action expert from the VLM backbone to preserve pretrained language/vision knowledge, improving training convergence, task performance, and language following [20]. Recent solutions like VLM co-training and discrete action tokenizers have helped, but the core challenge remains: grounding language into physical action from limited robot data.
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogVLM2VLA frames this directly as catastrophic forgetting during the VLM-to-VLA transition [55].
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogThe choice of backbone impacts the full training and evaluation pipeline, from training recipe and data mixture to inference optimizations. Given the cost of running these models at scale, most teams will likely have to prioritize one direction (VLA or WAM) first rather than fully pursuing both in parallel.
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogMany teams are building on the traditional VLA recipe established by Pi-0 [2] and later refined by Pi-0.5 [4], using VLM backbones as the starting point for policy learning. This VLM-backbone recipe appears in public work from teams including NVIDIA GR00T [5], Xiaomi Robotics [27], Being-H0.5 [28], and others.
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical BlogTogether, these issues point to a single underlying problem: many latent-action objectives remain im- plicitly anchored to pixel variation, and thus learn dy- namics that are predictive but not necessarily action- centric
[2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelhe resulting “action” becomes se- mantically empty—useful for matching training losses, but not a meaningful factor for control.
[2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelBesides textbooks, I recommend reading PhD theses of researchers whose work interests you. PhD theses in ML usually are ordered as follows: (1) introductory and background material, (2) several papers that were previously published at conferences (it’s said that you just have to “staple together” your papers to write your thesis), and (3) a conclusion and outlook. You’re likely to benefit most from parts (1) and (3), since they contain a unifying view of the past and future of the field, written by an expert.
An Opinionated Guide to ML ResearchWill this be a 10% improvement or a 10X improvement?
An Opinionated Guide to ML ResearchSometimes, people who are both exceptionally smart and hard-working fail to do great research. In my view, the main reason for this failure is that they work on unimportant problems.
An Opinionated Guide to ML ResearchIn modern civilization particularly, no one can think fast enough to think their own thoughts.
Cached ThoughtsThing is to call this extra effort “taste”, but also see “quality” and “craftsmanship” for attempts to define the same.
personal principles for the new world - by Vedant NairYou want to position yourself near opportunities. You don’t have to be that perfect. You want to position yourself near the tree. Even if you don’t catch the apple before it hits the ground, so long as you’re the first one to pick it up. You want to position yourself close to the opportunities. That’s kind of a lot of my work, is positioning the company near opportunities, and the company having the skills to monetize each one of the steps along the way so that we can be sustainable. — Jensen Huang, CEO and co-founder of NVIDIA (on the Acquired Podcast)
personal principles for the new world - by Vedant NairHowever, this is not a well-defined task. It does not specify a strictly falsifiable input-output behavior: The inputs are images or video, sure—but what are the outputs?
The flavor of the bitter lesson for computer vision - Vincent Sitzmannultimate uncaused cause. (See Four causes.)
Why is there anything at all?Some have suggested the possibility of an infinite regress, where, if an entity cannot come from nothing and this concept is mutually exclusive from something, there must have always been something that caused the previous effect, with this causal chain (either deterministic or probabilistic) extending infinitely back in time.[14][15][16]
Why is there anything at all?process reward models and teacher-forcing on the reasoning sequences might make a comeback.
As Rocks May Think | Eric JangIn early 2024, Yao et al. combined the deductive inference of tree search to try to boost reasoning capabilities by giving an explicit way for LLMs to parallelize and backtrack on reasoning steps, much like how the AlphaGo game tree works. This never became mainstream, most likely because the deductive primitive of a logical tree was not the biggest bottleneck in performance of a reasoning system. Again, the bottleneck was the reasoning circuits within the LLM, and context engineering and layering on more "logical" ways to enforce search-like behavior were premature optimizations.
As Rocks May Think | Eric JangThere are two broad categories of it: deductive inference and inductive inference.
As Rocks May Think | Eric Jangcultivate the conditions for good
What Is It Like To Be A Writer? - by Jasmine SunBlogging is not only a search query but a kind of prompt for human wisdom: the world will respond with the seriousness that you put into it.
What Is It Like To Be A Writer? - by Jasmine Sunrationalist’s precision and a critic’s glamour.
What Is It Like To Be A Writer? - by Jasmine Sunomplete picture of who they are and what they believe
What Is It Like To Be A Writer? - by Jasmine SunThe best interviews feel vaguely spiritual
What Is It Like To Be A Writer? - by Jasmine SunThis sort of spiritual stuff is core to my interest in real thinking. But I also worry that it sounds a bit … pious. As though: “ah, yes, real thinking; real thinking is very spiritually superior; it is very virtuous; thou shalt do real thinking.” The way, perhaps, one “should” meditate, but actually it’s boring and dry and you don’t want to. Indeed: because real thinking is more work, it’s at risk of seeming like a chore, especially if the main motivation is some vague, spiritual “supposed to.”
Fake thinking and real thinking - Joe Carlsmith's SubstackBut I don’t think the resonance is accidental.
Fake thinking and real thinking - Joe Carlsmith's SubstackIndeed, I think curiosity is maybe the most paradigm “real thinking” vibe.
Fake thinking and real thinking - Joe Carlsmith's SubstackAnd relatedly: real thinking feels somehow “right.” As though (as with sincerity), something falls into its proper place. Some deep structure coheres and mobilizes and starts to work as it should. There is a sense of grip.
Fake thinking and real thinking - Joe Carlsmith's SubstackThis is why Jesus speaks in parables throughout the New Testament — in ways that stick with you long after you’ve read them — rather than just stating the abstract principle.
How To Understand Things - Nabeel S. Qureshi