Tasha Pais
16 followers · 21 following · 773 views
on the atlas — 82
- All About Transformer Inference | How To Scale Your Model2 savers
- GO EAST — Writing Fellowship Application3 savers
- Introduction - SITUATIONAL AWARENESS: The Decade Ahead50 savers
- Wiring Capital to Compute1 savers
- MatX AI Chip Startup Raises $500M to Challenge Nvidia | 20261 savers
- Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI | NVIDIA Technical Blog1 savers
- adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up ·4 savers
- All About Rooflines | How To Scale Your Model10 savers
- Reintroducing Sieve2 savers
- The future of AI is already written | Mechanize Inc.8 savers
- Surgeons Should Not Look Like Surgeons | by Nassim Nicholas Taleb | INCERTO | Medium1 savers
- life updates, 2026 q0.5 - by Hardeep Gambhir3 savers
- vivek on X: "how to be good at research" / X5 savers
- The Old World Is Dying: Advice for 2026 graduates42 savers
- Sulaiman Khan Ghori | Thiel Fellowship1 savers
- The month that triathlon took over my life | by Grace Gerwe | Medium1 savers
- Khan Space Industries1 savers
- Google’s Genie 3 Is What Science Fiction Looks Like1 savers
- Genie 3: A new frontier for world models — Google DeepMind1 savers
- How To Scale Your Model34 savers
- Noam Brown | Innovators Under 351 savers
- Style-Aware Generative Models · Gwern.net1 savers
- Learning the integral of a diffusion model – Sander Dieleman4 savers
- How I Got a Job at Google DeepMind (No ML Degree) | Medium3 savers
- adam-maj/robotics: A deep dive on the history of robotics and the future of humanoids ·4 savers
- How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and Lessons1 savers
- How to Land a Frontier Lab Job19 savers
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation1 savers
- Nikolaus West on X: "The data layer tax for robot learning" / X1 savers
- Careers | Fulcrum1 savers
- Keller Jordan on X: "Modded-NanoGPT Optimization Benchmark Hundreds of neural network optimizers have been proposed in the literature, recently including dozens citing Muon: MARS, SWAN, REG, ADANA, Newton-Muon, TrasMuon, AdaMuon, HTMuon, COSMOS, Conda, ASGO, SAGE, and Magma, to name a few. The https://t.co/y6RykqhzL2" / X1 savers
- Introducing HealthBench | OpenAI1 savers
- openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! ·1 savers
- Gaming Worlds Could Be The Answer To AI’s Data Problem1 savers
- Research Engineer/Scientist - Human Alignment, Consumer Devices @ OpenAI1 savers
- SoC #5: The Computer Was Only a Transition1 savers
- On Optimism for Interpretability3 savers
- Chess-GPT’s Internal World Model | Adam Karvonen2 savers
- Gmail106 savers
- Writing for LLMs So They Listen · Gwern.net3 savers
- Curius / Bookmarks for the extremely curious182 savers
- The Bitter Lesson78 savers
- i wish we’d grown up on the same advice - by vincent huang61 savers
- How I've run major projects | benkuhn.net52 savers
- finding the right people - by Nicole - startingfromnix49 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- No one can teach you to have conviction | benkuhn.net39 savers
- Andrej's advice for success39 savers
- Advice for Early Career — Celine Halioua39 savers
- The First Fully General Computer Action Model | blog35 savers
- On the Biology of a Large Language Model32 savers
- Why We Think | Lil'Log25 savers
- You don't have to be busy to be prolific | thesephist.com22 savers
- Zoom In: An Introduction to Circuits22 savers
- The Gentle Singularity - Sam Altman21 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- Dario Amodei — The Urgency of Interpretability17 savers
- Neural Networks, Manifolds, and Topology -- colah's blog16 savers
- The Shigalyovist Turn16 savers
- The Era of Experience Paper.pdf15 savers
- The discomfort of intimacy - by Kasra - Bits of Wonder15 savers
- the agony of eros: dating - by Ava - bookbear express14 savers
- Tips for Empirical Alignment Research — AI Alignment Forum14 savers
- As Rocks May Think | Eric Jang13 savers
- Demystifying evals for AI agents \ Anthropic12 savers
- cdixon | The idea maze11 savers
- On neural scaling and the quanta hypothesis10 savers
- Reward Hacking in Reinforcement Learning | Lil'Log10 savers
- The Artificial Intelligence Revolution: Part 2 - Wait But Why9 savers
- GenAI Handbook8 savers
- An Ambitious Vision for Interpretability — AI Alignment Forum8 savers
- How not to do research - Rajan Agarwal7 savers
- The Zero-Day Flaw in AI Companies7 savers
- The Big LLM Architecture Comparison6 savers
- Futarchy: Vote Values, But Bet Beliefs6 savers
- State of RL for reasoning LLMs | A. Weers6 savers
- Man-Computer Symbiosis5 savers
- What Google Learned From Its Quest to Build the Perfect Team - The New York Times5 savers
- ⭐️ Fast LLM Inference From Scratch4 savers
- Leaving 1X | Eric Jang3 savers
- Exclusive: Emmett Shear Is Back With a New Company and A Lot of Alignment3 savers
- Features as Rewards: Using Interpretability to Reduce Hallucinations2 savers
highlights — 362
$500 billion into a combined data center and supercomputer over a four-year horizon.
Wiring Capital to ComputeSRAM operates orders of magnitude faster than HBM, but it’s not space-efficient.
MatX AI Chip Startup Raises $500M to Challenge Nvidia | 2026That flexibility enables the Rubin GPU to deliver up to 50 petaflops of NVFP4 inference performance while preserving accuracy.
Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI | NVIDIA Technical BlogModern GPUs use pipelining to stream execution of multiple sequential instructions at once while ensuring that instructions with dependencies on each other still get executed sequentially.
adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up ·While most instructions can be executed synchronously, these load-store operations are asynchronous, meaning the rest of the instruction execution has to be built around these long wait times.
adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up ·Programming language design has come up as an ancillary area of study to accelerate kernel development.
How to Land a Frontier Lab JobAnd your coding agent will always beat you when you set up a question that introduces these concepts.
How to Land a Frontier Lab JobThis comes with a few notable caveats we’ll explore in the problems below, particularly with respect to quantization (e.g., if we quantize our activations but still do full-precision FLOPs), but it’s a good rule to remember.
All About Rooflines | How To Scale Your ModelOn an H100, this is about 3.35TB/s and on TPU v6e this is about 1.6TB/s.
All About Rooflines | How To Scale Your ModelMost research teams evaluating a data partner care about the same four things: quality, scale, speed, and diversity.
Reintroducing SieveOur view was that video data operates at a level of granularity and scale where collection, QA, and labeling have to be built tech-first.
Reintroducing SieveOur APIs were functioning as highly effective annotation systems for research teams
Reintroducing SieveIt has only been about one human generation since human cloning became technologically feasible.
The future of AI is already written | Mechanize Inc.The true test of whether humanity can control technology lies in its experience with technologies that provide unique, irreplaceable capabilities.
The future of AI is already written | Mechanize Inc.The vertebrate eye and the cephalopod eye evolved completely independently, yet both converged on a remarkably similar camera-type design.
The future of AI is already written | Mechanize Inc.It is the singletons - discoveries made only once in the history of science - that are the residual cases, requiring special explanation.” This pattern suggests that technologies emerge almost spontaneously when the necessary conditions are in place.
The future of AI is already written | Mechanize Inc.Upon closer examination, however, it becomes clear that this is a false choice. Autonomous agents that fully substitute for human labor will inevitably be created because they will provide immense utility that mere AI tools cannot.
The future of AI is already written | Mechanize Inc.In math and physics, a result posted on arXiv (with a minimum hurdle) is fine.
Surgeons Should Not Look Like Surgeons | by Nassim Nicholas Taleb | INCERTO | MediumBut that’s their jobs: as I keep reminding the reader, counter to the common belief, executives are different from entrepreneurs and are supposed to look like actors.
Surgeons Should Not Look Like Surgeons | by Nassim Nicholas Taleb | INCERTO | MediumWhy? Simply the one who doesn’t look the part, conditional of having made a (sort of) successful career in his profession, had to have much to overcome in terms of perception.
Surgeons Should Not Look Like Surgeons | by Nassim Nicholas Taleb | INCERTO | Mediumthe returns arrive sideways, months later, as the collaboration or the reference or the role you couldn't have applied for.
vivek on X: "how to be good at research" / Xevery experiment should be reproducible from its config, and comparing two runs should take seconds, not an afternoon of archaeology.
vivek on X: "how to be good at research" / Xthe stories about alec radford rarely involve a single stroke of genius. they involve volume. more runs per day, more wrong ideas discarded per week, a model of reality that updated faster than anyone else's.
vivek on X: "how to be good at research" / Xa working sense of how gpus actually move memory tells you which architecture papers are doomed before the benchmarks do.
vivek on X: "how to be good at research" / XYou must imagine yourself happy in any number of worlds.
The Old World Is Dying: Advice for 2026 graduatesThe people who will win are those who can remake themselves again and again: to summit one peak, descend it, and then hike up the next.
The Old World Is Dying: Advice for 2026 graduatesRegret. As far as I'm aware, we have one shot. Any and all stumbles do cost us, even if marginally so. With projects of great stature, the significance of any error scales exponentially with time.
Sulaiman Khan Ghori | Thiel FellowshipI didn’t realize that after about 90mins, you use up all your glycogen stores and if you aren’t consuming calories, you will literally start to malfunction.
The month that triathlon took over my life | by Grace Gerwe | MediumHe told me I was a strong swimmer and that I shouldn’t bother with breastroke, which got me super hyped up, and lead to me officially buying the ticket for the race on Wednesday.
The month that triathlon took over my life | by Grace Gerwe | MediumWe believe there is potential for intelligent alien life to already exist, but we see no evidence and thus operate on the assumption that they are not currently out there.
Khan Space IndustriesAI pundits believe scaling LLMs will eventually bypass these obstacles and they might be right
Google’s Genie 3 Is What Science Fiction Looks LikeThis ability also increases the breadth of counterfactual, or “what if” scenarios, that can be used by agents learning from experience to handle unexpected situations.
Genie 3: A new frontier for world models — Google DeepMindHowever, generating an environment auto-regressively is generally a harder technical problem than generating an entire video, since inaccuracies tend to accumulate over time.
Genie 3: A new frontier for world models — Google DeepMindBut today even “small” models run so close to hardware limits that doing novel research requires you to think about efficiency at scale.
How To Scale Your ModelWhen you introduce hidden information, all these past techniques just fall apart
Noam Brown | Innovators Under 35This suggests a different target for style-aware generative models, for which I propose a concrete implementation strategy using contrastive learning.
Style-Aware Generative Models · Gwern.netWhile diffusion models describe paths between noise and data by predicting the tangent direction at each point along the path, flow maps are instead able to predict any point on a path from any other point on that same path.
Learning the integral of a diffusion model – Sander Dielemana professor at the university of michigan’s space engineering program made a list of his 10 best students of all time — 5 out of 10 were working at spacex.
life updates, 2026 q0.5 - by Hardeep GambhirResearch every single one of your interviewers. Understand where they are coming from and what they do, so that you can ask relevant questions.
How I Got a Job at Google DeepMind (No ML Degree) | MediumDeep Deterministic Policy Gradient (DDPG) and Soft Actor-Critic (SAC) that were more sample efficient due to reusing the same data multiple times.
adam-maj/robotics: A deep dive on the history of robotics and the future of humanoids ·I also added ruler grids to the table. Why is this important? So I can run eval on the same episode over and over again and compare between different models, as the performance difference will be truly model differences instead of eval episode setup differences.
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsIt turns out that they have identical serial numbers, so I ended up using the physical USB path to differentiate between them two and encoded them as udev rules.
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsThis helped with detecting overfitting and selecting checkpoints.
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsLimited data diversity
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsIt turns out that I accidentally lost my old calibration files during a power cycle as they were saved in a temporary folder!
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsI found that my cameras had, in fact, moved a little between training and testing
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsI realized it was because my two webcams were identical to each other. That confused my computer and so the camera’s USB paths would be randomly changed every once in a while.
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and LessonsBy considering a new variable, one that’s obvious now in hindsight, such as memory bandwidth, we realize that the operation can be restructured to avoid materializing intermediate values in slow HBM memory.
How to Land a Frontier Lab JobSupervised probe (ceiling)
3 Challenges and 2 Hopes for the Safety of Unsupervised ElicitationCCS. Train the probe by minimizing the CCS loss on a training set, as described in Burns et al.
3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation