Julian Quevedo
12 followers · 15 following · 1576 views
on the atlas — 102
- Game Emulation via Neural Network2 savers
- Xianbang Wang — research & writing2 savers
- game-emulation-via-neural-network-dataset/Example_Diffusion_World_Model.ipynb at main · madebyollin/game-emulation-via-neural-network-dataset1 savers
- World Emulation via Neural Network1 savers
- LeNet-5: Unusual Patterns1 savers
- MNIST Demos on Yann LeCun's website1 savers
- Newmu/dcgan_code: Deep Convolutional Generative Adversarial Networks1 savers
- Evolved Virtual Creatures1 savers
- Prafulla Dhariwal1 savers
- Neural Networks for Machine Perception1 savers
- Pantograph2 savers
- A General Goal-Conditioned Minecraft Model - Pantograph4 savers
- Ultrasound imaging of the brain — Aleph23 savers
- Kelsey Pool1 savers
- Alex Gajewski11 savers
- Aleph6 savers
- Verbalizable Representations Form a Global Workspace in Language Models24 savers
- 1959 – "Starship Troopers" Power Suits (Fiction) – Robert Heinlein (American) - cyberneticzoo.com1 savers
- 1953 - "Creakyfoot" Power Suit - E.R. James (British) - cyberneticzoo.com1 savers
- David Buckley1 savers
- HomeMadeGarbage - 電子工作、音楽制作、日常など、家族ブログ。1 savers
- Lena @ Things Of Interest16 savers
- Speeding up RL with high-leverage samples | Applied Compute2 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- Scalable Behavior Cloning with Open Data, Training, and Evaluation1 savers
- joanna's home1 savers
- Top of mind -1 savers
- Ben Reinhardt | About1 savers
- Artificial intelligence by mimicking natural intelligence1 savers
- A little bit of slope makes up for a lot of y-intercept2 savers
- Elon Litman | Elements of a Vector Space1 savers
- Some mistakes I made as a new manager | benkuhn.net2 savers
- Social learning - Linda Tong5 savers
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog24 savers
- The Generalist2 savers
- GPU Performance Background User's Guide - NVIDIA Docs2 savers
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Research6 savers
- CHM Releases AlexNet Source Code - CHM1 savers
- Outperforming cuBLAS on H100: a Worklog2 savers
- Bare-bones Diffusion Models2 savers
- How to use distributed shared memory in CUDA for inter-thread-block communication1 savers
- wangzyon/NVIDIA_SGEMM_PRACTICE: Step-by-step optimization of CUDA SGEMM1 savers
- Reverse-Engineering cuBLAS1 savers
- How to Start Google8 savers
- Stanford’s War on Social Life15 savers
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blog2 savers
- What Every Programmer Should Know About Memory7 savers
- Genie: Generative Interactive Environments - Google DeepMind1 savers
- the world’s largest distributed LLM training job on TPU v5e | Google Cloud Blog1 savers
- Claire Baum | LinkedIn1 savers
- Anna Ivanova1 savers
- Soft Actor-Critic — Spinning Up documentation4 savers
- Groq Inference Tokenomics: Speed, But At What Cost?2 savers
- cs140e-24win/labs/13-ss-equiv/README.md at main · dddrrreee/cs140e-24win1 savers
- Memory-Limited Layers User's Guide - NVIDIA Docs1 savers
- PT2 Architecture - Google Docs1 savers
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blog2 savers
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Models5 savers
- Building for builders - Linda Tong2 savers
- Sithwards Induction, or: The Dark Side for Grad Students - Econlib2 savers
- The Bitter Lesson78 savers
- Nat Friedman57 savers
- Here's to the fools who dream54 savers
- advice - nabeelqu43 savers
- crushes are often just misplaced ambition - by Isabel42 savers
- Andrej's advice for success39 savers
- How Tech Hype is Ruining College Students in Tech37 savers
- how to avoid half-heartedness - by Ava - bookbear express36 savers
- Half-assing it with everything you've got33 savers
- We Need a New Science of Progress - The Atlantic25 savers
- Teach Yourself Computer Science25 savers
- cdixon | Climbing the wrong hill24 savers
- Kernel | all the better to see you with19 savers
- [2304.03442] Generative Agents: Interactive Simulacra of Human Behavior19 savers
- Life is Short15 savers
- The Unreasonable Effectiveness of Recurrent Neural Networks15 savers
- Don't Major in Computer Science - by Varun Shenoy15 savers
- quiet confidence13 savers
- Pmarchive · Pmarca Guide to Career Planning: Opportunity12 savers
- The Annotated S411 savers
- Kernel | Trees Won’t Save Us11 savers
- A friendly introduction to machine learning compilers and optimizers10 savers
- A GPS for the mind | thesephist.com9 savers
- Kullback–Leibler divergence9 savers
- Planning for AGI and beyond9 savers
- Language Models, World Models, and Human Model-Building8 savers
- A Recipe for Training Neural Networks7 savers
- Semantic reconstruction of continuous language from non-invasive brain recordings | Nature Neuroscience7 savers
- quantitative finance - by nancy - general zuo's chicken7 savers
- Chip Huyen7 savers
- Hyperlink maximalism | thesephist.com7 savers
- Stop Treating Women Like Men6 savers
- "Should" considered harmful5 savers
- Simple advice for academic publishing - Marginal REVOLUTION5 savers
- Mosaic LLMs (Part 2): GPT-3 quality for <$500k4 savers
- ACT-1: Transformer for Actions4 savers
- Stanford Alpaca Model Release4 savers
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness3 savers
- Do I wish to revise my time management tips? - Marginal REVOLUTION3 savers
- The Perfect Match - Lightspeed Magazine3 savers
highlights — 848
So, if traditional game worlds are paintings, neural worlds are photographs. Information flows from sensor to screen without passing through human hands.
World Emulation via Neural NetworkSo for this post, to demonstrate what makes neural networks truly special, I wanted to train a neural network on gameplay videos of the actual world.
World Emulation via Neural NetworkSince neural networks are programs, the answer is always yes! They can mimic anything. As long as you specify the behavior you want, really, really precisely.
Game Emulation via Neural NetworkThe networks that fit in your web browser in 2030 will probably be capable of more. Like, a lot more.
Game Emulation via Neural NetworkIn a nutshell, we were genuinely surprised by how far a minimalist, direct-pixel approach could take us. We hope MiniT2I can serve as a clean baseline for text-to-image generation: small enough to debug, strong enough to reveal real failures, and simple enough to reproduce, modify, and challenge.
Xianbang Wang — research & writingBecause the task is so much broader, succeeding at T2I is a much stronger signal that a method genuinely scales — and a more direct reflection of what we actually ask generative models to do.
Xianbang Wang — research & writinglearnt a latent space in a completely unsupervised fashion where ROTATIONS ARE LINEAR in this latent space. WHHHAAATT????!!!!!!
Newmu/dcgan_code: Deep Convolutional Generative Adversarial NetworksFigure 11, trained on imagenet has a plane with bird legs. so cooool.
Newmu/dcgan_code: Deep Convolutional Generative Adversarial NetworksGoal-conditioned models are often more general than models trained to maximize a numerical score. For example, Dreamer 41010Danijar Hafner, Wilson Yan, and Timothy Lillicrap, Training Agents Inside of Scalable World Models, 2025. ↩ is trained on 20 different tasks, each with a manually specified reward function. This limits the range of applicability of the resulting agent to just those 20 tasks, and adding a new task would entail re-training on the new reward function. Our goal-conditioned agents, on the other hand, are trained once, and are able to achieve new goals that we come up with at i…
A General Goal-Conditioned Minecraft Model - PantographContravariance Principle
The Theory of Contravariance, Part 0 - by Dan YaminsThe next step is to communicate in latents
AlephWe believe end-to-end machine learning, trained on large enough datasets, will recover far more signal than current methods can see.
Ultrasound imaging of the brain — AlephOf course, it is impossible to get bitwise identical results between training and inference if we can’t even get bitwise identical results from two identical inference requests. Then, deterministic inference enables us to also modify our training stack to obtain bitwise identical results between sampling and training, thus resulting in true on-policy RL.
Defeating Nondeterminism in LLM Inference - Thinking Machines LabABC-DiT inference at 27.6 Hz on a single 5090.
Scalable Behavior Cloning with Open Data, Training, and Evaluation— learning the ABCs of behavior cloning, together —
Scalable Behavior Cloning with Open Data, Training, and EvaluationBut if you just think about your slope and don't worry about where you start out you'll end up some place nice.
A little bit of slope makes up for a lot of y-interceptEvery leaf on every tree.
Elon Litman | Elements of a Vector Spacesqrt_one_minus_alphas_cumprod
Bare-bones Diffusion ModelsIt’s easy to set out to build a DARPA but end up building a Skunkworks. The dominant paradigm when starting a research organization is to do internalized R&D unless you’re a charitable foundation. In reality, R&D orgs lie on a spectrum between externalized and internalized, with DARPA on one end and Lockheed Skunkworks on the other. There’s nothing inherently wrong with building a Skunkworks –it just means that there are different tradeoffs and the statement “DARPA for X” is misleading.
Why does DARPA work?For seedling projects, as long as it’s below roughly $500k program managers can just write a check. In the past, there was almost no oversight over PM spending after the director authorized the money. This is how J.C.R. Licklider was able to “Johnny Appleseed” computing groups all over the country in only a year.
Why does DARPA work?Easily re-deployed funds lower the overhead to starting something risky and differentiate DARPA from other funding agencies, philanthropies, and venture capital.
Why does DARPA work?DARPA programs have a 5—10% success rate and have included things like jetpacks, earthworm robots, creating fusion with sound waves, spider-man wall climbing, and bomb detecting bees.
Why does DARPA work?“I never really felt constrained by money,” (former DARPA director) Tether says. “I was more constrained by ideas.”
Why does DARPA work?In this sense, the canonical system that’s not a world model is exactly a lookup table.
Language Models, World Models, and Human Model-Buildingby representing some important aspect of the data-generating process—even if incompletely or imperfectly—we can avoid pre-computing all answers to all questions, within whatever restricted set we need to answer.
Language Models, World Models, and Human Model-BuildingWith these differences in affordances come differences in the complexity required to implement each model.
Language Models, World Models, and Human Model-BuildingAnd the simulator lets us answer counterfactual questions of the system by representing something close to its true underlying dynamics (but requires us to do substantially more work to specify the initial conditions for these counterfactuals).
Language Models, World Models, and Human Model-BuildingWe’ve hard-coded much less information than the orrery (we no longer need to bake in the fact that orbits are elliptical—instead this emerges from simulation). As a result, we can answer even more complex questions (what would happen in three years if Jupiter and Saturn changed places today?). Indeed, we can now evaluate a pretty general class of counterfactual questions that involve reasoning about states of the world that are unreachable from its current state.
Language Models, World Models, and Human Model-BuildingNevertheless, we can’t seem to agree on a definition of “world model” any more than we agreed on “meaning representation”.
Language Models, World Models, and Human Model-BuildingI will live my life for me and toward you.
And still - by Evana - arbiter of distasteBut you weren’t, and I was no small planet. And what we both knew is that I didn’t want that, and neither did you.
And still - by Evana - arbiter of distasteSacred sites are often mundane.
And still - by Evana - arbiter of distasteCompilation: As training progresses, the computation and communication that happen across the chips need to be managed effectively by a high-performance compiler.
the world’s largest distributed LLM training job on TPU v5e | Google Cloud BlogThat's it, just two things, build stuff and do well in school.
How to Start Googlesomething your friends actually want.
How to Start Googleif you're young and good at technology, then your unconscious instincts about what's interesting are better than your conscious ideas about what would be a good company.
How to Start GoogleNvidia is not a static target.
Groq Inference Tokenomics: Speed, But At What Cost?Another challenge for Groq is that speculative decoding and techniques such as Medusa are getting better at a rapid clip. Tree/branch speculation approaches are leading to upwards of 3x speedups with speculative decoding. If these can be deployed efficiently on production grade systems, then an 8x H100 system could achieve over 600 tokens per second. That alone would blow away Groq’s advantage in speed.
Groq Inference Tokenomics: Speed, But At What Cost?We struggle to see how Groq could ever implementing extremely large context lengths given the KVCache size requirements. This would require systems of tens of thousands of chips, instead of 10s or 100s of chips as is used with Google, Nvidia, and AMD based inference solutions.
Groq Inference Tokenomics: Speed, But At What Cost?Currently the largest MoE models sit in the 1-2 trillion parameter range, but we expect Google and OpenAI to launch >10 trillion parameter models over the next year that will require inference systems of hundreds of GPUs and 10s of TB of memory.
Groq Inference Tokenomics: Speed, But At What Cost?Groq claims to have a power advantage, but we can’t see that. Even with the most pessimistic assumptions for H100 servers, at 10kW, which would include the CPU and all 8 NICs running full blast, it is more efficient than the 576 chip Groq server which requires 230kW, or 3.2kW per 8 chip server. Groq claimed a performance per watt advantage, but we do not see how that is calculated.
Groq Inference Tokenomics: Speed, But At What Cost?Groq effectively purchases its systems at cost.
Groq Inference Tokenomics: Speed, But At What Cost?With that said, for some reason, whether it be a lack of buffers or the VLIW architecture, Groq’s FLOPS utilization is lower than Nvidia’s even with next weeks push of batch size 3 implemented.
Groq Inference Tokenomics: Speed, But At What Cost?Due to the memory wall, an H100-based inference system typically has low FLOPS utilization, while Groq’s architecture sidesteps the memory wall by having on chip SRAM.
Groq Inference Tokenomics: Speed, But At What Cost?Groq is not competitive architecturally at all for throughput optimized scenarios.
Groq Inference Tokenomics: Speed, But At What Cost?Nvidia buys 80GB of HBM from SK Hynix for ~$1,150 for each H100 chip.
Groq Inference Tokenomics: Speed, But At What Cost?In the case of the Mixtral model, Groq had to connect 8 racks of 9 servers each with 8 chips per server. That’s a total of 576 chips to build up the inference unit and serve the Mixtral model. Compare that to Nvidia where a single H100 can fit the model at low batch sizes, and two chips have enough memory to support large batch sizes.
Groq Inference Tokenomics: Speed, But At What Cost?Because each chip only has 230MB of SRAM, no useful models can actually fit on a single chip. Instead, they must utilize many chips to fit the model and network them together.
Groq Inference Tokenomics: Speed, But At What Cost?The same cannot be said for others offering Mixtral APIs. They are either lying about quantization, or lighting VC money on fire to acquire a customer base.
Groq Inference Tokenomics: Speed, But At What Cost?As for serving many users with extremely high batch sizes, IE throughput and cost optimized, GPUs are king.
Groq Inference Tokenomics: Speed, But At What Cost?