flâneur

Compiling Models to Megakernels - by Luminal and Joe Fioti

blog.luminal.com · 2,776 words · saved by 1 readers

Fine-grained synchronization, deep pipelines, and zero kernel launch overheads, automatically.

Compiling Models to Megakernels Fine-grained synchronization, deep pipelines, and zero kernel launch overheads, automatically. Luminal and Joe Fioti Jan 09, 2026 8 2 Share Luminal is an inference compiler, and as such we’re interested in driving inference right up to the physical limits of the hardware. Inference has two fundamental limitations: compute (flops) and bandwidth (TB/s). Increasing these two requires buying much more expensive hardware, so we want to make sure we’re using all the compute and bandwidth we have available to us! This basically boils down to: anytime the GPU is not loa

saved by

related reading