linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training ·
github.com · 2,903 words · saved by 1 readers
Efficient Triton Kernels for LLM Training
Stable Nightly Discord Installation | Getting Started | Examples | High-level APIs | Low-level APIs | Cite our work Latest News 🔥 [2025/12/19] We announced a liger kernel discord channel at https://discord.gg/X4MaxPgA; We will be hosting Liger Kernel x Triton China Meetup in mid of January 2026 [2025/03/06] We release a joint blog post on TorchTune × Liger - Peak Performance, Minimized Memory: Optimizing torchtune’s performance with torch.compile & Liger Kernel [2024/12/11] We release v0.5.0: 80% more memory efficient post training losses (DPO, ORPO, CPO, etc)! [2024/12/5] We release…
saved by
related reading
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- Together AI | The AI Native Cloudtogether.ai
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Tinkerthinkingmachines.ai
- GitHub - NVIDIA-NeMo/Automodel: 🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face supportgithub.com
- API Reference — TensorRT LLMnvidia.github.io
- Llama 2 · Hugging Facehuggingface.co
- Hugging Face – The AI community building the future.huggingface.co
- PostTrainBenchposttrainbench.com
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co
- GitHub - triton-lang/triton: Development repository for the Triton language and compiler · GitHubgithub.com
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co