✳flâneur — a map of the web's best reading
huggingface/nanotron: Minimalistic large language model 3D-parallelism training ·
github.com · 816 words · saved by 1 readers
Minimalistic large language model 3D-parallelism training
⚡️ Nanotron Installation • Quick Start • Features • Benchmarks • Contributing Pretraining models made easy Nanotron is a library for pretraining transformer models. It provides a simple and flexible API to pretrain models on custom datasets. Nanotron is designed to be easy to use, fast, and scalable. It is built with the following principles in mind: Simplicity : Nanotron is designed to be easy to use. It provides a simple and flexible API to pretrain models on custom datasets. Performance : Optimized for speed and scalability, Nanotron uses the latest techniques to train models faster and mor
Explore this link on the map →related reading
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically · GitHubgithub.com
- Replicate - Run AI with an APIreplicate.com
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- Hugging Face · GitHubgithub.com
- Annotated Research Paper Implementations: Transformers, StyleGAN, Stable Diffusion, DDPM/DDIM, LayerNorm, Nucleus Sampling and morenn.labml.ai
- nemotron-3-super-120b-a12b Model by NVIDIA | NVIDIA NIMbuild.nvidia.com
- PostTrainBenchposttrainbench.com