flâneur — a map of the web's best reading

Memory Efficient RL | Unsloth Documentation

docs.unsloth.ai · 2,946 words · saved by 1 readers

We're excited to introduce more efficient reinforcement learning (RL) in Unsloth with multiple algorithmic advancements: 1.2 to 1.7x increased context lengths with no slowdown and no extra memory usage! 10% faster RL training runs with revamped kernels and async data movements 2x faster torch.compile times during model loading Unsloth already increases RL training speed, context window and reduces VRAM usage by 50–90% vs. all other setups with FA2, but now Unsloth's Standby improves this even further. Our Standby feature uniquely limits speed degradation compared to other implementations and sometimes makes training even faster! Now, Qwen3-32B LoRA 16-bit can attain 6,144 context lengths vs 3,600 (1.7x longer) before on 1xH100 80GB GPU. Llama-3.1-8B QLoRA 4bit can attain 47,500 lengths vs 42,000 before (1.13x longer). We made RL runs 10% faster through various kernel optimizations, and removed the LoRA communication channel between the CPU and GPU when switching from training to infer

Memory Efficient RL | Unsloth Documentation 🇺🇸 English Get Started 🦥 Homepage 🔮 Models ⭐ Beginner? 📒 Unsloth Notebooks 📥 Installation 🧬 Fine-tuning Guide 💡 Reinforcement Learning 🌀 7x Longer Context RL 👁️‍🗨️ Vision RL 🎱 FP8 RL ⚡ Tutorial: GRPO Training 🧩 Advanced RL Docs Memory Efficient RL 🏆 DPO, ORPO, KTO New Unsloth for AMD 🦥 Introducing Unsloth Studio Unsloth Updates Models Complete LLM Directory GLM-5.2 💜 Qwen3.6 ✨ Gemma 4 DeepSeek-V4 Inkling 🪽 Run MTP Models DiffusionGemma 💜 Qwen3.5 Basics Unsloth Start Unsloth API 🖥️ Inference & Deployment Claude Code OpenAI Codex 🦥

Explore this link on the map →

related reading