Defeating the Training-Inference Mismatch via FP16
arxiv.org · 5,403 words · saved by 1 readers
N/A
Defeating the Training-Inference Mismatch via FP16 Penghui Qi*†1,2 , Zichen Liu*1,2 , Xiangxin Zhou*1 , Tianyu Pang1 , Chao Du1 , Wee Sun Lee2 , Min Lin1 1 Sea AI Lab 2 National…
saved by
related reading
- The 4-bitter Lesson | humans&humansand.ai
- RL is even more information inefficient than you thoughtdwarkesh.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2506.09501] Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inferencearxiv.org
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Modelsarxiv.org
- RL Post-Training on Macs | Pluralis Researchpluralis.ai
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- Composer2.pdfcursor.com