[2503.01840] EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2503.01840] EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2503.01840 (cs) [Submitted on 3 Mar 2025 ( v1 ), last revised 23 Apr 2025 (this version, v3)] Title: EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test Authors: Yuhui Li , Fangyun Wei , Chao Zhang , Hongyang Zhang View a PDF of the paper titled EAGLE-3: Scaling u
Explore this link on the map →related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- Composer2.pdfcursor.com
- More Efficient In-Context Learning with GLaMblog.research.google
- The State of LLM Reasoning Model Inferencesebastianraschka.com
- Speculative decodingaarnphm.xyz
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- LLM Resourcesforrestbicker.com
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- o3 — LessWronglesswrong.com
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com