Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon University
Figure 1: Training models to optimize test-time compute and learn "how to discover" correct responses, as opposed to the traditional learning paradigm of learning "what answer" to output. The major strategy to improve large language models (LLMs) thus far has been to use more and more high-qualit
Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning Research Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem Authors Amrith Setlur by Amrith Setlur --> Affiliations Published January 8, 2025 DOI Figure 1 : Training models to optimize test-time compute and learn “how to discover” correct responses, as opposed to the traditional learning paradigm of lea
saved by
related reading
- Composer2.pdfcursor.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- As Rocks May Think | Eric Jangevjang.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- LLM Resourcesforrestbicker.com
- Why test-time training? – Rabbitholessarahpannn.github.io
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com