Human Data is (Probably) More Expensive Than Compute for Training Frontier LLMs
Post-training techniques (e.g., supervised fine-tuning and reinforcement learning with verifiable rewards) are crucial to recent advances in LLMs. Unlike pre-training, post-training relies heavily on annotated data provided by humans, often requiring expert input. Fine-tuning with reinforcement learning, the core technique powering today’s most advanced reasoning models, demands not only high-quality data but also verifiable answers. “Scale AI expects to more than double sales to $2 billion in 2025. The startup generated revenue of about $870 million last year,” reported by Bloomberg. The incredible demand for high-quality human-annotated data is fueling soaring revenues of data labeling companies. In tandem, the cost of human labor has been consistently increasing. We estimate that obtaining high-quality human data for LLM post-training is more expensive than the marginal compute itself1 and will only become even more expensive. In other words, high-quality human data will be the bott
Post-training techniques (e.g., supervised fine-tuning and reinforcement learning with verifiable rewards) are crucial to recent advances in LLMs. Unlike pre-training, post-training relies heavily on annotated data provided by humans, often requiring expert input. Fine-tuning with reinforcement learning, the core technique powering today’s most advanced reasoning models, demands not only high-quality data but also verifiable answers. “Scale AI expects to more than double sales to $2 billion in 2025. The startup generated revenue of about $870 million last year,” reported by Bloomberg. The…
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Limits to narrow LLM complementarityosmarks.net
- A World of Verifiable Domainsseancai.com
- How much does it cost to train frontier AI models? | Epoch AIepoch.ai
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Mosaic LLMs: GPT-3 quality formosaicml.com
- AI capabilities can be significantly improved without expensive retraining | Epoch AIepoch.ai
- Thinking about High-Quality Human Data | Lil'Loglilianweng.github.io
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Cheap RL tasks will waste compute | Mechanize, Inc.mechanize.work
- The Future of Meta Superintelligence: A 1 Year Progress Updatenewsletter.semianalysis.com