[2509.24372] Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Abstract:Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art LLMs. In contrast, evolution strategies (ES) has largely been overlooked due to the widespread belief that it does not scale to modern model sizes. This paper overturns this assumption by demonstrating the first successful application of ES to full-parameter fine-tuning of LLMs at the billion-parameter scale, without dimensionality reduction. ES can indeed search over extremely high-dimensional parameter spaces and outperform established RL implementations across multiple axes, including improved tolerance to long-horizon and delayed rewards, robustness across diverse base LLMs, reduced susceptibility to reward hacking, and improved training stability. These findings suggest that ES is not merely a viable alternative to RL, but a fundamentally different and powerful backpropagation-free post-training paradigm that opens a new direction for LLM fine-tuning beyond current RL-based approaches. The source codes are provided at: this https URL.
[2509.24372] Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Skip to main content Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2509.24372 (cs) [Submitted on 29 Sep 2025 ( v1 ), last revised 14 Jul 2026 (this version, v3)] Title: Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Authors: Xin Qiu , Yulu Gan , Conor F. Hayes , Qiyao Liang , Yinggan Xu , Roberto Dailey , Elliot Meyerson , Babak Hodjat , Risto Miikkulainen View a PDF of the paper titled Evolution Strategies at Scale:
Explore this link on the map →related reading
- GitHub - emparu/Evolution-Strategies-LLMs: Evolutionary Strategies for RL in LLMs. · GitHubgithub.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org
- GenAI Handbookgenai-handbook.github.io
- [2511.16652] Evolution Strategies at the Hyperscalearxiv.org
- LLM Resourcesforrestbicker.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Recent Advances in Language Model Fine-tuningruder.io