✳flâneur — a map of the web's best reading
Can LLMs invent better ways to train LLMs?
sakana.ai · 1,603 words · saved by 1 readers
Can LLMs invent better ways to train LLMs?
--> Summary At Sakana AI, we harness nature-inspired ideas such as evolutionary optimization to develop cutting-edge foundation models. The development of deep learning has historically relied on extensive trial-and-error by AI researchers and their theoretical insights. This is especially true for preference optimization algorithms, which are crucial for aligning Large Language Models (LLMs) with human preferences. Meanwhile, LLMs themselves have grown increasingly capable of generating hypotheses and writing code. This raises an intriguing question: can we leverage AI to automate the process
Explore this link on the map →related reading
- LLM Daydreaming · Gwern.netgwern.net
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- GenAI Handbookgenai-handbook.github.io
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learningarxiv.org
- [2506.13131] AlphaEvolve: A coding agent for scientific and algorithmic discoveryarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- Best explanations of how LLMs work | Roman Vorushinvorushin.github.io