combinatorial reasoning environments for LLMs and RL | Alex Wa's Blog
Can RL agents learn to play spatial reasoning puzzle games as well as, or better than, LLMs? We develop a complete RL pipeline by developing an environment for fruit box (a grid-based reasoning game) using Prime Intellect’s verifiers library, benchmarking LLMs like gpt-5.1 and gemini-3-pro, and training RL agents with SFT and GRPO to play. Repo here.
Can RL agents learn to play spatial reasoning puzzle games as well as, or better than, LLMs? We develop a complete RL pipeline by developing an environment for fruit box (a grid-based reasoning game) using Prime Intellect’s verifiers library, benchmarking LLMs like gpt-5.1 and gemini-3-pro, and training RL agents with SFT and GRPO to play. Repo here. This blog is structured chronologically. First, writing scripted policies that reveal strategy and help validate the core environment mechanics. Second, developing the environment and integrating it with Prime Intellect’s verifiers library to…
saved by
related reading
- As Rocks May Think | Eric Jangevjang.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- DeepSeek-R1arxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Explore | alphaXivalphaxiv.org
- Language Models can Solve Computer Tasksarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chessarxiv.org
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning - ACL Anthologyaclanthology.org