EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning - ACL Anthology
Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, Junge Zhang [EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning](https://aclanthology.org/2025.acl-long.747/) (Liu et al., ACL 2025) ACL materials are Copyright © 1963–2026 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License. The ACL Anthology is managed and built by the ACL Anthology team of volunteers. Site last built on 07 January 2026 at 01:10 UTC with commit ed95914.
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning - ACL Anthology EPO : Explicit Policy Optimization for Strategic Reasoning in LLM s via Reinforcement Learning Xiaoqian Liu , Ke Wang , Yongbin Li , Yuchuan Wu , Wentao Ma , Aobo Kong , Fei Huang , Jianbin Jiao , Junge Zhang Correct Metadata for Use this form to create a GitHub issue with structured data describing the correction. You will need a GitHub account. Once you create that issue, the correction will be reviewed by a staff member. ⚠️ Mobile Users: Submitting this form to create a new issue wil
Explore this link on the map →saved by
related reading
- Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up - ACL Anthologyaclanthology.org
- 2025.acl-long.896.pdfaclanthology.org
- RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models - ACL Anthologyaclanthology.org
- The Role of Deductive and Inductive Reasoning in Large Language Models - ACL Anthologyaclanthology.org
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Explore | alphaXivalphaxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- [2606.10346] Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learningarxiv.org
- Language Models can Solve Computer Tasksarxiv.org