flâneur — a map of the web's best reading

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning - ACL Anthology

aclanthology.org · 1,676 words · saved by 1 readers

Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, Junge Zhang [EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning](https://aclanthology.org/2025.acl-long.747/) (Liu et al., ACL 2025) ACL materials are Copyright © 1963–2026 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License. The ACL Anthology is managed and built by the ACL Anthology team of volunteers. Site last built on 07 January 2026 at 01:10 UTC with commit ed95914.

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning - ACL Anthology EPO : Explicit Policy Optimization for Strategic Reasoning in LLM s via Reinforcement Learning Xiaoqian Liu , Ke Wang , Yongbin Li , Yuchuan Wu , Wentao Ma , Aobo Kong , Fei Huang , Jianbin Jiao , Junge Zhang Correct Metadata for Use this form to create a GitHub issue with structured data describing the correction. You will need a GitHub account. Once you create that issue, the correction will be reviewed by a staff member. ⚠️ Mobile Users: Submitting this form to create a new issue wil

Explore this link on the map →

saved by

related reading