flâneur

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

arxiv.org · 6,932 words · saved by 1 readers

N/A

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold 1 1 2 3 1 2 Amrith Setlur , Saurabh Garg , Xinyang (Young) Geng , Naman Garg , Virginia Smith and Aviral Kumar 1 2 3 Carnegie Mellon…

related reading