flâneur

ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? | alphaXiv

alphaxiv.org · 1,957 words · saved by 1 readers

ZeroCoder introduces a label-free co-evolutionary framework for Large Language Models, significantly enhancing both code and test generation capabilities b

Submitted 28 Jun 2026 LF Lishui Fan MC Mouxiang Chen TZ Tingwei Zhu KL Kui LiuXin Xia SL Shanping Li ZL Zhongxin Liu Abstract Code generation is important in software engineering, and Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm to improve it through execution-based feedback. However, most RLVR pipelines rely on human-curated tests, making progress bottlenecked by scarce and costly supervision. Existing work tried to use self-generated tests to ground rewards, but the lack of discriminative tests constrains the effect due to the sub-optimal…

saved by

related reading