LiCoEval: Evaluating LLMs on License Compliance in Code Generation
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Recent advances in Large Language Models (LLMs) have revolutionized code generation, leading to widespread adoption of AI coding tools by developers. However, LLMs can generate license-protected code without providing the necessary license information, leading to potential intellectual property violations during software production. This paper addresses the critical, yet underexplored, issue of license compliance in LLM-generated code by establishing a benchmark to evaluate the ability of LLMs to provide accurate license information for their generated code. To establish this benchmark, we conduct an empirical study to identify a reasonable standard for “striking similarity” that excludes the possibility of independent creation, indicating a copy relatio
LiCoEval: Evaluating LLMs on License Compliance in Code Generation Weiwei Xu1, Kai Gao4, Hao He2, Minghui Zhou13 3Minghui Zhou is the corresponding author. 1 School of Computer Science, Peking University, Beijing, China 1 Key Laboratory of High Confidence Software Technologies, Ministry of Education, China 4 University of Science and Technology Beijing, Beijing, China 2 Carnegie Mellon University, Pittsburgh, USA xuww@stu.pku.edu.cn, kai.gao@ustb.edu.cn, haohe@andrew.cmu.edu, zhmh@pku.edu.cn Abstract Recent advances in Large Language Models (LLMs) have revolutionized code generation, leading t
Explore this link on the map →saved by
related reading
- 2025: The year in LLMssimonwillison.net
- Here’s how I use LLMs to help me write codesimonwillison.net
- crawshaw - 2025-01-06crawshaw.io
- Introducing FrontierCode | Cognitioncognition.ai
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingsarxiv.org
- Coding Models Are Doing Too Much | whnrehiew.github.io
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- [2203.07814] Competition-Level Code Generation with AlphaCodearxiv.org
- The L in "LLM" Stands for Lying — Acko.netacko.net
- 2502.18449arxiv.org
- Introducing FrontierCode | Cognitioncognition.com
- Scaling LLMs to larger codebases - Kieran Gillblog.kierangill.xyz