Data Management For Large Language Models: A Survey
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Data plays a fundamental role in the training of Large Language Models (LLMs). Effective data management, particularly in the formulation of a well-suited training dataset, holds significance for enhancing model performance and improving training efficiency during pretraining and supervised fine-tuning phases. Despite the considerable importance of data management, the curren
Data Management For Large Language Models: A Survey Zige Wang 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Wanjun Zhong 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Yufei Wang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Qi Zhu 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Fei Mi 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Baojun Wang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Lifeng Shang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Xin Jiang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Qun Liu 2
related reading
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- A Bitter Lesson for Data Filteringarxiv.org
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org
- Large Language Diffusion Modelsarxiv.org
- 2403.09611.pdfarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2409.19759] Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMsarxiv.org
- DataRater: Meta-Learned Dataset Curationarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- [2505.04741] When Bad Data Leads to Good Modelsarxiv.org