Data Management For Large Language Models: A Survey
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Data plays a fundamental role in the training of Large Language Models (LLMs). Effective data management, particularly in the formulation of a well-suited training dataset, holds significance for enhancing model performance and improving training efficiency during pretraining and supervised fine-tuning phases. Despite the considerable importance of data management, the curren
Data Management For Large Language Models: A Survey Zige Wang 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Wanjun Zhong 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Yufei Wang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Qi Zhu 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Fei Mi 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Baojun Wang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Lifeng Shang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Xin Jiang 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Qun Liu 2
Explore this link on the map →related reading
- Large Language Diffusion Modelsarxiv.org
- A Bitter Lesson for Data Filteringarxiv.org
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- DataRater: Meta-Learned Dataset Curationarxiv.org
- 2403.09611.pdfarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2605.15220] Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Timearxiv.org
- The bitter lesson of LLM evalsparsed.com
- [2505.04741] When Bad Data Leads to Good Modelsarxiv.org