[2403.02327] Model Lakes
Abstract:Given a set of deep learning models, it can be hard to find models appropriate to a task, understand the models, and characterize how models are different one from another. Currently, practitioners rely on manually-written documentation to understand and choose models. However, not all models have complete and reliable documentation. As the number of models increases, the challenges of finding, differentiating, and understanding models become increasingly crucial. Inspired from research on data lakes, we introduce the concept of model lakes. We formalize key model lake tasks, including model attribution, versioning, search, and benchmarking, and discuss fundamental research challenges in the management of large models. We also explore what data management techniques can be brought to bear on the study of large model management.
Model Lakes Koyena Pal David Bau Renée J. Miller Northeastern University Northeastern University Northeastern & U. Waterloo Boston, Massachusetts, USA Boston, Massachusetts, USA Waterloo, ON, Canada…
saved by
related reading
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- GenAI Handbookgenai-handbook.github.io
- The Little Book of Deep Learningfleuret.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- laguna-m1-xs2-technical-report.pdfpoolside.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Training Language Models to Explain Their Own Computationsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Training Language Models to Explain Their Own Computationsarxiv.org
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org