flâneur — a map of the web's best reading

Model Merging — a biased overview | GLADIA

gladia-research-group.github.io · 4,146 words · saved by 1 readers

I recently attended an Estimathon game. This is basically a quiz where you estimate for some hard-to-quantify question like “How many cabs are there in New York?”without using any tools. I was a total disaster. But let me ask you one Estimathon-style question I just made up: “How many models were there on HuggingFace one year ago?” Take a moment to think before scrolling. Got your answer? Nice. Click to reveal the answer. Were you close? No? Okay, another chance: “How many models are there on HuggingFace today?” As above, think about it. You’re basically guessing the growth rate of HuggingFace itself. Almost got it this time? Cool, you win a t-shirt or something. As you can see, the number has more than doubled in just one year! With this explosion of models, a natural question comes up: should we keep making new ones, or spend more effort reusing what we already have? If, like me, you lean toward the latter in many practical cases, this blogpost is for you. Apparently, that’s also the

Model Merging — a biased overview | GLADIA Model Merging — a biased overview A friendly tour of model merging, suspiciously aligned with my own research. The HuggingFace Universe. Credits to the Workshop on Weight Space Learning Disclaimer : This is not a survey. It's more of a tour where my own work keeps getting suspiciously good seats. I promise I'll try a more balanced and comprehensive one in the future. Motivation I recently attended an Estimathon game. This is basically a quiz where you estimate for some hard-to-quantify question like “How many cabs are there in New York?” Around 12,000

Explore this link on the map →

related reading