flâneur — a map of the web's best reading

Part 1: The Map Was Wrong — Nemostation

nemostation.com · 3,952 words · saved by 1 readers

A 2B open-source video model audits the standard video-caption benchmarks and finds 70% of the ground-truth captions are wrong.

A month ago, we came across the Tarsier model , which was specifically trained for dense video captioning. Something felt off, though. The captions were dense, but there was no information about the time frames where those events occurred. We searched through the existing video-understanding benchmarks for models that were good at this, but the results were subpar. Most benchmarks don't focus on the temporal axis at all, and the ones that do try to measure it use text-based judges, which introduce several biases. So we built Marlin-2B and a new audit pipeline. The rest of this post is what we

Explore this link on the map →

saved by

related reading