Part 1: The Map Was Wrong — Nemostation
nemostation.com · 3,952 words · saved by 1 readers
A 2B open-source video model audits the standard video-caption benchmarks and finds 70% of the ground-truth captions are wrong.
A month ago, we came across the Tarsier model , which was specifically trained for dense video captioning. Something felt off, though. The captions were dense, but there was no information about the time frames where those events occurred. We searched through the existing video-understanding benchmarks for models that were good at this, but the results were subpar. Most benchmarks don't focus on the temporal axis at all, and the ones that do try to measure it use text-based judges, which introduce several biases. So we built Marlin-2B and a new audit pipeline. The rest of this post is what we
saved by
related reading
- gpt-4.pdfcdn.openai.com
- Claim-Level Rubric Rewards for Video Caption Reinforcement Learningarxiv.org
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times - ACL Anthologyaclanthology.org
- Composer2.pdfcursor.com
- Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Textharvey-fin.github.io
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Video models are zero-shot learners and reasonersarxiv.org
- PostTrainBenchposttrainbench.com
- Video Generation Models Explosion 2024yenchenlin.me
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com