✳flâneur — a map of the web's best reading
Part 1: The Map Was Wrong — Nemostation
nemostation.com · 3,952 words · saved by 1 readers
A 2B open-source video model audits the standard video-caption benchmarks and finds 70% of the ground-truth captions are wrong.
A month ago, we came across the Tarsier model , which was specifically trained for dense video captioning. Something felt off, though. The captions were dense, but there was no information about the time frames where those events occurred. We searched through the existing video-understanding benchmarks for models that were good at this, but the results were subpar. Most benchmarks don't focus on the temporal axis at all, and the ones that do try to measure it use text-based judges, which introduce several biases. So we built Marlin-2B and a new audit pipeline. The rest of this post is what we
Explore this link on the map →saved by
related reading
- gpt-4.pdfcdn.openai.com
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times - ACL Anthologyaclanthology.org
- Composer2.pdfcursor.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Textharvey-fin.github.io
- Video models are zero-shot learners and reasonersarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- 2025: The year in LLMssimonwillison.net
- Paper AI Tigersgleech.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org