physics-IQ-benchmark/README.md at main · google-deepmind/physics-IQ-benchmark
Physics-IQ is a high-quality, realistic, and comprehensive benchmark dataset for evaluating physical understanding in generative video models. Project website: physics-iq.github.io The best possible score on Physics-IQ is 100.0%, this score would be achieved by physically realistic videos that differ only in physical randomness but adhere to all tested principles of physics. If you test your model on Physics-IQ and would like your score/paper/model to be featured here in this table, feel free to open a pull request that adds a row to the table and we'll be happy to include it! Note to early adopters of the benchmark: results from the paper were finalized on February 19, 2025; if you used the toolbox before please re-run since we changed and improved a few aspects. Likewise, if you downloaded the dataset before that date, it is recommended to re-download it, ensuring the ground truth video masks have a duration of five seconds. Visit the Google Cloud Storage link to download the dataset
Explore this link on the map →saved by
related reading
- Solving Physics Olympiad via Reinforcement Learning on Physics Simulatorssim2reason.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Physical Intelligence (@physical_int) / Xtwitter.com
- Explore | alphaXivalphaxiv.org
- Replicate - Run AI with an APIreplicate.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Contra Labs - Powered by Contracontralabs.com
- [2411.02385] How Far is Video Generation from World Model: A Physical Law Perspectivearxiv.org
- Physics of Language Modelsphysics.allen-zhu.com
- benchmarks.bio — Agentic AI benchmarks on messy, real-world biological databenchmarks.bio
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com