✳flâneur — a map of the web's best reading
Humanity's Last Exam
lastexam.ai · 7,364 words · saved by 1 readers
Humanity's Last Exam Dataset
Humanity's Last Exam Humanity's Last Exam Paper Nature Arxiv Dataset load_dataset(" cais/hle ") GitHub HLE-Rolling Live Submission Dashboard Latest News [01/28/2026] : HLE is published on Nature (Nature 649, 1139–1146). [10/08/2025] : We release a dynamic fork version HLE-Rolling ( update logs ). To contribute, please send us an email at agibenchmark@safe.ai See full notes [04/03/2025] : HLE has been finalized with 2,500 questions. Questions flagged in the bug bounty program and searchable questions have been removed and replaced with other questions. [03/21/2025] : Bug Bounty closed. Thank yo
Explore this link on the map →related reading
- Humanity's Last Machinehumanityslastmachine.com
- Elicit: AI for scientific researchelicit.org
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Elicit: AI for scientific researchelicit.com
- benchmarks.bio — Agentic AI benchmarks on messy, real-world biological databenchmarks.bio
- Dataset list - A list of the biggest machine learning datasetsdatasetlist.com
- ML Contestsmlcontests.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- Humanoid Atlas | Humanoid Robot Supply Chain Map, OEM Database & Industry Analysishumanoids.fyi
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com