✳flâneur — a map of the web's best reading
openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering ·
github.com · 2,393 words · saved by 1 readers
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
Explore this link on the map →saved by
related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- PostTrainBenchposttrainbench.com
- ML Contestsmlcontests.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- DX Research Archivegetdx.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- EconEvals: Benchmarks and Litmus Tests for LLM Agents in Unknown Environmentsarxiv.org