✳flâneur — a map of the web's best reading
EconEvals: Benchmarks and Litmus Tests for LLM Agents in Unknown Environments
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Cookbookcookbook.openai.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- GitHub - agno-agi/agno: Build, run, and manage agent platforms. · GitHubgithub.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- OpenAI | Research & Deploymentopenai.com