Inspect
inspect.ai-safety-institute.org.uk · 1,243 words · saved by 1 readers
Open-source framework for large language model evaluations
Inspect Welcome Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs . Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding. Core features of Inspect include: Composable building blocks—datasets, agents, tools, and scorers—that make evaluations easy to write and reuse. A collection of over 200 pre-built evaluations ready to run on any model. Extensive tooling, including a web-based Inspect View tool for monitoring and visualizing evaluation
related reading
- GitHub - UKGovernmentBEIS/inspect_ai: Inspect: A framework for large language model evaluations · GitHubgithub.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Together AI | The AI Native Cloudtogether.ai
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- Langfuselangfuse.com
- AI agent evaluation frameworks for production - Vercelvercel.com
- Tasks – Inspectinspect.aisi.org.uk
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Goodfire AIgoodfire.ai