✳flâneur — a map of the web's best reading
Inspect
inspect.ai-safety-institute.org.uk · 1,243 words · saved by 1 readers
Open-source framework for large language model evaluations
Inspect Welcome Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs . Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding. Core features of Inspect include: Composable building blocks—datasets, agents, tools, and scorers—that make evaluations easy to write and reuse. A collection of over 200 pre-built evaluations ready to run on any model. Extensive tooling, including a web-based Inspect View tool for monitoring and visualizing evaluation
Explore this link on the map →related reading
- GitHub - UKGovernmentBEIS/inspect_ai: Inspect: A framework for large language model evaluations · GitHubgithub.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Tasks – Inspectinspect.aisi.org.uk
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Langfuselangfuse.com
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- GitHub - EleutherAI/lm-evaluation-harness: A framework for few-shot evaluation of language models. · GitHubgithub.com
- The bitter lesson of LLM evalsparsed.com
- Model optimization | OpenAI APIplatform.openai.com