flâneur — a map of the web's best reading

Inspect

inspect.ai-safety-institute.org.uk · 1,243 words · saved by 1 readers

Open-source framework for large language model evaluations

Inspect Welcome Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs . Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding. Core features of Inspect include: Composable building blocks—datasets, agents, tools, and scorers—that make evaluations easy to write and reuse. A collection of over 200 pre-built evaluations ready to run on any model. Extensive tooling, including a web-based Inspect View tool for monitoring and visualizing evaluation

Explore this link on the map →

related reading