UKGovernmentBEIS/inspect_ai: Inspect: A framework for large language model evaluations ·
github.com · 400 words · saved by 1 readers
Inspect: A framework for large language model evaluations
Welcome to Inspect, a framework for large language model evaluations created by the UK AI Security Institute . Inspect provides many built-in components, including facilities for prompt engineering, tool usage, multi-turn dialog, and model graded evaluations. Extensions to Inspect (e.g. to support new elicitation and scoring techniques) can be provided by other Python packages. To get started with Inspect, please see the documentation at https://inspect.aisi.org.uk/ . Inspect also includes a collection of over 200 pre-built evaluations ready to run on any model (learn more at https://inspect.a
related reading
- Inspectinspect.ai-safety-institute.org.uk
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Together AI | The AI Native Cloudtogether.ai
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- PromptLayer — Prompt Management, Evals & Observabilitypromptlayer.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Langfuselangfuse.com
- LangChain: the open agent platform to own your intelligencelangchain.com
- Goodfire AIgoodfire.ai