Arize-ai/phoenix: AI Observability & Evaluation ·
github.com · 1,313 words · saved by 1 readers
AI Observability & Evaluation
Arize Phoenix is Arize's open-source AI observability platform designed for experimentation, evaluation, and troubleshooting. For managed production workflows, Arize also offers Arize AX. Phoenix provides: Tracing - Trace your LLM application's runtime using OpenTelemetry-based instrumentation. Evaluation - Leverage LLMs to benchmark your application's performance using response and retrieval evals. Datasets - Create versioned datasets of examples for experimentation, evaluation, and fine-tuning. Experiments - Track and evaluate changes to prompts, LLMs, and retrieval. Playground-…
saved by
related reading
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- Together AI | The AI Native Cloudtogether.ai
- Langfuselangfuse.com
- PromptLayer — Prompt Management, Evals & Observabilitypromptlayer.com
- Cookbookcookbook.openai.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- LangChain: the open agent platform to own your intelligencelangchain.com
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- Goodfire AIgoodfire.ai
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Modelsgithub.com
- The vector database to build knowledgeable AI | Pineconepinecone.io