Prompt-to-Leaderboard
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Prompt-to-Leaderboard Evan Frick ∗ , Connor Chen ∗ , Joseph Tennyson ∗ , Tianle Li ∗ , Wei-Lin Chiang ∗ , Anastasios N. Angelopoulos ∗ , Ion Stoica {evanfrick, connorchen, josephtennyson, tianleli, weichiang, angelopoulos, istoica}@berkeley.edu (University of California, Berkeley March 10, 2025 *equal contribution ) Abstract Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and prompt-specific variations in model performance. To address this, we propose Prompt-to-Leade
related reading
- Bradley–Terry model - Wikipediaen.wikipedia.org
- Composer2.pdfcursor.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF | John Lambertjohnwlambert.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- BalatroBenchbalatrobench.com
- PostTrainBenchposttrainbench.com
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Modelarxiv.org