Prompt-to-Leaderboard
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Prompt-to-Leaderboard Evan Frick ∗ , Connor Chen ∗ , Joseph Tennyson ∗ , Tianle Li ∗ , Wei-Lin Chiang ∗ , Anastasios N. Angelopoulos ∗ , Ion Stoica {evanfrick, connorchen, josephtennyson, tianleli, weichiang, angelopoulos, istoica}@berkeley.edu (University of California, Berkeley March 10, 2025 *equal contribution ) Abstract Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and prompt-specific variations in model performance. To address this, we propose Prompt-to-Leade
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Modelarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF Bookrlhfbook.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Bradley–Terry model - Wikipediaen.wikipedia.org
- castform - the training platform for the ai engineercgft.io
- PostTrainBenchposttrainbench.com