flâneur — a map of the web's best reading

Prompt-to-Leaderboard

arxiv.org · 14,604 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Prompt-to-Leaderboard Evan Frick ∗ , Connor Chen ∗ , Joseph Tennyson ∗ , Tianle Li ∗ , Wei-Lin Chiang ∗ , Anastasios N. Angelopoulos ∗ , Ion Stoica {evanfrick, connorchen, josephtennyson, tianleli, weichiang, angelopoulos, istoica}@berkeley.edu (University of California, Berkeley March 10, 2025 *equal contribution ) Abstract Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and prompt-specific variations in model performance. To address this, we propose Prompt-to-Leade

Explore this link on the map →

related reading