flâneur — a map of the web's best reading

Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

arxiv.org · 13,739 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker’s weekly cartoon caption contest over the past eight years. This unique dataset supports the development and evaluation of multimodal large language models and preference-based fine-tuning algorithms for humorous caption generation. We propose novel benchmarks for judging the quality of model-generated captions, utilizing both GPT4 and human judgments to establish ranking-based evaluation strategies. Our experimental results highlight the limitations of current fine-tuning methods, such as RLHF and DPO, when applied to creative tasks. Furthermore

Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning Jifan Zhang 1 , Lalit Jain 2∗ , Yang Guo 1∗ , Jiayi Chen 1 , Kuan Lok Zhou 1† , Siddharth Suresh 1 , Andrew Wagenmaker 2 , Scott Sievert 1 , Timothy Rogers 1 , Kevin Jamieson 2 , Robert Mankoff 3 , Robert Nowak 1 1 University of Wisconsin-Madison, 2 University of Washington, Seattle, 3 Air Mail and Cartoon Collections lalitj@uw.edu, {jifan,yguo}@cs.wisc.edu Equal contribution.Equal contribution. Abstract We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million h

Explore this link on the map →

related reading