[2309.14556] Art or Artifice? Large Language Models and the False Promise of Creativity
Abstract:Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.
View PDF HTML (experimental) Abstract:Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality,…
saved by
related reading
- Creativity Has Left the Chat: The Price of Debiasing Language Modelsarxiv.org
- Torrance Tests of Creative Thinkingen.wikipedia.org
- [2312.03746] Evaluating Large Language Model Creativity from a Literary Perspectivearxiv.org
- The Human Skill That Eludes AI - The Atlantictheatlantic.com
- Why A.I. Isn’t Going to Make Art | The New Yorkernewyorker.com
- Lluminatejoelsimon.net
- Why LLMs Are Bad Writers But Good Editors - by Jasmine Sunjasmi.news
- Measuring LLM Novelty As The Frontier Of Original And High-Quality Outputarxiv.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Modifying Large Language Model Post-Training for Diverse Creative Writingarxiv.org
- Contra Labs - Powered by Contracontralabs.com
- Writing for LLMs So They Listen · Gwern.netgwern.net