✳flâneur — a map of the web's best reading
[2201.02387] The Defeat of the Winograd Schema Challenge
arxiv.org · saved by 1 readers
The Winograd Schema Challenge -- a set of twin sentences involving pronoun reference disambiguation that seem to require the use of commonsense knowledge -- was proposed by Hector Levesque in 2011. By 2019, a number of AI systems, based on large pre-trained transformer-based language models and fine-tuned on these kinds of problems, achieved better than 90% accuracy. In this paper, we review the history of the Winograd Schema Challenge and assess its significance.
Explore this link on the map →related reading
- Winograd schema challenge - Wikipediaen.wikipedia.org
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Language Models can Solve Computer Tasksarxiv.org
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- Group | Sherry Tongshuang Wucs.cmu.edu
- What Does It Mean for AI to Understand? | Quanta Magazinequantamagazine.org
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- NLPContributionGraph -- Structuring Scholarly NLP Contributions in the Open Research Knowledge Graphncg-task.github.io
- Chain-of-Thought Promptinglearnprompting.org
- The Universe of Discourseblog.plover.com