[2310.08118] Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2310.08118] Can Large Language Models Really Improve by Self-critiquing Their Own Plans? Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Artificial Intelligence arXiv:2310.08118 (cs) [Submitted on 12 Oct 2023] Title: Can Large Language Models Really Improve by Self-critiquing Their Own Plans? Authors: Karthik Valmeekam , Matthew Marquez , Subbarao Kambhampati View a PDF of the paper titled Can Large Language Models Really Improve by Self-critiquing Their Own Plans?, by Karthik Val
Explore this link on the map →related reading
- Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) on X: "So my👇 thread about our papers investigating the verification and self-critiquing inabilities of GPT4 has apparently resonated with a lot of folks. Here is a quick response to several issuetwitter.com
- [2605.20873] PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Modelsarxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Can LLMs Critique and Iterate on Their Own Outputs? | Eric Jangevjang.com
- The bitter lesson of LLM evalsparsed.com
- LLMs Can Self-Improvearxiv.org
- Language Models can Solve Computer Tasksarxiv.org
- pdfopenreview.net
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org
- Large Language Model: world models or surface statistics?thegradient.pub
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io