Can LLMs Critique and Iterate on Their Own Outputs? | Eric Jang
Avi Singh told me yesterday about a recent arxiv preprint, Reflexion, that proposes the following idea: use a LLM to examine whether the output of another generative model is “on the right track” during generation. According to the paper, “the reflection loop aims to help the agent correct common cases of hallucination and inefficiency through trial and error.” Reflexion extends the ReAct architecture to predict whether the agent should stop generating, pause, and reflect on its entire generated trajectory. Should the agent decide to self-reflect with a LLM, it restarts the generation process with its LLM critique of its past trajectory loaded in-context. The paper is evaluated against text-based decision-making benchmarks like AlfWorld, HotPotQA, and WebShop. If it helps the intuition, you can think of this like someone sending you a text, then quickly “unsending” it and then sending a new one.
Avi Singh told me yesterday about a recent arxiv preprint, Reflexion , that proposes the following idea: use a LLM to examine whether the output of another generative model is "on the right track" during generation. According to the paper, "the reflection loop aims to help the agent correct common cases of hallucination and inefficiency through trial and error." Reflexion extends the ReAct architecture to predict whether the agent should stop generating, pause, and reflect on its entire generated trajectory. Should the agent decide to self-reflect with a LLM, it restarts the generation process
Explore this link on the map →related reading
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- 1b44b878bb782e6954cd888628510e90-Paper-Conference.pdfproceedings.neurips.cc
- Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) on X: "So my👇 thread about our papers investigating the verification and self-critiquing inabilities of GPT4 has apparently resonated with a lot of folks. Here is a quick response to several issuetwitter.com
- [2310.08118] Can Large Language Models Really Improve by Self-critiquing Their Own Plans?arxiv.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- Writing for LLMs So They Listen · Gwern.netgwern.net
- Meta-Prompt: A Simple Self-Improving Language Agentnoahgoodman.substack.com
- crawshaw - 2025-01-06crawshaw.io
- [2310.11511] Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflectionarxiv.org