Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) on X: "So my👇 thread about our papers investigating the verification and self-critiquing inabilities of GPT4 has apparently resonated with a lot of folks. Here is a quick response to several issues raised (either in replies or other quote-tweet threads). [Using one of 'em long "vanity…" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore Notifications Messages Lists Bookmarks Top Articles Communities Verified Orgs Profile More Post Eden Chan @onlychans1 Post See new posts Conversation Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) @rao2z So my thread about our papers investigating the verification and self-critiquing inabilities of GPT4 has apparently resonated with a lot of folks. Here is a quick response to several issues raised (either in replies or other quote-tweet threads). [Using one of 'em long "vanity tweets" ] 1. I am not an LLM luddite--I think they are amazing "idea generators" (in either language or code form). They just can't do their own planning/reasoning with any guarantees. So they are best used in LLM-Modulo settings (with either a sound reasoner or an expert human in the loop). Self-critiquing needs verification, which is a form of reasoning (..and so I was surprised at all the claims about LLM self-critiquin
Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) @rao2z So my👇 thread about our papers investigating the verification and self-critiquing inabilities of GPT4 has apparently resonated with a lot of folks. Here is a quick response to several issues raised (either in replies or other quote-tweet threads). [Using one of 'em long "vanity tweets" 😛] 1. I am not an LLM luddite--I think they are amazing "idea generators" (in either language or code form). They just can't do their own planning/reasoning with any guarantees. So they are best used in LLM-Modulo settings (with either a sound reasoner or an ex
Explore this link on the map →related reading
- [2310.08118] Can Large Language Models Really Improve by Self-critiquing Their Own Plans?arxiv.org
- gpt-4.pdfcdn.openai.com
- Can LLMs Critique and Iterate on Their Own Outputs? | Eric Jangevjang.com
- GPT-4openai.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org
- The bitter lesson of LLM evalsparsed.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- LLMs Can Self-Improvearxiv.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- gpt-4-system-card.pdfcdn.openai.com
- Self-Verification, The Key to AIincompleteideas.net