flâneur — a map of the web's best reading

AI #23: Fundamental Problems with RLHF - by Zvi Mowshowitz

thezvi.substack.com · saved by 1 readers

After several jam-packed weeks, things slowed down to allow everyone to focus on the potential room temperature superconductor, check Polymarket to see how likely it is we are so back and bet real money, or Manifold for chats and better graphs and easier but much smaller trading. The main thing I would highlight this week are an excellent paper laying out many of the fundamental difficulties with RLHF, and a systematic new exploit of current LLMs that seems to reliably defeat RLHF. I’d also note that GPT-4 fine tuning is confirmed to be coming. That should be fun. Introduction. Table of Contents. Language Models Offer Mundane Utility. Here’s what you’re going to do. Language Models Don’t Offer Mundane Utility. Universal attacks on LLMs. Fun With Image Generation. Videos might be a while. Deepfaketown and Botpocalypse Soon. An example of doing it right. They Took Our Jobs. What, me worry? Get Involved. If you share more opportunities in comments I’ll include next week. Introducing. A bi

After several jam-packed weeks, things slowed down to allow everyone to focus on the potential room temperature superconductor, check Polymarket to see how likely it is we are so back and bet real money, or Manifold for chats and better graphs and easier but much smaller trading. The main thing I would highlight this week are an excellent paper laying out many of the fundamental difficulties with RLHF, and a systematic new exploit of current LLMs that seems to reliably defeat RLHF. I’d also note that GPT-4 fine tuning is confirmed to be coming. That should be fun. Introduction. Table of Conten

Explore this link on the map →