Deft Fixing LLM Writing with Distribution Fine-Tuning
Rosmine's research paper on Distribution Fine-Tuning: the measurement stack, SFT failure modes, DFT results, model comparisons, methods, data, and limitations.
Slop. It's not just annoying — it's exhausting. You're absolutely right to be annoyed by it, and in this blog I will delve into a solution. You've probably noticed most models have their favorite words or phrases they overuse, like "—", "it's not X, it's Y", or "delve". Before investigating the solution, I first address the metrics I use to measure output quality. Instead of measuring "quality" itself, which is not well defined, I measure similarity to human writing samples. Metrics N-Gram Token Distribution L2 Distance This metric captures word choice similarity, and is useful for…
saved by
related reading
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Writing for LLMs So They Listen · Gwern.netgwern.net
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Modifying Large Language Model Post-Training for Diverse Creative Writingarxiv.org
- Productizing Large Language Modelsblog.replit.com
- The bitter lesson of LLM evalsparsed.com
- PostTrainBenchposttrainbench.com
- Large Language Diffusion Modelsarxiv.org
- dim sum paperarxiv.org
- Fine-tuning a LLM on my blog posts | Didier Lopesdidierlopes.com
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generationpub.sakana.ai