✳flâneur — a map of the web's best reading
Learning to Summarize with Human Feedback
openai.com · saved by 1 readers
We've applied reinforcement learning from human feedback to train language models that are better at summarization. Our models generate summaries that are better than summaries from 10x larger models trained only with supervised learning. Even though we train our models on the Reddit TL;DR dataset, the same models transfer
Explore this link on the map →