✳flâneur — a map of the web's best reading
Llama 2 follow-up: too much RLHF, GPU sizing, technical details
interconnects.ai · 3,062 words · saved by 1 readers
The community reaction to Llama 2 and all of the things that I didn't get to in the first issue.
Llama 2 follow-up: too much RLHF, GPU sizing, technical details The community reaction to Llama 2 and all of the things that I didn't get to in the first issue. Nathan Lambert Jul 21, 2023 22 Share Following all of the Llama 2 news in the last few days would've been beyond a full-time job. The information networks truly were overflowing with takes, experiments, and updates. It'll still be like this for another week at least, but there are already some crucial points. In this post, I will clarify a couple of corrections I made to the original post on all things Llama 2, and then I will continue
Explore this link on the map →related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- rlhfbook.com/book.pdfrlhfbook.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- LLM Training: RLHF and Its Alternativesmagazine.sebastianraschka.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- Gemma 2: Improving Open Language Models at a Practical Sizearxiv.org
- How RLHF actually works - by Nathan Lambertinterconnects.ai