Stanford Alpaca Model Release
Instruction-following models such as GPT-3.5 (text-davinci-003), ChatGPT, Claude, and Bing Chat have become increasingly powerful. Many users now interact with these models regularly and even use them for work. However, despite their widespread deployment, instruction-following models still have many deficiencies: they can generate false information, propagate social stereotypes, and produce toxic language. To make maximum progress on addressing these pressing problems, it is important for the academic community to engage. Unfortunately, doing research on instruction-following models in academia has been difficult, as there is no easily accessible model that comes close in capabilities to closed-source models such as OpenAI’s text-davinci-003. We are releasing our findings about an instruction-following language model, dubbed Alpaca, which is fine-tuned from Meta’s LLaMA 7B model. We train the Alpaca model on 52K instruction-following demonstrations generated in the style of self-instr
Stanford CRFM Alpaca: A Strong, Replicable Instruction-Following Model Authors: Rohan Taori* and Ishaan Gulrajani* and Tianyi Zhang* and Yann Dubois* and Xuechen Li* and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto We introduce Alpaca 7B , a model fine-tuned from the LLaMA 7B model on 52K instruction-following demonstrations. On our preliminary evaluation of single-turn instruction following, Alpaca behaves qualitatively similarly to OpenAI’s text-davinci-003, while being surprisingly small and easy/cheap to reproduce (<600$). Checkout our code release on GitHub . Update: The pub
Explore this link on the map →saved by
related reading
- gpt-4.pdfcdn.openai.com
- Alignment faking in large language modelsarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- 2506.17298arxiv.org
- GPT-3 - Wikipediaen.wikipedia.org
- Large Language Diffusion Modelsarxiv.org
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Unsupervised Elicitationalignment.anthropic.com
- 2405.01470arxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com