RedPajama, a project to create leading open-source models, starts by reproducing LLaMA training dataset of over 1.2 trillion tokens — TOGETHER
RedPajama is a project to create a set of leading, fully open-source models. Today, we are excited to announce the completion of the first step of this project: the reproduction of the LLaMA training dataset of over 1.2 trillion tokens.
Foundation models such as GPT-4 have driven rapid improvement in AI. However, the most powerful models are closed commercial models or only partially open. RedPajama is a project to create a set of leading, fully open-source models. Today, we are excited to announce the completion of the first step of this project: the reproduction of the LLaMA training dataset of over 1.2 trillion tokens. The most capable foundation models today are closed behind commercial APIs, which limits research, customization, and their use with sensitive data. Fully open-source models hold the promise of removing thes
Explore this link on the map →saved by
related reading
- Google "We Have No Moat, And Neither Does OpenAI"semianalysis.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Olmo from Ai2allenai.org
- Catching up on the weird world of LLMssimonwillison.net
- Large Language Diffusion Modelsarxiv.org
- Accelerating LLaMA with Fabric: A Comprehensive Guide to Training and Fine-Tuning LLaMA - Lightning AIlightning.ai
- Stanford CRFMcrfm.stanford.edu
- Introducing Marin: An Open Lab for Building Foundation Models | Marinmarin.community
- Gemma 2: Improving Open Language Models at a Practical Sizearxiv.org
- 2310.10631arxiv.org