RedPajama, a project to create leading open-source models, starts by reproducing LLaMA training dataset of over 1.2 trillion tokens — TOGETHER
RedPajama is a project to create a set of leading, fully open-source models. Today, we are excited to announce the completion of the first step of this project: the reproduction of the LLaMA training dataset of over 1.2 trillion tokens.
Foundation models such as GPT-4 have driven rapid improvement in AI. However, the most powerful models are closed commercial models or only partially open. RedPajama is a project to create a set of leading, fully open-source models. Today, we are excited to announce the completion of the first step of this project: the reproduction of the LLaMA training dataset of over 1.2 trillion tokens. The most capable foundation models today are closed behind commercial APIs, which limits research, customization, and their use with sensitive data. Fully open-source models hold the promise of removing thes
Explore this link on the map →saved by
related reading
- Google "We Have No Moat, And Neither Does OpenAI"semianalysis.com
- GitHub - openlm-research/open_llama: OpenLLaMA, a permissively licensed open source reproduction of Meta AI’s LLaMA 7B trained on the RedPajama datasetgithub.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- GitHub - eugeneyan/open-llms: 📋 A list of open LLMs available for commercial use.github.com
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- Olmo from Ai2allenai.org
- Catching up on the weird world of LLMssimonwillison.net
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Large Language Diffusion Modelsarxiv.org
- Introducing Marin: An Open Lab for Building Foundation Models | Marinmarin.community
- Accelerating LLaMA with Fabric: A Comprehensive Guide to Training and Fine-Tuning LLaMA - Lightning AIlightning.ai
- Stanford CRFMcrfm.stanford.edu