DeepSeek-R1 and exploring DeepSeek-R1-Distill-Llama-8B
DeepSeek are the Chinese AI lab who dropped the best currently available open weights LLM on Christmas day, DeepSeek v3. That model was trained in part using their unreleased R1 …
DeepSeek-R1 and exploring DeepSeek-R1-Distill-Llama-8B Simon Willison’s Weblog Subscribe Sponsored by: Atlassian - Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more DeepSeek-R1 and exploring DeepSeek-R1-Distill-Llama-8B 20th January 2025 DeepSeek are the Chinese AI lab who dropped the best currently available open weights LLM on Christmas day , DeepSeek v3. That model was trained in part using their unreleased R1 “reasoning” model. Today they’
related reading
- deepseek-r1ollama.com
- DeepSeek-R1arxiv.org
- The Complete Guide to DeepSeek Models: V3, R1, V4 and Beyondbentoml.com
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com
- Deepseek: The Quiet Giant Leading China’s AI Racechinatalk.media
- DeepSeek: The View from Chinachinatalk.media
- Dario Amodei — On DeepSeek and Export Controlsdarioamodei.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- The Illustrated DeepSeek-R1newsletter.languagemodels.co
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- DeepSeek and the Day Before New Year'swheremachinesthink.substack.com
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsarxiv.org