2401.02415.pdf
arxiv.org · 6,101 words · saved by 1 readers
N/A
LL A MA P RO: Progressive LLaMA with Block Expansion Chengyue Wu1,2 Yukang Gan2 Yixiao Ge2 * 3 Zeyu Lu Jiahao Wang1 Ye Feng4 Ying Shan2 Ping Luo1 1 2…
saved by
related reading
- Language Model Compositionarxiv.org
- Composer2.pdfcursor.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Large Language Diffusion Modelsarxiv.org
- 2310.10631arxiv.org
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Productizing Large Language Modelsblog.replit.com
- LLM Post-Training: A Deep Dive into Reasoning Large Language Modelsarxiv.org
- [2606.07527] Post-training is (Massive) Supervised Learningarxiv.org
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceblog.redwoodresearch.org
- LLM Resourcesforrestbicker.com