✳flâneur — a map of the web's best reading
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- Recent Advances in Language Model Fine-tuningruder.io
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org
- GLaM: Efficient Scaling of Language Models with Mixture-of-Expertsarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- Language Models are Unsupervised Multitask Learnerscdn.openai.com
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine Translation - ACL Anthologyaclanthology.org
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Llama 2 · Hugging Facehuggingface.co