✳flâneur — a map of the web's best reading
Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- Language Models are Unsupervised Multitask Learnerscdn.openai.com
- Efficiently Scaling Transformer Inferencearxiv.org
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- Neuronpedianeuronpedia.org
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- GitHub - ml5js/training-styletransfer: Style Transfer training and using the model in ml5js · GitHubgithub.com
- TensorTonic | Learn ML through codetensortonic.com