✳flâneur — a map of the web's best reading
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
beichenhuang.github.io · saved by 1 readers
N/A
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- Neuronpedianeuronpedia.org
- Xiaomi MiMo, Explore and Lovemimo.xiaomi.com
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- MatX: High-throughput chips for LLMsmatx.com
- Future leakage in block-quantized attention | MatXmatx.com
- GitHub - EveryInc/compound-engineering-plugin: Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more · GitHubgithub.com