JustVugg/colibri: Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦 ·
github.com · 5,260 words · saved by 2 readers
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Website · Discord · English · 简体中文 · 繁體中文 · Italiano Tiny engine, immense model. Run frontier MoE models — 744B to 2.8T parameters — on consumer and heterogeneous hardware, in pure C with zero engine dependencies, by treating storage, RAM, and VRAM as a single inference hierarchy (AI memory multitiering). Eight families run today: GLM-5.2/5.3 (744B), GLM-5.3-Flash (321B, with vision), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash (284B), Qwen3.8-Flash-Next (125B + 51B n-gram), Qwen3.6 (35B-A3B) and OLMoE (7B) — one C file each, the same coli chat / coli serve / coli web front end. Full…
saved by
related reading
- Xiaomi MiMo, Explore and Lovemimo.xiaomi.com
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai
- Together AI | The AI Native Cloudtogether.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Composer2.pdfcursor.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Modal: High-performance AI infrastructuremodal.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- MatX: High-throughput chips for LLMsmatx.com
- Goodfire AIgoodfire.ai