✳flâneur — a map of the web's best reading
Vanilla GPT-3 quality from an open source model on a single machine: GLM-130B - The Full Stack
fullstackdeeplearning.com · 2,130 words · saved by 1 readers
Notes from deploying GLM-130B, a large language model from Tsinghua KEG
Vanilla GPT-3 quality from an open source model on a single machine: GLM-130B By Charles Frye . tl;dr GLM-130B is a GPT-3-scale and quality language model that can run on a single 8xA100 node without too much pain. Kudos to Tang Jie and the Tsinghua KEG team for open-sourcing a big, powerful model and the tricks it takes to make it run on reasonable hardware. Results are roughly what you might expect after reading the paper : similar to the original GPT-3 175B, worse than the InstructGPTs. I've really been spoiled by OpenAI's latest models: easier to prompt, higher quality generations. And It'
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Mosaic LLMs: GPT-3 quality formosaicml.com
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Inference characteristics of Llama · Cursorcursor.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GPT-3 - Wikipediaen.wikipedia.org
- More Efficient In-Context Learning with GLaMblog.research.google
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- 2506.17298arxiv.org