flâneur — a map of the web's best reading

Vanilla GPT-3 quality from an open source model on a single machine: GLM-130B - The Full Stack

fullstackdeeplearning.com · 2,130 words · saved by 1 readers

Notes from deploying GLM-130B, a large language model from Tsinghua KEG

Vanilla GPT-3 quality from an open source model on a single machine: GLM-130B By Charles Frye . tl;dr GLM-130B is a GPT-3-scale and quality language model that can run on a single 8xA100 node without too much pain. Kudos to Tang Jie and the Tsinghua KEG team for open-sourcing a big, powerful model and the tricks it takes to make it run on reasonable hardware. Results are roughly what you might expect after reading the paper : similar to the original GPT-3 175B, worse than the InstructGPTs. I've really been spoiled by OpenAI's latest models: easier to prompt, higher quality generations. And It'

Explore this link on the map →

related reading