✳flâneur — a map of the web's best reading
Building the heap: racking 30 petabytes of hard drives for pretraining | blog
si.inc · 3,297 words · saved by 17 readers
We built the heap, a 30 petabyte data storage cluster in downtown San Francisco, and spent under $500,000.
We built a storage cluster in downtown SF to store 90 million hours worth of video data. Why? We’re pretraining models to solve computer use. Compared to text LLMs like LLaMa-405B, which require ~60 TB of text data to train, videos are sufficiently large that we need 500 times more storage. Instead of paying the $12 million / yr it would cost to store all of this on AWS, we rented space from a colocation center in San Francisco to bring that cost down ~40x to $354k per year, including depreciation. Why Our use case for data is unique. Most cloud providers care highly about redundancy, availabi
Explore this link on the map →saved by
- Elizabeth Qiu
- Ratan Kaliani
- Hardeep Gambhir
- Asher P
- Neel Redkar
- Jianmin Chen
- Trevor Ze Trinh
- Lydia Nottingham
- Shiza Charania
- Vincent Cheng
- agniv sarkar
- Arjun Khandelwal
related reading
- How to Rack 30 Petabytes of Storage | blogsi.inc
- Century-Scale Storagelil.law.harvard.edu
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- Building and operating a pretty big storage system called S3 | All Things Distributedallthingsdistributed.com
- From bare metal to a 70B model: infrastructure set-up and scripts - Imbueimbue.com
- Roadmap: The AI data center stack - Bessemer Venture Partnersbvp.com
- To Boldly Go: The Case for Space Datacentersnewsletter.semianalysis.com
- IO devices and latency — PlanetScaleplanetscale.com
- Fast, scalable, clean, and cheap enoughoffgridai.us
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- Valhalla - Archiveteamwiki.archiveteam.org
- Will We Really Put Data Centers in Space?forethought.org