flâneur — a map of the web's best reading

A World of Verifiable Domains

seancai.com · 5,172 words · saved by 2 readers

In the mid-2010s, Scale AI pioneered the third-party data labeling market at a time when most enterprises were unsophisticated buyers of AI data. Back then, few companies had in-house machine learning (ML) expertise or applied ML teams, so they relied on external services to curate and annotate training datasets. Scale AI's initial focus was on labeling "hard tech" data (like images for computer vision in self-driving cars) and other straightforward annotation tasks. Cutting-edge ML was largely academic and not yet powerful enough for most commercial use cases, which meant early enterprise demand for data was limited and naive. This began to change with the transformer revolution (circa 2018–2020) and the rise of foundation models. As models like GPT-3 demonstrated astonishing capabilities from large-scale training, the need for ever-increasing "frontier" data surged. The most sophisticated early buyers of such data were Mag7 labs like OpenAI, DeepMind, and Anthropic who suddenly neede

Historical Context: Data Labelling from Scale AI days In the mid-2010s, Scale AI pioneered the third-party data labeling market at a time when most enterprises were unsophisticated buyers of AI data. Back then, few companies had in-house machine learning (ML) expertise or applied ML teams, so they relied on external services to curate and annotate training datasets. Scale AI's initial focus was on labeling "hard tech" data (like images for computer vision in self-driving cars) and other straightforward annotation tasks. Cutting-edge ML was largely academic and not yet powerful enough for most

Explore this link on the map →

saved by

related reading