User-Owned Foundation Models
Today, GPT-3 is trained on publicly scraped data from the internet. What if it were also trained on private data contributed directly by users, allowing for access to datasets like all Facebook messages, Instagram posts and Gmail emails that are usually siloed within a platform? I originally presented this memo at Vana all-hands. I'm including this slack message for context: GPT-3 is trained on these text datasets: Common Crawl (publicly scraped webpages), WebText2 (publicly available URLs posted on reddit with more than 3 upvotes), Books1 and Books2 datasets (unclear which exact book dataset used), and Wikipedia. Source: Wikipedia GPT-3 page But people estimate <0.1% of the internet is publicly scrapable. The rest of the internet requires permissions or a sign-in to access, so it’s not used as training data today. I would like to explore including permissioned data in training models by directly asking users to contribute their platform data. Proposed data sources: From a representati
Today, GPT-3 is trained on publicly scraped data from the internet. What if it were also trained on private data contributed directly by users, allowing for access to datasets like all Facebook messages, Instagram posts and Gmail emails that are usually siloed within a platform? I originally presented this memo at Vana all-hands. I'm including this slack message for context: GPT-3 is trained on these text datasets: Common Crawl (publicly scraped webpages), WebText2 (publicly available URLs posted on reddit with more than 3 upvotes), Books1 and Books2 datasets (unclear which exact book dataset
Explore this link on the map →