flâneur

WebText Dataset | Papers With Code

paperswithcode.com · 1,676 words · saved by 1 readers

Some tasks are inferred based on the benchmarks list. The benchmarks section lists all benchmarks using a given dataset or any of its variants. We use variants to distinguish between results evaluated on slightly different versions of the same dataset. For example, ImageNet 32⨉32 and ImageNet 64⨉64 are variants of the ImageNet dataset. WebText is an internal OpenAI corpus created by scraping web pages with emphasis on document quality. The authors scraped all outbound links from Reddit which received at least 3 karma. The authors used the approach as a heuristic indicator for whether other users found the link interesting, educational, or just funny. WebText contains the text subset of these 45 million links. It consists of over 8 million documents for a total of 40 GB of text. All Wikipedia documents were removed from WebText since it is a common data source for other datasets.

new Get trending papers in your email inbox once a day! Get trending papers in your email inbox! Subscribe byAK and the research community Trending Papers Submitted by wanng AutoResearch: Insight In, Hallucination Out AutoResearch is a two-stage autonomous system that grounds research ideas through integrated generation and evidence-based execution to improve experimental reliability and measurable outcomes. EvoMap · Published on Aug 23, 2026 Submitted by wanng AutoResearch: Insight In, Hallucination Out AutoResearch is a two-stage autonomous system that grounds research ideas…

related reading