The WARC Format
Web sites and web pages emerge and disappear from the world wide web every day. For the past ten years, memory organizations have tried to find the most appropriate ways to collect and keep track of this vast quantity of important material using web-scale tools such as web crawlers. A web crawler is a program that browses the web in an automated manner according to a set of policies; starting with a list of URLs, it saves each page identified by a URL, finds all the hyperlinks in the page (e. g. links to other pages, images, videos, scripting or style instructions, etc.), and adds them to the list of URLs to visit recursively. Storing and managing the billions of saved web page objects itself presents a challenge. At the same time, those same organizations have a rising need to archive large numbers of digital files not necessarily captured from the web (e.g., entire series of electronic journals, or data generated by environmental sensing equipment). A general requirement that appears
The WARC Format Standard About Title The WARC Format 1.0 standard Latest version See version 1.1 + community annotations Previous version None Issues View issues on GitHub Contents Scope Normative references Terms, definitions and acronyms Terms and definitions WARC record WARC record content block WARC record payload WARC record header WARC named fields WARC logical record Acronyms File and record model Named fields General WARC-Record-ID (mandatory) Content-Length (mandatory) WARC-Date (mandatory) WARC-Type (mandatory) Content-Type WARC-Concurrent-To WARC-Block-Digest WARC-Payload-Digest WAR
Explore this link on the map →saved by
related reading
- Curiuscurius.app
- Webpage archivearchive.ph
- Archiving URLs · Gwern.netgwern.net
- Redirecting to correct wddwbenjaminreinhardt.com
- bookbear express | Ava | Substackava.substack.com
- Wayback Machinemaa.org
- This Page is Designed to Last: A Manifesto for Preserving Content on the Webjeffhuang.com
- The Kool Aid Factory :: All Zineskoolaidfactory.com
- Wayback Machine - Wikipediaen.wikipedia.org
- Archiving websites – a new challenge to archives | International Council on Archives, Kuala Lumpur 2008web.archive.org
- curius graphcurius.hapi.foundation
- Wayback Machinemanuelohan.com