An Awesome List for getting started with web archiving
-
Updated
Aug 17, 2026
An Awesome List for getting started with web archiving
Wayback Machine API interface & a command-line tool
WARC + AI - Experimental Retrieval Augmented Generation Pipeline for Web Archive Collections.
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
A list of things related to software, literature, and other content for 🕣 Memento
Parse And Create Web ARChive (WARC) files with node.js
Various Jupyter notebooks about Common Crawl data
A dockerized, queued high fidelity web archiver based on Squidwarc
Awesome list dedicated to digital and data preservation tools, sources, services and so on.
Quick Cache and Archive search buttons
metawarc: a command-line tool for metadata extraction from files from WARC (Web ARChive)
A social media open post web archiving tool
Digital Preservation of HTTP in documentary heritage.
Decentralized web archiving
A tool for detecting viruses and NSFW material in WARC files
🗄 File-Based Reference Filing System.
A javascript for fighting link rot and content drift using link decoration and web archives.
Seeder - Czech webarchive curating tool and public site
Parser for WARC (aka WebArchive) files
To associate your repository with the webarchiving topic, visit your repo's landing page and select "manage topics."