A toolkit for CDX indices such as Common Crawl and the Internet Archive's Wayback Machine
-
Updated
Sep 7, 2026 - Python
A toolkit for CDX indices such as Common Crawl and the Internet Archive's Wayback Machine
Summarize web archive capture index (CDX) files.
Python tools to retrieve text from CommonCrawl WARC files based on cdx index.
Fetch and normalize parameterized URLs from the Wayback CDX API (OSINT, inspired by ParamSpider).
Enables Mac roundtrip editing for ChemDraw scheme-contaning PowerPoints made in Windows
Zero-install, server-preserving web archiver and ISO 28500 forensic engine for offline replication and digital preservation.
View cdx and warc files, caching them locally as needed
A web crawler built from scratch in Python, one concept per stage: frontier, robots.txt, politeness, async workers, dedup, rendering, freshness, a distributed frontier, WARC archives, a CDX index and replay. Thirteen stages, 222 tests, verified against a corpus that declares its own ground truth.
The solution to extend the deadline for the virtual machines on CDX.
To associate your repository with the cdx topic, visit your repo's landing page and select "manage topics."