Skip to content
Strategic SolutionsLegal research · evidence discovery2026Client engagement

Turning pages that no longer exist into citable evidence

The problem

Proving what an organisation advertised years ago means working from pages that have since been edited or deleted. Done by hand it is one person opening archived snapshots one at a time — slow, incomplete, and impossible to demonstrate as systematic afterwards.

The insight

The value is in the coverage, not the best example

A manual trawl finds the single most damning page and stops. That is the weakest possible form of the argument, because the response is always that one page was an outlier. What holds up is a sweep that can state its own scope: every archived capture in the date range, every page matching the criteria, all put through the same classifier, with anything found carrying the archive URL and capture timestamp that lets a reader verify it independently. The run is checkpointed, so it is resumable and repeatable rather than a one-off nobody can reproduce.

What we built

  • Enumeration of every archived capture for a domain and date range through the Wayback CDX index, with retry, backoff, and rate limiting tuned to what the archive actually tolerates
  • A relevance gate so a page only produces findings if it genuinely concerns the subject in scope
  • A tiered claim taxonomy — specific and falsifiable claims separated from merely relevant ones and from supporting context — so reviewer time goes where the evidence is strongest
  • Output as a reviewable spreadsheet where every row carries its archive URL and capture date as the citation
  • Checkpointing, so a long run resumes rather than restarts and the same scope can be re-run later

Have something with a domain problem underneath it?

Those are the ones worth talking about. Describe it and you will get a straight answer about what it would take.

Start a conversation