EXPERIMENT 04 / CRAWLER HONEYPOT

Signal Vault.

An open archive. A trail of links.
A place to see what follows.

Twelve research documents invite web bots and AI retrieval agents into a finite crawl path. A reserved download records visits to a robots-disallowed canary.

Public research • 12 archive pages • No sign-in required

THE PUBLIC COLLECTION

Follow the research trail

Download manifest →
01
Identity

AI crawler identity: claims and evidence

How to interpret a crawler name without mistaking a declared identity for proof.

02
Datasets

Open crawler telemetry: JSON and CSV exports

Public aggregate data for studying automated requests to the observatory.

03
Discovery

Machine-readable crawl manifests and resource indexes

A compact inventory of the archive, public datasets, and research notes.

04
Protocol

Robots exclusion and a reserved archive canary

A decoy resource for observing requests to a robots-disallowed path.

05
Discovery

Canonical URLs and finite archive boundaries

Why one permanent address per document keeps a crawl interpretable.

06
Retrieval

HTML link discovery without JavaScript

Archive navigation is present in the first server-rendered response.

07
Discovery

Search discovery with sitemaps and IndexNow

How the observatory announces public pages when they change.

08
Retrieval

AI agent retrieval and task context

A fetched research page does not reveal the task that brought a client here.

09
Measurement

Privacy-minimized request observations

The experiment records route and classification fields, not visitor profiles.

10
Measurement

Request counts, sampling, and missing observations

How to read a quiet dashboard without inventing crawler activity.

11
Measurement

Reproducible crawl graphs and archive coverage

A small, inspectable corpus for comparing fetch patterns over time.

12
End of archive

Archive boundary and public research resources

The final document in the Signal Vault crawl path.

RESERVED RESEARCH CANARY

The extra door

The full-export link is an intentional decoy. Automated crawlers are asked to leave it alone in robots.txt. A request records a canary hit; the response is a small fixture containing no private data.

Human clicks and user-directed fetches can also reach it, so a hit alone does not establish abuse.

Full crawl export (canary) →