Skip to content

How Data Is Collected

Logan Besecker edited this page Oct 1, 2026 · 1 revision

How Data Is Collected

Three scheduled processes keep listings current. None of them is triggered by a webhook — see why polling and not webhooks.

Process Reads How often
Official registry sync each listing's server.json every 6 hours (incremental), full pass weekly
Prober the tools, prompts and resources a live server reports each remote server weekly
Document fetcher llms.txt and AGENTS.md each URL weekly

Official registry sync

Copies the official MCP Registry's catalogue, asking only for what changed since the last run. A listing upstream marks deleted is removed here; a full pass that sees far fewer servers than expected is treated as an upstream problem and deletes nothing. Locally: mix registry.sync_official.

Prober

Connects to each remote listing, performs the MCP handshake and calls tools/list, prompts/list and resources/list.

  • Nothing is executed. Packaged (stdio) listings are never probed: finding their tools would mean running a stranger's code.
  • A failed probe never erases data. Over a third of remote endpoints require auth. An auth-gated server keeps whatever tools its publisher declared — a locked door is not evidence of an empty room.
  • Lists are capped at 500 items, and a name over 2,000 characters is dropped rather than truncated, since half a URI is a wrong URI.
  • It is deliberately slow — a few hundred endpoints per five minutes — so that it looks like a visitor rather than a scan.

Locally: mix probe.tools.

Document fetcher

Fetches each listing's llms.txt (from its website, or the domain its namespace was verified for) and AGENTS.md (from the root of its GitHub repository). One URL that many listings share is fetched once.

Because website_url is whatever a publisher typed, every fetch is guarded:

  • https only;
  • the host must resolve to a public address, checked again at every redirect;
  • bodies stream with a 256 KB cap;
  • an HTML or JSON page at /llms.txt counts as no file, because many sites answer every path with a 200.

Unchanged files are asked for conditionally and answer 304 with no body.

Logos

Chosen in this order, from what a publisher has already put on the public record:

  1. The icon declared in server.json — preferring SVG, then the smallest image still sharp at double size.
  2. The GitHub avatar of an io.github.* namespace owner, whose identity the official registry verified.
  3. The avatar of the GitHub repository's owner.
  4. Otherwise the listing's monogram.

Logos are never guessed from a name or scraped from a website: a wrong logo implies an affiliation that isn't there. If an image fails to load it removes itself and the monogram shows instead.

Clone this wiki locally