Crawled — a first-party ledger of which machines read this website, which pages they take, and how long after publication they arrive

Crawled

Every machine that has read this website, what it took, and how long after the page existed it turned up.

Fetches recorded
206since 2026-08-30 11:45Z
By the answer layer
130training, AI search, live assistant fetches
By search engines
76the control column
Distinct machines
11named agents only
Pages reached
26/27read at least once by the answer layer

The machines

Split by what each operator says its crawler is for. The purpose matters more than the count: a page fetched by ChatGPT-User seconds after somebody asked a question is a different event from the same page being swept into a training corpus.

AgentOperatorLayerWhat it is forFetchesLast seen
BytespiderByteDanceanswerBuilding a training corpus572026-09-01 06:06Z
ClaudeBotAnthropicanswerBuilding a training corpus352026-09-01 07:59Z
GPTBotOpenAIanswerBuilding a training corpus152026-09-01 00:49Z
OAI-SearchBotOpenAIanswerIndexing for AI answers142026-09-01 07:19Z
AmazonbotAmazonanswerIndexing for AI answers52026-09-01 07:00Z
ChatGPT-UserOpenAIanswerFetching live to answer somebody42026-08-31 16:51Z
bingbotMicrosoftsearchIndexing for a search engine482026-09-01 06:41Z
PetalBotHuaweisearchIndexing for a search engine112026-09-01 07:34Z
DuckDuckBotDuckDuckGosearchIndexing for a search engine92026-09-01 07:08Z
ApplebotApplesearchIndexing for a search engine42026-09-01 07:57Z
GooglebotGooglesearchIndexing for a search engine42026-08-31 15:37Z

First contact

How long each experiment’s page existed before a machine read it. Publication time is the launch recorded in the control plane, not the day somebody decided to announce it. A dash means it has not happened.

PageArmPublishedFirst read by the answer layerAfterFetches
Neededbrief2026-08-31ClaudeBot · 2026-09-01 07:24Z15.5 h2
Slopholdout2026-08-30GPTBot · 2026-08-31 02:09Z8.5 h3
Donebrief2026-08-30GPTBot · 2026-08-31 02:09Z9.2 h3
Up Peakholdout2026-08-30GPTBot · 2026-08-31 02:09Z9.7 h4
Crawledbrief2026-08-30GPTBot · 2026-08-31 02:09Z12.2 h4
Out of Tenholdout2026-08-30GPTBot · 2026-08-31 02:09Z17.2 h3
Perfect Scorebrief2026-08-29GPTBot · 2026-08-31 02:09Z40.2 h2
Insulaholdout2026-08-28ClaudeBot · 2026-09-01 07:24Z4 d2
Exactlybrief2026-08-26ClaudeBot · 2026-09-01 07:24Z6 d2
Spentholdout2026-08-25ClaudeBot · 2026-09-01 07:25Z7 d2
Sievebrief2026-08-24Bytespider · 2026-08-30 12:16Z6 d5
Hold the Noteholdout2026-08-22Bytespider · 2026-08-30 17:26Z8 d2
Signal Weatherbrief2026-05-20Bytespider · 2026-08-31 10:45Z103 d3
AI-deasbrief2026-05-11Bytespider · 2026-08-31 07:10Z112 d2
Built From a Phoneholdout2026-05-04ClaudeBot · 2026-09-01 07:26Z120 d3
Micro-Breakholdout2026-05-04Bytespider · 2026-08-31 08:31Z119 d4
Code of Lawsholdout2026-05-04Bytespider · 2026-08-31 01:00Z118 d3
LLM Price Volatilitybrief2026-05-04Bytespider · 2026-08-30 12:34Z118 d4
Voidleholdout2026-05-04Bytespider · 2026-08-30 18:54Z118 d4
Velocity Cardsbrief2026-05-04Bytespider · 2026-08-30 18:13Z118 d2
Instruction Elasticity Indexholdout2026-05-04ClaudeBot · 2026-09-01 07:25Z120 d1
The Replacement Indexbrief2026-05-04Bytespider · 2026-08-30 22:19Z119 d2
Signal from the Voidholdout2026-05-04Bytespider · 2026-08-30 14:28Z118 d2
Subscribe to Nothingbrief2026-05-04Bytespider · 2026-08-31 04:51Z119 d3
Make Interestingholdout2026-05-04Bytespider · 2026-08-31 00:49Z118 d5
On Mebrief2026-08-30not yet1
Mutation Observerbrief2026-05-04ClaudeBot · 2026-09-01 07:24Z120 d1

Brief against holdout

The answer layer has reached 93% of the pages that publish a brief and 100% of the pages that do not. Median time from publishing to first machine read: 2473.8h with a brief, 2837.5h without. This is a difference between two small groups on one small site, not a finding about the web.

Both arms have passed 3 read pages, so the columns are shown side by side. They are still two small groups on one small site.

ArmPagesReadMedian time to first read
Machine-readable brief1413103 d
HTML only (holdout)1313118 d

Does anything read the machine-readable surfaces?

/llms.txt and the per-experiment briefs exist because a great deal of advice says they should. Whether any machine has ever fetched one is a question with an answer, and this is it.

SurfaceFetchesRead by
/llms.txt1GPTBot

The tail

The last 120 fetches, newest first. Nothing here is live-updating; reload it.

  • ClaudeBot/experiments/crawled
  • ClaudeBot/experiments/done
  • ClaudeBot/experiments/insula
  • ClaudeBot/experiments/needed
  • Applebot/robots.txt
  • PetalBot/robots.txt
  • ClaudeBot/e/out-of-ten
  • ClaudeBot/e/signal-weather
  • ClaudeBot/
  • ClaudeBot/e/signal-from-the-void
  • ClaudeBot/e/llm-price-volatility
  • ClaudeBot/e/built-from-a-phone
  • ClaudeBot/e/velocity-cards
  • ClaudeBot/e/the-replacement-index
  • ClaudeBot/e/code-of-laws
  • ClaudeBot/e/subscribe-to-nothing
  • ClaudeBot/e/ai-deas
  • ClaudeBot/e/instruction-elasticity
  • ClaudeBot/e/up-peak
  • ClaudeBot/e/hold-the-note
  • ClaudeBot/e/perfect-score
  • ClaudeBot/e/slop
  • ClaudeBot/e/sieve
  • ClaudeBot/e/spent
  • ClaudeBot/e/micro-break
  • ClaudeBot/e/mutation-observer
  • ClaudeBot/e/make-interesting
  • ClaudeBot/e/exactly
  • ClaudeBot/e/crawled
  • ClaudeBot/e/done
  • ClaudeBot/experiments/exactly
  • ClaudeBot/e/insula
  • ClaudeBot/e/voidle
  • ClaudeBot/e/needed
  • ClaudeBot/robots.txt
  • OAI-SearchBot/robots.txt
  • DuckDuckBot/
  • Amazonbot/experiments/needed
  • ClaudeBot/
  • ClaudeBot/robots.txt
  • bingbot/
  • Bytespider/robots.txt
  • Bytespider/robots.txt
  • bingbot/e/code-of-laws
  • PetalBot/sitemap.xml
  • bingbot/e/llm-price-volatility
  • bingbot/sitemap.xml
  • bingbot/robots.txt
  • bingbot/robots.txt
  • bingbot/e/sieve
  • Bytespider/experiments/velocity-cards
  • PetalBot/e/crawled
  • PetalBot/experiments/insula
  • bingbot/e/on-me
  • bingbot/e/needed
  • bingbot/sitemap.xml
  • Bytespider/robots.txt
  • PetalBot/experiments/exactly
  • Bytespider/robots.txt
  • GPTBot/sitemap.xml
  • OAI-SearchBot/robots.txt
  • bingbot/experiments/sieve
  • bingbot/e/built-from-a-phone
  • PetalBot/robots.txt
  • bingbot/e/make-interesting
  • bingbot/experiments/signal-from-the-void
  • OAI-SearchBot/
  • OAI-SearchBot/
  • OAI-SearchBot/experiments/signal-from-the-void
  • OAI-SearchBot/robots.txt
  • OAI-SearchBot/robots.txt
  • bingbot/experiments/code-of-laws
  • Amazonbot/
  • OAI-SearchBot/robots.txt
  • bingbot/e/micro-break
  • bingbot/e/voidle
  • OAI-SearchBot/robots.txt
  • Bytespider/robots.txt
  • Applebot/robots.txt
  • Bytespider/robots.txt
  • ChatGPT-User/
  • ChatGPT-User/
  • Applebot/robots.txt
  • Googlebot/robots.txt
  • DuckDuckBot/
  • PetalBot/e/up-peak
  • PetalBot/e/insula
  • DuckDuckBot/
  • Amazonbot/experiments/i-built-a-game-from-my-phone
  • DuckDuckBot/
  • Bytespider/e/signal-weather
  • DuckDuckBot/
  • ChatGPT-User/
  • bingbot/e/sieve
  • Googlebot/robots.txt
  • bingbot/e/signal-weather
  • Bytespider/e/micro-break
  • PetalBot/e/exactly
  • Bytespider/e/ai-deas
  • bingbot/robots.txt
  • bingbot/experiments/sieve
  • bingbot/
  • ChatGPT-User/
  • Bytespider/robots.txt
  • bingbot/e/make-interesting
  • PetalBot/e/spent
  • Bytespider/e/subscribe-to-nothing
  • bingbot/e/subscribe-to-nothing
  • bingbot/experiments/signal-from-the-void
  • bingbot/experiments/code-of-laws
  • PetalBot/robots.txt
  • DuckDuckBot/
  • GPTBot/llms.txt
  • GPTBot/experiments/done
  • GPTBot/experiments/slop
  • GPTBot/experiments/perfect-score
  • GPTBot/experiments/up-peak
  • GPTBot/experiments/crawled
  • GPTBot/experiments/out-of-ten
  • GPTBot/e/perfect-score
Take the whole ledger as JSON, CC BY 4.0 →
What this is, if you want it after the fact

For thirty years the question was where a page ranked. It is becoming a different question: whether a machine reads the page at all, and whether what it takes away is what the page actually said. Almost everything written about that second question is advice. Very little of it is measurement, because measuring it requires a site, a control group, and a willingness to publish the result when it is boring.

This is the smallest honest instrument for it. Every request to a tracked page is matched against a table of known crawlers, and one row is written: which agent, which path, when. From that come three things that are hard to get any other way — which operators actually fetch a small site, how long a page waits between being published and being read, and whether the answer layer behaves at all differently from the search engines crawling the same pages in the same window.

Half the experiment pages publish a machine-readable brief. Half do not, on purpose, by a rule fixed before there was anything to see. That holdout is the only reason any sentence on this page can be a finding rather than an anecdote, and it costs something real: half the portfolio is deliberately published in the worse condition, if the worse condition turns out to be worse.

What this deliberately does not measure

A fetch is not a citation. A citation is not a recommendation. A recommendation is not an agent choosing to use something. Those are four different events and only the first is observable from a server log.

It would be easy to invent scores for the other three — a retrievability index, a citation-worthiness rating, an agent-readiness percentage out of a hundred. They would look authoritative, they would be screenshot-friendly, and they would be fabricated, because there is no ground truth on this site to check them against. Anything on this page that is a number came from a request that actually arrived.

The honest version of the missing half is a second instrument that asks real models real questions on a schedule and records whether this site comes back — mention, citation, and whether the description is accurate. That is measurable, it is not cheap, and it is not built yet. It is the obvious next thing, and it stays unbuilt rather than being simulated in the meantime.

The method, in full
  • Where the count happens. Edge middleware, before the cache. A story page is served from the edge for sixty seconds at a time, so a crawler fetch never reaches a route handler — anywhere later in the stack, most of this traffic is invisible.
  • Who is counted. Only agents in a published table, matched on a substring of the user-agent. Unrecognised agents are not counted at all: an unknown-bot bucket would be mostly scrapers and uptime checks, and putting that beside GPTBot would flatter the result.
  • What is stored. Agent, path, timestamp. No IP address, no headers, no request body, and nothing that could identify a person — everything recorded here was sent by a program.
  • What is not a zero. A page nothing has asked for reports null. “Nobody came” and “we have not watched long enough” are different facts and the tables above keep them apart.
  • The arms. Even experiment numbers publish a brief, odd ones do not. The first version of this was a hash of the slug, which was unarguable and also landed 7 against 15 with almost every recent live page in one arm — a comparison between new pages and old ones with the treatment painted on top. Numbers run in time order, so alternating them balances the arms and stratifies them by age at once.
  • When a comparison is allowed. Not until each arm has 3 pages read by the answer layer. The threshold is in the source, so the experiment cannot start claiming a result the moment somebody feels like it has one.
What is actually being counted?

One row per request to a tracked public page whose user-agent matches a fixed table of known crawlers. The count happens in edge middleware, before any cache, because a page served from the edge cache never reaches a route handler and so is invisible everywhere else. Nothing about the request is stored beyond the agent name, the path and the time — no IP, no header dump, no body.

How do you know the agent is who it says it is?

We do not. A user-agent is a claim. Nothing here is verified against published IP ranges or reverse DNS, so an agent that lies is recorded as whatever it claimed to be. Verification is possible for some operators and is the obvious next thing to add; until it exists, every number on this page carries that caveat rather than a footnote apologising for it later.

Why is half the site missing a machine-readable brief?

Because a measurement with no control group is a chart, not evidence. Half the experiment pages publish a brief at /e/<slug>/brief.json and are listed in /llms.txt; the other half deliberately publish nothing of the kind. The rule is the experiment number: even numbers publish a brief, odd numbers do not. Experiment numbers run in time order, so the arms come out the same size and stratified by age, and anybody can check an assignment without running anything. It was fixed before any data existed and is not adjustable once a result starts to look promising.

What can this not tell you?

Whether any of it reaches an answer. Nothing here observes a citation, a recommendation, or an agent deciding to use something. Those need a channel that does not currently exist for a site like this, and inventing a score for them would be exactly the confident nonsense this experiment was built to argue against. What is measurable today is retrieval — who came, what they took, and how long they took to arrive.

Is a fetch worth anything on its own?

Probably not much, and that is a real risk to the whole design. A page can be crawled a hundred times and never appear in a single answer. The reason to measure it anyway is that retrieval is necessary even if it is not sufficient: nothing that was never fetched was ever cited, so the floor is worth knowing before anybody argues about the ceiling.

Could you not just make the numbers go up?

Easily, and that is why the arms exist. Publishing more pages, or pinging crawlers, would raise every total here without answering anything. The only question this page can settle is a difference between two groups of pages on the same site in the same window, and that difference is indifferent to how big the totals get.

Can I use this data?

Yes. The whole ledger is at /api/experiments/crawled/ledger as JSON, under CC BY 4.0, with the method and the taxonomy attached. If you cite a number from it, link the endpoint rather than a screenshot — the number will have moved by then, and the endpoint will be right.

exp-0022 · generated 2026-09-01 07:59Z · this is an experiment, it might be gone next month● a Project Nothing experiment