Crawled — a first-party ledger of which machines read this website, which pages they take, and how long after publication they arrive
Crawled
Every machine that has read this website, what it took, and how long after the page existed it turned up.
- Fetches recorded
- 206since 2026-08-30 11:45Z
- By the answer layer
- 130training, AI search, live assistant fetches
- By search engines
- 76the control column
- Distinct machines
- 11named agents only
- Pages reached
- 26/27read at least once by the answer layer
The machines
Split by what each operator says its crawler is for. The purpose matters more than the count: a page fetched by ChatGPT-User seconds after somebody asked a question is a different event from the same page being swept into a training corpus.
| Agent | Operator | Layer | What it is for | Fetches | Last seen |
|---|---|---|---|---|---|
| Bytespider | ByteDance | answer | Building a training corpus | 57 | 2026-09-01 06:06Z |
| ClaudeBot | Anthropic | answer | Building a training corpus | 35 | 2026-09-01 07:59Z |
| GPTBot | OpenAI | answer | Building a training corpus | 15 | 2026-09-01 00:49Z |
| OAI-SearchBot | OpenAI | answer | Indexing for AI answers | 14 | 2026-09-01 07:19Z |
| Amazonbot | Amazon | answer | Indexing for AI answers | 5 | 2026-09-01 07:00Z |
| ChatGPT-User | OpenAI | answer | Fetching live to answer somebody | 4 | 2026-08-31 16:51Z |
| bingbot | Microsoft | search | Indexing for a search engine | 48 | 2026-09-01 06:41Z |
| PetalBot | Huawei | search | Indexing for a search engine | 11 | 2026-09-01 07:34Z |
| DuckDuckBot | DuckDuckGo | search | Indexing for a search engine | 9 | 2026-09-01 07:08Z |
| Applebot | Apple | search | Indexing for a search engine | 4 | 2026-09-01 07:57Z |
| Googlebot | search | Indexing for a search engine | 4 | 2026-08-31 15:37Z |
First contact
How long each experiment’s page existed before a machine read it. Publication time is the launch recorded in the control plane, not the day somebody decided to announce it. A dash means it has not happened.
| Page | Arm | Published | First read by the answer layer | After | Fetches |
|---|---|---|---|---|---|
| Needed | brief | 2026-08-31 | ClaudeBot · 2026-09-01 07:24Z | 15.5 h | 2 |
| Slop | holdout | 2026-08-30 | GPTBot · 2026-08-31 02:09Z | 8.5 h | 3 |
| Done | brief | 2026-08-30 | GPTBot · 2026-08-31 02:09Z | 9.2 h | 3 |
| Up Peak | holdout | 2026-08-30 | GPTBot · 2026-08-31 02:09Z | 9.7 h | 4 |
| Crawled | brief | 2026-08-30 | GPTBot · 2026-08-31 02:09Z | 12.2 h | 4 |
| Out of Ten | holdout | 2026-08-30 | GPTBot · 2026-08-31 02:09Z | 17.2 h | 3 |
| Perfect Score | brief | 2026-08-29 | GPTBot · 2026-08-31 02:09Z | 40.2 h | 2 |
| Insula | holdout | 2026-08-28 | ClaudeBot · 2026-09-01 07:24Z | 4 d | 2 |
| Exactly | brief | 2026-08-26 | ClaudeBot · 2026-09-01 07:24Z | 6 d | 2 |
| Spent | holdout | 2026-08-25 | ClaudeBot · 2026-09-01 07:25Z | 7 d | 2 |
| Sieve | brief | 2026-08-24 | Bytespider · 2026-08-30 12:16Z | 6 d | 5 |
| Hold the Note | holdout | 2026-08-22 | Bytespider · 2026-08-30 17:26Z | 8 d | 2 |
| Signal Weather | brief | 2026-05-20 | Bytespider · 2026-08-31 10:45Z | 103 d | 3 |
| AI-deas | brief | 2026-05-11 | Bytespider · 2026-08-31 07:10Z | 112 d | 2 |
| Built From a Phone | holdout | 2026-05-04 | ClaudeBot · 2026-09-01 07:26Z | 120 d | 3 |
| Micro-Break | holdout | 2026-05-04 | Bytespider · 2026-08-31 08:31Z | 119 d | 4 |
| Code of Laws | holdout | 2026-05-04 | Bytespider · 2026-08-31 01:00Z | 118 d | 3 |
| LLM Price Volatility | brief | 2026-05-04 | Bytespider · 2026-08-30 12:34Z | 118 d | 4 |
| Voidle | holdout | 2026-05-04 | Bytespider · 2026-08-30 18:54Z | 118 d | 4 |
| Velocity Cards | brief | 2026-05-04 | Bytespider · 2026-08-30 18:13Z | 118 d | 2 |
| Instruction Elasticity Index | holdout | 2026-05-04 | ClaudeBot · 2026-09-01 07:25Z | 120 d | 1 |
| The Replacement Index | brief | 2026-05-04 | Bytespider · 2026-08-30 22:19Z | 119 d | 2 |
| Signal from the Void | holdout | 2026-05-04 | Bytespider · 2026-08-30 14:28Z | 118 d | 2 |
| Subscribe to Nothing | brief | 2026-05-04 | Bytespider · 2026-08-31 04:51Z | 119 d | 3 |
| Make Interesting | holdout | 2026-05-04 | Bytespider · 2026-08-31 00:49Z | 118 d | 5 |
| On Me | brief | 2026-08-30 | not yet | — | 1 |
| Mutation Observer | brief | 2026-05-04 | ClaudeBot · 2026-09-01 07:24Z | 120 d | 1 |
Brief against holdout
The answer layer has reached 93% of the pages that publish a brief and 100% of the pages that do not. Median time from publishing to first machine read: 2473.8h with a brief, 2837.5h without. This is a difference between two small groups on one small site, not a finding about the web.
Both arms have passed 3 read pages, so the columns are shown side by side. They are still two small groups on one small site.
| Arm | Pages | Read | Median time to first read |
|---|---|---|---|
| Machine-readable brief | 14 | 13 | 103 d |
| HTML only (holdout) | 13 | 13 | 118 d |
Does anything read the machine-readable surfaces?
/llms.txt and the per-experiment briefs exist because a great deal of advice says they should. Whether any machine has ever fetched one is a question with an answer, and this is it.
| Surface | Fetches | Read by |
|---|---|---|
| /llms.txt | 1 | GPTBot |
The tail
The last 120 fetches, newest first. Nothing here is live-updating; reload it.
- ClaudeBot/experiments/crawled
- ClaudeBot/experiments/done
- ClaudeBot/experiments/insula
- ClaudeBot/experiments/needed
- Applebot/robots.txt
- PetalBot/robots.txt
- ClaudeBot/e/out-of-ten
- ClaudeBot/e/signal-weather
- ClaudeBot/
- ClaudeBot/e/signal-from-the-void
- ClaudeBot/e/llm-price-volatility
- ClaudeBot/e/built-from-a-phone
- ClaudeBot/e/velocity-cards
- ClaudeBot/e/the-replacement-index
- ClaudeBot/e/code-of-laws
- ClaudeBot/e/subscribe-to-nothing
- ClaudeBot/e/ai-deas
- ClaudeBot/e/instruction-elasticity
- ClaudeBot/e/up-peak
- ClaudeBot/e/hold-the-note
- ClaudeBot/e/perfect-score
- ClaudeBot/e/slop
- ClaudeBot/e/sieve
- ClaudeBot/e/spent
- ClaudeBot/e/micro-break
- ClaudeBot/e/mutation-observer
- ClaudeBot/e/make-interesting
- ClaudeBot/e/exactly
- ClaudeBot/e/crawled
- ClaudeBot/e/done
- ClaudeBot/experiments/exactly
- ClaudeBot/e/insula
- ClaudeBot/e/voidle
- ClaudeBot/e/needed
- ClaudeBot/robots.txt
- OAI-SearchBot/robots.txt
- DuckDuckBot/
- Amazonbot/experiments/needed
- ClaudeBot/
- ClaudeBot/robots.txt
- bingbot/
- Bytespider/robots.txt
- Bytespider/robots.txt
- bingbot/e/code-of-laws
- PetalBot/sitemap.xml
- bingbot/e/llm-price-volatility
- bingbot/sitemap.xml
- bingbot/robots.txt
- bingbot/robots.txt
- bingbot/e/sieve
- Bytespider/experiments/velocity-cards
- PetalBot/e/crawled
- PetalBot/experiments/insula
- bingbot/e/on-me
- bingbot/e/needed
- bingbot/sitemap.xml
- Bytespider/robots.txt
- PetalBot/experiments/exactly
- Bytespider/robots.txt
- GPTBot/sitemap.xml
- OAI-SearchBot/robots.txt
- bingbot/experiments/sieve
- bingbot/e/built-from-a-phone
- PetalBot/robots.txt
- bingbot/e/make-interesting
- bingbot/experiments/signal-from-the-void
- OAI-SearchBot/
- OAI-SearchBot/
- OAI-SearchBot/experiments/signal-from-the-void
- OAI-SearchBot/robots.txt
- OAI-SearchBot/robots.txt
- bingbot/experiments/code-of-laws
- Amazonbot/
- OAI-SearchBot/robots.txt
- bingbot/e/micro-break
- bingbot/e/voidle
- OAI-SearchBot/robots.txt
- Bytespider/robots.txt
- Applebot/robots.txt
- Bytespider/robots.txt
- ChatGPT-User/
- ChatGPT-User/
- Applebot/robots.txt
- Googlebot/robots.txt
- DuckDuckBot/
- PetalBot/e/up-peak
- PetalBot/e/insula
- DuckDuckBot/
- Amazonbot/experiments/i-built-a-game-from-my-phone
- DuckDuckBot/
- Bytespider/e/signal-weather
- DuckDuckBot/
- ChatGPT-User/
- bingbot/e/sieve
- Googlebot/robots.txt
- bingbot/e/signal-weather
- Bytespider/e/micro-break
- PetalBot/e/exactly
- Bytespider/e/ai-deas
- bingbot/robots.txt
- bingbot/experiments/sieve
- bingbot/
- ChatGPT-User/
- Bytespider/robots.txt
- bingbot/e/make-interesting
- PetalBot/e/spent
- Bytespider/e/subscribe-to-nothing
- bingbot/e/subscribe-to-nothing
- bingbot/experiments/signal-from-the-void
- bingbot/experiments/code-of-laws
- PetalBot/robots.txt
- DuckDuckBot/
- GPTBot/llms.txt
- GPTBot/experiments/done
- GPTBot/experiments/slop
- GPTBot/experiments/perfect-score
- GPTBot/experiments/up-peak
- GPTBot/experiments/crawled
- GPTBot/experiments/out-of-ten
- GPTBot/e/perfect-score
What this is, if you want it after the fact
For thirty years the question was where a page ranked. It is becoming a different question: whether a machine reads the page at all, and whether what it takes away is what the page actually said. Almost everything written about that second question is advice. Very little of it is measurement, because measuring it requires a site, a control group, and a willingness to publish the result when it is boring.
This is the smallest honest instrument for it. Every request to a tracked page is matched against a table of known crawlers, and one row is written: which agent, which path, when. From that come three things that are hard to get any other way — which operators actually fetch a small site, how long a page waits between being published and being read, and whether the answer layer behaves at all differently from the search engines crawling the same pages in the same window.
Half the experiment pages publish a machine-readable brief. Half do not, on purpose, by a rule fixed before there was anything to see. That holdout is the only reason any sentence on this page can be a finding rather than an anecdote, and it costs something real: half the portfolio is deliberately published in the worse condition, if the worse condition turns out to be worse.
What this deliberately does not measure
A fetch is not a citation. A citation is not a recommendation. A recommendation is not an agent choosing to use something. Those are four different events and only the first is observable from a server log.
It would be easy to invent scores for the other three — a retrievability index, a citation-worthiness rating, an agent-readiness percentage out of a hundred. They would look authoritative, they would be screenshot-friendly, and they would be fabricated, because there is no ground truth on this site to check them against. Anything on this page that is a number came from a request that actually arrived.
The honest version of the missing half is a second instrument that asks real models real questions on a schedule and records whether this site comes back — mention, citation, and whether the description is accurate. That is measurable, it is not cheap, and it is not built yet. It is the obvious next thing, and it stays unbuilt rather than being simulated in the meantime.
The method, in full
- Where the count happens. Edge middleware, before the cache. A story page is served from the edge for sixty seconds at a time, so a crawler fetch never reaches a route handler — anywhere later in the stack, most of this traffic is invisible.
- Who is counted. Only agents in a published table, matched on a substring of the user-agent. Unrecognised agents are not counted at all: an unknown-bot bucket would be mostly scrapers and uptime checks, and putting that beside GPTBot would flatter the result.
- What is stored. Agent, path, timestamp. No IP address, no headers, no request body, and nothing that could identify a person — everything recorded here was sent by a program.
- What is not a zero. A page nothing has asked for reports
null. “Nobody came” and “we have not watched long enough” are different facts and the tables above keep them apart. - The arms. Even experiment numbers publish a brief, odd ones do not. The first version of this was a hash of the slug, which was unarguable and also landed 7 against 15 with almost every recent live page in one arm — a comparison between new pages and old ones with the treatment painted on top. Numbers run in time order, so alternating them balances the arms and stratifies them by age at once.
- When a comparison is allowed. Not until each arm has 3 pages read by the answer layer. The threshold is in the source, so the experiment cannot start claiming a result the moment somebody feels like it has one.
What is actually being counted?
One row per request to a tracked public page whose user-agent matches a fixed table of known crawlers. The count happens in edge middleware, before any cache, because a page served from the edge cache never reaches a route handler and so is invisible everywhere else. Nothing about the request is stored beyond the agent name, the path and the time — no IP, no header dump, no body.
How do you know the agent is who it says it is?
We do not. A user-agent is a claim. Nothing here is verified against published IP ranges or reverse DNS, so an agent that lies is recorded as whatever it claimed to be. Verification is possible for some operators and is the obvious next thing to add; until it exists, every number on this page carries that caveat rather than a footnote apologising for it later.
Why is half the site missing a machine-readable brief?
Because a measurement with no control group is a chart, not evidence. Half the experiment pages publish a brief at /e/<slug>/brief.json and are listed in /llms.txt; the other half deliberately publish nothing of the kind. The rule is the experiment number: even numbers publish a brief, odd numbers do not. Experiment numbers run in time order, so the arms come out the same size and stratified by age, and anybody can check an assignment without running anything. It was fixed before any data existed and is not adjustable once a result starts to look promising.
What can this not tell you?
Whether any of it reaches an answer. Nothing here observes a citation, a recommendation, or an agent deciding to use something. Those need a channel that does not currently exist for a site like this, and inventing a score for them would be exactly the confident nonsense this experiment was built to argue against. What is measurable today is retrieval — who came, what they took, and how long they took to arrive.
Is a fetch worth anything on its own?
Probably not much, and that is a real risk to the whole design. A page can be crawled a hundred times and never appear in a single answer. The reason to measure it anyway is that retrieval is necessary even if it is not sufficient: nothing that was never fetched was ever cited, so the floor is worth knowing before anybody argues about the ceiling.
Could you not just make the numbers go up?
Easily, and that is why the arms exist. Publishing more pages, or pinging crawlers, would raise every total here without answering anything. The only question this page can settle is a difference between two groups of pages on the same site in the same window, and that difference is indifferent to how big the totals get.
Can I use this data?
Yes. The whole ledger is at /api/experiments/crawled/ledger as JSON, under CC BY 4.0, with the method and the taxonomy attached. If you cite a number from it, link the endpoint rather than a screenshot — the number will have moved by then, and the endpoint will be right.