# Pre-registration — Research Paper No. 03
## Can AI crawlers read press releases?

**Registered:** 2026-10-07, by the PPN World Research Center, before any
robots.txt file, release page or archived copy was fetched for this study. The
git commit that adds this file is the timestamp.

**What had been looked at before registration** (feasibility only): the list of
websites that host the releases in the population (§2), the number of releases
per website and per publisher type, and the response time of one Wayback
Machine lookup (its status code only). No robots.txt rule, page header or
archived copy was read.

**What we already knew, disclosed.** Paper No. 01 (2026-10-06) checked the
robots.txt files behind a stratified sample of 13,125 release URLs and found 6%
of them disallowed at least one AI crawler, against 0.6% for Googlebot. That
result informed H1 below, which is therefore a replication at full scale, not a
discovery. The crawler documentation of OpenAI, Anthropic, Perplexity, Google,
Apple and Meta was read to build the crawler list (§4).

---

### 1. Question

Press releases are written to be repeated. AI assistants and answer engines are
now one of the places they are repeated. Do the websites that publish press
releases let AI crawlers read them, and do they treat crawlers that collect
training data differently from crawlers that fetch pages to answer questions?

Whether press releases are visible to AI has been argued about in the
communications trade, but, to our knowledge, it has not been measured across a
complete population of press-release publishers: companies, distributors and
public bodies, in many countries.

### 2. Population

Every release in PPN World's archive with a stated publication date from
2026-07-01 to 2026-10-06 (UTC), published by one of the 1,497 sources PPN World
serves on 2026-10-07, with a valid http(s) link: **500,129 releases on 1,503
websites**. A website is an origin (scheme, host and port), because that is the
scope of a robots.txt file.

Each website takes the publisher type of the sources that link to it (the type
contributing most releases where two types share one website; two websites are
in that case). Types are read from the served registry: government,
municipality, company, wire (commercial distributors), university, league,
nonprofit, political, other.

- **Public bodies** = government + municipality (781 websites).
- **Commercial publishers** = wire + company (346 websites).

### 3. Measurement

**robots.txt.** For each website, one `GET {origin}/robots.txt`, following up
to five redirects, 20-second timeout, with an honest client identity
(`PPNWorldResearch/1.0 (+https://ppnworld.com/research)`). If that request
fails or answers 401, 403, 429 or 5xx, it is repeated once as a desktop browser
and the browser's answer is used. A website whose file is served to a browser
but refused to the identified client is recorded as **walled**.

Each response is classified per RFC 9309 §2.3.1:

| Response | Class | Meaning applied |
|---|---|---|
| 2xx | **rules** | the file is parsed (a page served at that address parses to no rules) |
| 400–499 | **no file** | everything allowed |
| 5xx, timeout, DNS or TLS failure (both passes) | **unreachable** | excluded from rates (RFC: assume complete disallow — robustness R1) |

**Matching (RFC 9309).** A crawler's product token is matched
case-insensitively against each group's `User-agent` lines (a line's value is
cut at the first `/`). Every group naming the crawler applies, combined; if
none does, every `*` group applies, combined; if none exists, nothing is
disallowed. The longest matching `Allow`/`Disallow` path wins; on a tie, Allow
wins; `*` and `$` are honoured; `/robots.txt` itself is always allowed. The
rule is applied to **the path and query of each release URL**, not to the site
root, because a site can close one directory and leave the rest open.

**Website-level outcome.** A website *disallows* a crawler if more than half
of its in-window release URLs are disallowed to it.

**Release-level outcome.** Each release is disallowed or not, per crawler.

### 4. Crawlers

Fixed now, in four families:

| Family | Tokens |
|---|---|
| Search (reference) | `Googlebot`, `Bingbot` |
| AI training | `GPTBot`, `ClaudeBot`, `Google-Extended`, `Applebot-Extended`, `CCBot`, `Meta-ExternalAgent`, `Bytespider` |
| AI search (answer-engine indexes) | `OAI-SearchBot`, `Claude-SearchBot`, `PerplexityBot` |
| AI user-triggered fetchers | `ChatGPT-User`, `Claude-User`, `Perplexity-User` |

`Google-Extended` and `Applebot-Extended` are control tokens, not crawlers:
they govern use in model training (and, for Google-Extended, grounding in
Gemini), not search. OpenAI and Perplexity state that their user-triggered
fetchers may not follow robots.txt. The paper reports what each file *asks*,
not what each company does.

"Any AI crawler" = disallowed to at least one token in the three AI families.

### 5. Hypotheses and tests

Unit: the website. α = 0.05; Holm correction across H1–H3, all three reported.

**H1 — AI crawlers are turned away more often than search crawlers.** More
websites disallow `GPTBot` than disallow `Googlebot`.
*Test:* exact McNemar test on the paired outcome per website. *Supported* if
GPTBot is disallowed by more websites and the Holm-adjusted p < 0.05. The same
comparison is reported for every AI token (descriptive).

**H2 — training is refused more often than answering.** Websites disallow
OpenAI's training crawler (`GPTBot`) more often than its search crawler
(`OAI-SearchBot`).
*Test:* exact McNemar on the discordant websites (GPTBot-only vs
OAI-SearchBot-only). *Supported* if GPTBot-only websites outnumber
OAI-SearchBot-only websites and the Holm-adjusted p < 0.05. Anthropic's pair
(`ClaudeBot` vs `Claude-SearchBot`) is reported the same way, as a secondary
result.

**H3 — public bodies and commercial publishers differ.** The share of websites
disallowing `GPTBot` differs between public bodies and commercial publishers
(two-sided).
*Test:* Fisher's exact test; the difference in proportions is reported with a
95% Newcombe interval. *Supported* if the Holm-adjusted p < 0.05; the direction
is reported as found.

### 6. Secondary measures (reported, not tested)

- **S1.** Release-weighted versions of every rate ("share of press releases"),
  next to website-weighted ones, because a few distributors carry most
  releases (the largest website alone carries 16% of them).
- **S2. Twelve months earlier.** For each website, the Wayback Machine capture
  of its robots.txt (status 200) closest to 2025-10-07, within 2025-08-08 →
  2025-12-06, read raw and applied to the same release paths. Paired
  then-vs-now rates for GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended and
  any AI crawler, on the websites with a capture; coverage is reported.
- **S3.** Every crawler, every publisher type.
- **S4.** Content Signals: the share of files carrying `Content-Signal:` lines,
  and their `ai-train` / `ai-input` / `search` values.
- **S5.** Websites that disallow everything to every crawler (`User-agent: *`
  `Disallow: /` with no exception for Googlebot).
- **S6.** Walled and unreachable websites, by type.
- **S7. Page-level signals.** For each website, its most recent release in the
  window is fetched once (same two-pass protocol): the `X-Robots-Tag` header and
  `<meta name="robots|googlebot|…">` tags are read for `noindex`, `nosnippet`,
  `max-snippet:0`, `noai` and `noimageai`.
- **S8.** `llms.txt`: present if `{origin}/llms.txt` answers 2xx with a body
  that is not HTML and contains a line starting with `# `.
- **S9.** Countries with at least 20 websites (a website's country is the most
  frequent country of its releases).

### 7. Robustness (reported, none decides a hypothesis)

1. **R1.** Unreachable files counted as complete disallow (RFC 9309 §2.3.1.4).
2. **R2.** Release-weighted rates without the five largest websites.
3. **R3.** Identified client only (no browser retry).
4. **R4.** A website disallows a crawler if *any* of its release URLs is
   disallowed (instead of more than half).
5. **R5. Re-test.** Every robots.txt fetched a second time at least six hours
   after the first; the share of websites whose GPTBot outcome changed.

### 8. Reporting commitments

- The paper is published whatever the result.
- Deviations from this plan are listed, dated and explained.
- The per-website outcomes behind every figure are published as CSV, with the
  date each file was fetched.
- This document is published with the paper, unedited. Corrections are dated
  additions, never silent edits.
- Distributors are reported in aggregate and not named (PPN World does not name
  commercial sources it has no agreement with); public bodies may be named.
