Revisions
Published work changes: datasets get re-measured, methods get corrected, analyses get updated. Everything that changes is recorded here, including the corrections that made earlier numbers wrong.
Datasets & methods
22 entriesCorrection to what the refusal rate counts, not to any measurement. The published rate divided every excluded site by the sample — 14 of 38, 37% — when four of those fourteen exclusions are not refusals: Bombas and Chewy returned HTTP 429 (a rate limit, the distinction Rev D itself turned on), Lululemon failed to connect, and Patagonia returned HTTP 404 to the declared bot while serving a browser user-agent normally — a UA-conditional response, verified by requesting both ways, not a dead URL. The refusal rate is now computed from stated policy only — HTTP 401/403, bot-challenge walls, robots.txt disallows — and falls from 37% to 26% (10 of 38). No site was re-measured and no dataset changed; the classification of already-recorded reasons did. The sampling caveat separates the categories explicitly, and the collector now labels a 429 and a 404 as what they are so the conflation cannot recur at source.
The same counting correction as the ecommerce entry above, and it moves this sector's headline number most, because this sample carries the most non-refusal exclusions. The published 49% divided all 17 exclusions by the 35 sampled; five of those are connection failures (Costco, Kroger, Staples, Office Depot, Tractor Supply) and one is a rate limit (Chewy) — the network and the origin saying something, not the chain stating a policy. Counted as stated policy only, the refusal rate is 31% (11 of 35, all HTTP 403). That remains the highest refusal rate of any vertical measured here. The other figure this page has always led with is unchanged in substance: only 18 of 35 sampled chains could be measured at all, and the sheet still says the survivors are the chains willing to answer a non-browser client, not the retail sector. The analysis article led with the conflated figure in its headline — 'Why Half of Large Retail Homepages Refuse a Bot GET' — and has been retitled and patched to carry the reclassified numbers; its point that roughly half the sample cannot be measured stays true, but 'refused' now means refused.
Correction, and it retracts the headline finding of Rev C. Last week this sample measured 13 of 38 and the refusal rate was published as 66% — the highest recorded anywhere on this site. It did not reproduce. Twenty-four of 38 answered today and the refusal rate is 37%. Every one of the eleven brands that returned this week — Allbirds, Glossier, Casper, Away, Everlane, ThirdLove, Ritual, Brooklinen, Liquid Death, Olipop and Magic Spoon — was excluded last week for HTTP 429, and 429 is a rate-limit response, not a refusal. Rev C counted temporary throttling as sector policy and drew a conclusion from it. It should not have: a 403 is a decision about who may fetch the page, a 429 is a statement about how often, and only the first belongs in a refusal rate. Today's run was more aggressive, not less — all nine sectors were measured concurrently rather than one after another — so the recovery cannot be attributed to gentler collection. Headline figures move accordingly on the larger sample: median delivered HTML 1,257 KB → 878 KB, median third-party domains 12 → 16, consent mechanisms 62% → 71% (17 of 24), main landmark 85% → 92% (22 of 24), and 23 of 24 now serve the document compressed. The 429-heavy exclusions were concentrated in the Shopify-hosted segment, which is precisely the part of the sample Rev C dropped and therefore the part its figures were least able to describe. Refusal rates on this sample should be read with a run-to-run error bar of tens of points until the 429 behaviour is understood; it is not a stable sector property.
- Government observatoryRev C
First re-measurement of this sector under the corrected method, and the drop is the method rather than the sector. Government was last measured on 2026-07-28, before the collector stopped sending a Chrome user-agent and started reading robots.txt. Measured as a declared bot it falls from 44 of 51 to 41: Colorado, Michigan, Vermont and Nevada now refuse a request that identifies itself, Washington's robots.txt disallows the homepage and is excluded and named rather than fetched anyway, and Ohio returns HTTP 404 — the sampled URL no longer resolves to a page and needs re-checking rather than being reported as a refusal. Three portals moved the other way: Arkansas, Minnesota and New York all answered a declared bot having blocked a browser-shaped one a week ago. Refusal rate 14% → 20%. Consequent headline moves — main landmark 89% → 85% (35 of 41), recognised trackers 82% → 80%, HSTS 66% → 63% — follow the composition change and are not sector improvement or decline. Two genuine per-site changes among portals measured in both runs: Kentucky and Virginia both stopped sending an HSTS header, and Wisconsin's image alt-text coverage fell from 93% to 25%.
- Legal observatoryRev C
First re-measurement under the corrected method: 32 of 38, one more than Rev B. McDermott Will & Emery and Fenwick both answer a declared bot after blocking the previous run, and Norton Rose Fulbright is now excluded because its robots.txt disallows the homepage — an exclusion the old method would never have recorded, because the old method never asked. Refusal rate 18% → 16%. Headline movement is small and composition-driven: consent mechanisms 65% → 63% (20 of 32), pre-consent cookies 61% → 59%, and 30 of 32 serve the document compressed. Recognised trackers hold at 94%, still the highest of any sector measured here.
- Fintech observatoryRev C
Weekly re-measurement: 27 of 35, one fewer than Rev B. The single removal is Deel, whose robots.txt now disallows the homepage — it permitted the same fetch three days ago, so this is a change at Deel rather than a change in method. Refusal rate 20% → 23%. Every headline move follows that one removal: recognised trackers 68% → 67%, consent mechanisms 43% → 41%, pre-consent cookies 68% → 63%, main landmark 82% → 81%, median delivered HTML 547 KB → 494 KB. No surviving site changed a structural signal. These figures should be read as a change in sample composition, not as sector movement.
Weekly re-measurement: 18 of 35, one more than Rev B, taking the refusal rate from 51% to 49%. Target answered today after returning HTTP 429 last week — the same rate-limit-versus-refusal distinction that forced the ecommerce correction above, and a reminder that this sector's headline number is a refusal rate with real run-to-run variance in it. One genuine change among the chains measured in both runs: Ulta Beauty now serves its homepage compressed, taking compression to 18 of 18. Consent mechanisms 47% → 44% (8 of 18) and main landmark 76% → 78% (14 of 18) both follow the addition of Target rather than any site changing. This sample remains the chains willing to serve a non-browser client, not the retail sector.
- All observatories — weekly re-measurementNote
All nine registered observatories were re-measured today on the shared script, including the two retired as published verticals. Four returned an identical sample with no structural movement and therefore take no revision letter: SaaS (33 of 35), AI platforms (33 of 35), developer tools (33 of 35) and healthcare (29 of 35). Healthcare was unchanged on every structural signal. The others carry small per-site changes too minor to move a headline: Segment now sets a pre-consent cookie and Okta one fewer, taking SaaS pre-consent cookies from 26 of 33 to 27; Airtable added a tracker and lost image alt-text coverage. Median server response moved in every sector again — SaaS 230 ms → 186 ms, AI 128 ms → 120 ms, healthcare 230 ms → 272 ms, government 315 ms → 158 ms — and is again recorded as single-request timing variance rather than a finding, per the note of 2026-07-31. Search Console remains unconfigured, so no measured-position data was available to check this week's tracked rankings against.
- All observatories — method correctionCorrection
Correction to the measurement method itself, and a re-measurement of all seven sectors under it. Until today the collector sent a Chrome user-agent string and never read robots.txt. That is browser impersonation, and it is indefensible for a publication that criticises sites for refusing automated requests and publishes an article describing the Robots Exclusion Protocol as a voluntary convention — volunteering is the point. It also changed what the figures meant: a refusal rate gathered behind a browser user-agent measures what a site serves something that LOOKS like a browser, not what it serves a declared bot, which is what the articles said it measured. The collector now identifies as DigitariseBot/1.0 with a contact URL, fetches robots.txt first, and excludes-and-names any homepage that disallows it. The recorded method statement, which said only "single HTTPS homepage GET per site via curl", now states the user-agent and the robots.txt step. Re-measured honestly, six of seven sectors returned an identical sample: SaaS, AI platforms, developer tools, fintech, retail and healthcare all measured exactly as before. DTC ecommerce did not — it fell from 20 measured to 13, taking its refusal rate from 47% to 66%. Those seven sites served a browser user-agent and refuse a declared one, which is the single clearest demonstration of why the old method was wrong.
- All observatories — timing varianceNote
A finding from re-measuring everything on one afternoon, recorded because it bounds how any of these figures should be read. Median server response moved substantially in every sector on a re-run with no change in sample composition: SaaS 351 ms to 230 ms, AI platforms 282 ms to 128 ms, developer tools 255 ms to 133 ms, fintech 294 ms to 144 ms, healthcare 348 ms to 230 ms. These are single-request measurements taken from one location, and a 40% swing between runs is network and time-of-day variance, not sectors becoming faster in a week. The structural signals — protocol version, compression, landmarks, headings, tracker and consent presence — are stable across runs and are what these datasets are actually good for. Median server response should be read as indicative only, and comparisons between sectors measured at different times of day should not be made at all. Prior revisions quoted response-time movements as if they were sector changes; they were not.
- Industry setRev D
Sector coverage repointed at technology verticals. Five observatories added on first measurement: B2B SaaS (33 of 35), AI platforms (33 of 35), developer tools and cloud infrastructure (33 of 35), fintech and payments (28 of 35), and large US retail chains (17 of 35). Legal and Government were retired as published verticals; their hub and statistics routes now return 410 Gone, deliberately rather than 404, because the pages existed and were withdrawn. Both datasets remain registered and downloadable: published analyses still embed those observatories, and the benchmark corpus is more representative with them than without. The three tech samples share the lowest refusal rate measured anywhere on this site at 6%, which is itself a finding — these are sectors that do not treat a non-browser client as an attack.
- Retail chains observatoryRev A
Initial measurement, published with its limit stated in the title. 18 of 35 sampled chains refused a plain HTTPS request — a 51% refusal rate, the highest of any vertical measured here and higher than the DTC sample's 47%. Twelve returned 403, four 429 or 307, and four failed to connect at all. The 17 that answered are published as the chains willing to serve a non-browser client, not as the retail sector, and the statistics sheet leads on the refusal rate rather than on a median. Read alongside the DTC ecommerce observatory: digital-native brands answer, incumbent chains largely do not.
- Sector observatory widgetCorrection
Correction, and a material one. The widget resolved its dataset through a hand-maintained map separate from the sector registry, and that map never gained an `ecommerce` entry. Because an unknown key fell back to healthcare rather than failing, the DTC ecommerce hub and statistics sheet rendered the HOSPITAL sample — under an ecommerce heading, beneath a caption stating the figures above were recomputable from the table shown. They were not. The widget now derives its datasets from the single registry and renders an explicit notice for an unknown sector instead of substituting another sector's rows. The cross-sector comparator had the same duplicate-list weakness and is now derived too. Headline figures on the ecommerce statistics sheet were always computed from the correct dataset and did not move; the table beneath them did.
Weekly re-measurement. Sample composition unchanged — 29 of 35, the same six exclusions — and no structural signal moved: tracker, consent, main-landmark and heading figures are identical to Rev B. Median server response went 337 ms → 348 ms and the slowest homepage 1,692 ms → 1,101 ms, which is single-request timing variance rather than a change in the sector. Correction carried from Rev B: this sector was described as having the heaviest documents and the most third-party domains of any sector measured. That stopped being true when the DTC ecommerce dataset landed on 2026-07-20 at a 846 KB median document and 16 third-party domains against healthcare's 153 KB and 11. The description has been corrected.
- Legal observatoryRev B
Weekly re-measurement: 31 of 38 measured, one fewer than Rev A, because McDermott Will & Emery now answers an automated request with a bot-challenge page. That single removal accounts for every headline move — recognised trackers 91% → 94%, consent mechanisms 62% → 65%, main landmark 81% → 84%, and one fewer homepage shipping no first-level heading — because the firm scored negative on all four and no surviving site changed any of them. These figures should be read as a change in sample composition, not as sector improvement.
Weekly re-measurement: 20 of 38 measured. Bombas now returns HTTP 429, taking the sector's refusal rate from 45% to 47%. Two genuine changes among the surviving sites: Article added a first-level heading, and Ulta Beauty now serves its homepage compressed. The tracker (86% → 90%) and main-landmark (86% → 90%) moves are not sector change — both follow from Bombas, which had neither, leaving the denominator.
- Government observatoryRev B
Weekly re-measurement with sample composition unchanged at 44 of 51. One real change: Pennsylvania now sets a cookie before any consent interaction is possible, moving pre-consent cookies from 48% to 50% of measured portals. Median server response moved 386 ms → 315 ms; Kentucky returned 3,453 ms against 348 ms a week earlier, which is recorded as a single-request outlier and not treated as a finding.
- Government observatoryRev A
Initial measurement of 44 of 51 sampled US state and District of Columbia portals (7 excluded and named). Recorded here retrospectively: the dataset was published on 2026-07-26 without a revision entry, which this corrects.
Initial measurement. 17 of 38 sampled retail homepages refused an automated request — a 45% refusal rate concentrated on large enterprise retailers. The surviving 21 sites are therefore published as a digital-native DTC sample, not as a representative ecommerce sector, and the refusal rate is recorded as a measurement in its own right rather than a footnote.
- Legal observatoryRev A
Initial measurement of 32 large US law-firm homepages (38 sampled, 6 excluded and named).
Re-measured on the shared observatory script. Correction: the first run's header parsing understated document size and compression, and used a narrower tracker list — tracker and consent figures moved materially. Both sectors now use identical detection, which is what makes cross-sector comparison legitimate.
Initial measurement of 29 US hospital and health-system homepages.
Published analyses
17 issuedInitial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Initial issue; no revisions since publication.
Corrections are published, not silently patched. If a number changes, the reason is on this page.
Read the editorial policy