Research

Measured, not opined

Controlled experiments and evidence audits on what actually makes pages slow — and what speed is worth. Published with full methodology, raw data, and honest limitations. Free to read, free to cite, no email gate.

INP yield-gaming audit + registered design

Earlier paint, or less work?

Mobile INP improved across the web while lab Total Blocking Time rose 58% in a year. Some of that INP win is real work reduction; some is yielding — repainting sooner without doing less, which the metric can't tell apart. We define the line outcome-wise, register a three-study design over CrUX + HTTP Archive + lab replay, and ship the kit. No fabricated share.

Interaction anatomy · outcome classification · four-source triangulation · replication kit

INP-era framework + taxonomy

When does lab predict field?

A green Lighthouse score isn't a Core Web Vitals pass — 43% of 90+ pages failed a CWV even in the FID era, and INP made the responsiveness gap harder to see in the lab. We reframe it as a transportability problem: four gates, a falsifiable 12-class divergence taxonomy, and a calibrated reconciliation protocol over public CrUX + HTTP Archive. No fabricated coefficients.

Lab×field confusion matrix · transportability gates · reproducibility kit

Power analysis · tool + protocol

How many Lighthouse runs prove a change?

“We ran it five times and the score went up two points” usually isn't evidence. Median-of-5 stabilises an estimate but can't prove a small before/after change. We derive the power tables — at a typical SD of 5 points, a 5-point change needs ~16 runs and a 3-point change ~44 — ship a free required-runs calculator, and register the variance census.

Analytic CI/power tables · Heričko secondary evidence · required-runs calculator

Model + registered benchmark

Do soft navigations repay the slow first load?

SPAs trade a heavier first load for faster in-app navigations. Chrome's soft-navigation API (147–149) finally makes that bargain measurable. Using real RUM-Archive ratios, the break-even is unforgiving: at ~1 follow-on navigation per session, a 1,000 ms first-load penalty needs ~1,015 ms desktop / ~1,143 ms mobile saved per navigation just to break even. Source-derived model + the benchmark to close it — no fabricated lab numbers.

Break-even model · amortization grid · pre-frozen 50-pair benchmark registry

Registered protocol · adoption-cohort

Do WordPress speed plugins work in the field?

WP Rocket, NitroPack and LiteSpeed promise to fix Core Web Vitals — but the evidence offered is a Lighthouse score or a cross-sectional dashboard, neither of which shows what happens to real users after a site adopts one. We publish the full causal adoption-cohort event study (HTTP Archive × CrUX), including the delay-JavaScript INP risk, and a runnable replication package — with no fabricated effect estimates.

Staggered DiD (Callaway–Sant'Anna) · delay-JS INP cohort · SQL + R + Python pack

Cloaking census · pilot + protocol

How many sites cheat PageSpeed?

Some sites serve Lighthouse a stripped, fast page while real visitors get the heavy one. The techniques are documented; the prevalence never has been. We ran a real user-agent-flip pilot across the Tranco top 2,000 — 0 confirmed in 1,162 measured (a lower bound on the crudest cloaking; a naïve detector's 9 flags were all benign) — and publish the full preregistered census protocol. No fabricated number.

Tranco-frame pilot · A/A-controlled · replication pack + protocol

Evidence audit · 2026

The cost of a millisecond

Twenty years of speed-versus-revenue studies, graded by research design. The headline everyone quotes — “a few percent per 100 ms” — sits roughly an order of magnitude above the best causal estimate of ≈0.65% per 100 ms. Speed matters; the marketed ROI is systematically overstated.

10 studies, design-graded · study register (CSV) · reconciled causal anchor

Registered study · pilot

Do Lighthouse's promised savings materialize?

Lighthouse prints “Est savings: 1,200 ms.” Does it show up when you do the fix? We ran a controlled pilot — predicted vs realized for Lighthouse 13's insights — and publish the full protocol to measure it at web scale. On clean pages the savings materialized, and were conservative (realized 1.0–1.9× predicted).

8-case pilot · calibration figures · open protocol + workbook

Causal decomposition · 2026

Did the web actually get faster?

CrUX's “good Core Web Vitals” share rose ~13 points since 2023 — but the metric and the instrument both changed mid-stream. The metric-comparable gain is +10.1 points, not +12.9, and Google's own release notes prove a real share is measurement, browser and data-quality change, not sites optimizing. With an event ledger and the BigQuery decomposition design.

41 months of CrUX data · event ledger · SQL + R reproducibility pack

Study · 2026

The cost of a kilobyte

210 controlled Lighthouse runs on synthetic pages that vary one thing at a time: JavaScript size and placement, main-thread work, render-blocking CSS, hero image weight, webfont strategy, and request-chain depth. The headline: on the wire, every critical-path kilobyte costs the same ~0.5 s per 100 KB — placement, execution and round trips are where the real money is.

Full methodology · raw dataset (CSV) · reproducible harness

Free tool

Lighthouse required-runs calculator

How many Lighthouse runs do you actually need to prove a before/after change? Enter your run-level spread and target change; get the runs per condition, a precision sweep, and a red/green read on whether median-of-5 is enough for your claim.

Confidence + power · paired or independent · CI sweep

Free tool

Performance budget calculator

Turn an LCP target and a connection profile into per-resource byte budgets — modeled on TCP slow start instead of naive bandwidth division — and export a ready-to-use Lighthouse budget.json for CI.

RFC 6928 slow-start model · Lighthouse/WebPageTest profiles

Free tool

PageSpeed cloaking test

Check whether a site is “cheating” PageSpeed — showing Google a faster page than real visitors get. We compare Google's own PageSpeed Insights view against the page a normal visitor receives, and flag structural mismatches with the evidence.

PSI ground truth · lab-vs-field check · user-agent flip

How we publish

  • Design before numbers. Whether it's a controlled experiment or an audit of others' studies, the research design is graded first — effect sizes are only trusted in proportion to how they were measured.
  • Raw data included. Every study ships its complete dataset as CSV — the experiment's per-run results, or the audit's full study register — so you can check our math.
  • Limitations stated, not buried. Lab data and observational data are both models of reality. Each study says exactly where the model ends.
  • Reproducible. Methodology sections carry everything needed to re-run the work: versions and throttling profiles for experiments, provenance and grading for audits.

Questions about the data, or a replication to share? Get in touch — we link replications.