Research

Measured, not opined

Controlled experiments and evidence audits on what actually makes pages slow — and what speed is worth. Published with full methodology, raw data, and honest limitations. Free to read, free to cite, no email gate.

  • 10 studies
  • 5 free tools
  • Raw data or runnable code on every one
  • CC BY 4.0

Research library

Nine more studies, filed by the decision they settle

Start from what you're trying to decide, not from our taxonomy. Every card shows the result before the click, and links straight to the data behind it.

FIELD INP Good Poor LAB TBT Representative False confidence lab fine · users slow False alarm Representative Good Poor

INP-era framework + taxonomy

When does lab predict field?

A green Lighthouse score isn't a Core Web Vitals pass — 43% of 90+ pages failed a CWV even in the FID era, and INP made the responsiveness gap harder to see in the lab. We reframe it as a transportability problem: four gates, a falsifiable 12-class divergence taxonomy, and a calibrated reconciliation protocol over public CrUX + HTTP Archive. No fabricated coefficients.

Lab×field confusion matrix · transportability gates · reproducibility kit

Runs required per condition · σ = 5 median-of-5 16 44 detect 5 points detect 3 points

Power analysis · tool + protocol

How many Lighthouse runs prove a change?

“We ran it five times and the score went up two points” usually isn't evidence. Median-of-5 stabilises an estimate but can't prove a small before/after change. We derive the power tables for comparing run means — in a high-noise scenario (SD of 5 points), a 5-point change needs ~16 runs and a 3-point change ~44 — ship a free required-runs calculator that sizes for medians, and register the variance census.

Analytic CI/power tables · Heričko secondary evidence · required-runs calculator

Break-even saving per navigation 1,000 1,015 1,143 penalty desktop mobile

Model + registered benchmark

Do soft navigations repay the slow first load?

SPAs trade a heavier first load for faster in-app navigations. Chrome's soft-navigation API (147–149) finally makes that bargain measurable. Using real RUM-Archive ratios, the break-even is unforgiving: at ~1 follow-on navigation per session, a 1,000 ms first-load penalty needs ~1,015 ms desktop / ~1,143 ms mobile saved per navigation just to break even. Source-derived model + the benchmark to close it — no fabricated lab numbers.

Break-even model · amortization grid · pre-frozen 50-pair benchmark registry

Adoption-cohort event study PLUGIN FIRST DETECTED pre-trend field outcome HTTP Archive markers CrUX real users

Registered protocol · adoption-cohort

Do WordPress speed plugins work in the field?

WP Rocket, NitroPack and LiteSpeed promise to fix Core Web Vitals — but the evidence offered is a Lighthouse score or a cross-sectional dashboard, neither of which shows what happens to real users after a site adopts one. We publish the full causal adoption-cohort event study (HTTP Archive × CrUX), including the delay-JavaScript INP risk, and a runnable replication package — with no fabricated effect estimates.

Staggered DiD (Callaway–Sant'Anna) · delay-JS INP cohort · SQL + R + Python pack

Pilot detection funnel 2,000 frame 1,162 measured 9 flags 0 confirmed

Cloaking census · pilot + protocol

How many sites cheat PageSpeed?

Some sites serve Lighthouse a stripped, fast page while real visitors get the heavy one. The techniques are documented; the prevalence never has been. We ran a real user-agent-flip pilot across the Tranco top 2,000 — 0 confirmed in 1,162 measured (a lower bound on the crudest cloaking; a naïve detector's 9 flags were all benign) — and publish the full preregistered census protocol. No fabricated number.

Tranco-frame pilot · A/A-controlled · replication pack + protocol

Claim strength rises with research design case study observational quasi-causal ≈0.65% per 100 ms clean designs

Evidence audit

The cost of a millisecond

Twenty years of speed-versus-revenue studies, graded by research design. The headline everyone quotes — “a few percent per 100 ms” — sits roughly an order of magnitude above the best causal estimate of ≈0.65% per 100 ms. Speed matters; the marketed ROI is systematically overstated.

11 records, design-graded · study register (CSV) · reconciled reference + the one reported null

Realized ÷ predicted savings · 8 of 8 cases 1.0× parity 1.9× every controlled pilot case landed in this band

Registered study · pilot

Do Lighthouse's promised savings materialize?

Lighthouse prints “Est savings: 1,200 ms.” Does it show up when you do the fix? We ran a controlled pilot on Lighthouse 13's insights and publish the full protocol to measure it at web scale. On clean pages the promise was never bigger than the move in Lighthouse's own simulated metric — 1.0–1.9× — which is internal consistency, not yet an independently timed saving.

8-case pilot · calibration figures · open protocol + workbook

“Good CWV” gain — two different rulers +12.9 pp +14.3 pp spliced FID→INP one metric (INP)

Causal decomposition

Did the web actually get faster?

CrUX's “good Core Web Vitals” share rose ~13 points since 2023 — but that +12.9 splices a FID-era start onto an INP-era end. On a single metric the gain is +14.3 points, because the switch lowered the reported score — and Google's own release notes prove a real share of the climb is measurement, browser and data-quality change, not sites optimizing. With an event ledger and the BigQuery decomposition design.

41 months of CrUX data · event ledger · SQL + R reproducibility pack

Critical-path cost · tested Slow 4G 100 KB on the wire ≈0.5 s transfer cost placement execution round trips The bytes are the baseline; where and how they run decides the damage.

Controlled experiment · 210 runs

The cost of a kilobyte

210 controlled Lighthouse runs on synthetic pages that vary one thing at a time: JavaScript size and placement, main-thread work, render-blocking CSS, hero image weight, webfont strategy, and request-chain depth. The headline: on the wire, every critical-path kilobyte costs the same ~0.5 s per 100 KB — placement, execution and round trips are where the real money is.

Full methodology · raw dataset (CSV) · reproducible harness

Turn evidence into action

Every free tool comes out of a study above

No made-up multipliers and no folklore constants — each tool is the published method, made runnable. Free, and only one of the five needs a free account.

All free tools

Free tool

PageSpeed cloaking test

Check whether a site is “cheating” PageSpeed — showing Google a faster page than real visitors get. We compare Google's own PageSpeed Insights view against the page a normal visitor receives, and flag structural mismatches with the evidence.

PSI ground truth · lab-vs-field check · user-agent flip

Grounded in How many sites cheat PageSpeed?

Research standard

How we publish

  • Design before numbers. Whether it's a controlled experiment or an audit of others' studies, the research design is graded first — effect sizes are only trusted in proportion to how they were measured.
  • Raw data included. Every study ships its complete dataset as CSV — the experiment's per-run results, or the audit's full study register — so you can check our math.
  • Limitations stated, not buried. Lab data and observational data are both models of reality. Each study says exactly where the model ends.
  • Reproducible. Methodology sections carry everything needed to re-run the work: versions and throttling profiles for experiments, provenance and grading for audits.

Questions about the data, or a replication to share? Get in touch — we link replications.

The same standard, applied to one page instead of the whole web: read the audit framework.