INP yield-gaming audit + registered design
Mobile INP improved across the web while lab Total Blocking Time rose 58% in a year.
Some of that INP win is real work reduction; some is yielding — repainting sooner
without doing less, which the metric can't tell apart. We define the line outcome-wise, register a
three-study design over CrUX + HTTP Archive + lab replay, and ship the kit. No fabricated share.
Interaction anatomy · outcome classification · four-source triangulation · replication kit
INP-era framework + taxonomy
A green Lighthouse score isn't a Core Web Vitals pass — 43% of 90+ pages failed a CWV
even in the FID era, and INP made the responsiveness gap harder to see in the lab. We reframe it as
a transportability problem: four gates, a falsifiable 12-class divergence taxonomy,
and a calibrated reconciliation protocol over public CrUX + HTTP Archive. No fabricated coefficients.
Lab×field confusion matrix · transportability gates · reproducibility kit
Power analysis · tool + protocol
“We ran it five times and the score went up two points” usually isn't evidence. Median-of-5
stabilises an estimate but can't prove a small before/after change. We derive the power tables —
at a typical SD of 5 points, a 5-point change needs ~16 runs and a 3-point change
~44 — ship a free required-runs calculator, and register the variance census.
Analytic CI/power tables · Heričko secondary evidence · required-runs calculator
Model + registered benchmark
SPAs trade a heavier first load for faster in-app navigations. Chrome's soft-navigation API
(147–149) finally makes that bargain measurable. Using real RUM-Archive ratios, the break-even is
unforgiving: at ~1 follow-on navigation per session, a 1,000 ms first-load penalty needs
~1,015 ms desktop / ~1,143 ms mobile saved per navigation just
to break even. Source-derived model + the benchmark to close it — no fabricated lab numbers.
Break-even model · amortization grid · pre-frozen 50-pair benchmark registry
Registered protocol · adoption-cohort
WP Rocket, NitroPack and LiteSpeed promise to fix Core Web Vitals — but the evidence offered is a
Lighthouse score or a cross-sectional dashboard, neither of which shows what happens to real
users after a site adopts one. We publish the full causal adoption-cohort event study
(HTTP Archive × CrUX), including the delay-JavaScript INP risk, and a runnable replication
package — with no fabricated effect estimates.
Staggered DiD (Callaway–Sant'Anna) · delay-JS INP cohort · SQL + R + Python pack
Cloaking census · pilot + protocol
Some sites serve Lighthouse a stripped, fast page while real visitors get the heavy one. The
techniques are documented; the prevalence never has been. We ran a real user-agent-flip
pilot across the Tranco top 2,000 — 0 confirmed in 1,162 measured (a lower bound
on the crudest cloaking; a naïve detector's 9 flags were all benign) — and publish the full
preregistered census protocol. No fabricated number.
Tranco-frame pilot · A/A-controlled · replication pack + protocol
Evidence audit · 2026
Twenty years of speed-versus-revenue studies, graded by research design. The headline everyone
quotes — “a few percent per 100 ms” — sits roughly an order of magnitude above the best
causal estimate of ≈0.65% per 100 ms. Speed matters; the marketed ROI is
systematically overstated.
10 studies, design-graded · study register (CSV) · reconciled causal anchor
Registered study · pilot
Lighthouse prints “Est savings: 1,200 ms.” Does it show up when you do the fix? We ran a
controlled pilot — predicted vs realized for Lighthouse 13's insights — and publish the full
protocol to measure it at web scale. On clean pages the savings materialized, and were
conservative (realized 1.0–1.9× predicted).
8-case pilot · calibration figures · open protocol + workbook
Causal decomposition · 2026
CrUX's “good Core Web Vitals” share rose ~13 points since 2023 — but the metric and the
instrument both changed mid-stream. The metric-comparable gain is +10.1 points, not
+12.9, and Google's own release notes prove a real share is measurement, browser and
data-quality change, not sites optimizing. With an event ledger and the BigQuery decomposition design.
41 months of CrUX data · event ledger · SQL + R reproducibility pack
Study · 2026
210 controlled Lighthouse runs on synthetic pages that vary one thing at a time:
JavaScript size and placement, main-thread work, render-blocking CSS, hero image
weight, webfont strategy, and request-chain depth. The headline: on the wire,
every critical-path kilobyte costs the same ~0.5 s per 100 KB — placement,
execution and round trips are where the real money is.
Full methodology · raw dataset (CSV) · reproducible harness
Free tool
How many Lighthouse runs do you actually need to prove a before/after change? Enter your
run-level spread and target change; get the runs per condition, a precision sweep, and a
red/green read on whether median-of-5 is enough for your claim.
Confidence + power · paired or independent · CI sweep
Free tool
Turn an LCP target and a connection profile into per-resource byte budgets —
modeled on TCP slow start instead of naive bandwidth division — and export a
ready-to-use Lighthouse budget.json for CI.
RFC 6928 slow-start model · Lighthouse/WebPageTest profiles
Free tool
Check whether a site is “cheating” PageSpeed — showing Google a faster page than real
visitors get. We compare Google's own PageSpeed Insights view against the page a normal
visitor receives, and flag structural mismatches with the evidence.
PSI ground truth · lab-vs-field check · user-agent flip