Experiments / 12

Lab vs Field

Why does my 95 score not match what users feel?

  • Intermediate
  • 8 min
  • Impact ●●●

The problem

Every other experiment gave you a number: “load time 1.6 s”. It’s one number, for one device, on one connection, on one visit. A real site doesn’t have a load time. It has a distribution of them: a lot of fast visits, some slow ones, and a long tail of very slow ones.

That’s why a page can score 95 in Lighthouse and still be miserable for a quarter of its visitors. Neither number is wrong. They answer different questions: lab data answers “what happens under these exact conditions?”, field data answers “what happens to the people who actually visit?”.

Three thousand visits, and where the lab numbers fall

Start on “A typical mix”. Find the two dashed lab lines and the solid 75th-percentile line. Then change who visits.

Metric
Who visits
30%
20%

Mid-range phones: 50%

25%
35%
The page
600 KB
400 ms
40%
Fixes
Field, 75th percentile
3.03 sfails · threshold 2.50 s
Median
2.11 smean 2.57 s
95th percentile
6.61 sthe worst 1 in 20
Lab: Lighthouse mobile
5.15 sone run
3,000 visits

63% good · 21% needs improvement · 16% poor

75th percentile by audience
  • Desktop 30%1.44 s
  • Mid-range phone 50%2.91 s
  • Low-end phone 20%4.98 s
  • Fast (Wi-Fi/5G) 25%1.81 s
  • Average 4G 52%2.64 s
  • Slow 24%6.40 s

LCP at the 75th percentile is 3.03 s: needs improvement. This metric fails the Core Web Vitals check, which asks for the 75th percentile to be “good” (2.50 s or better). 63% of visits were good, 21% needed improvement and 16% were poor.

The lab numbers disagree with each other and with the field. Your laptop says 492 ms (good). Lighthouse’s mobile preset says 5.15 s (poor). The field’s 75th percentile is 3.03 s. Each is a true answer to a different question.

The mean (2.57 s) sits below the 75th percentile and well below the 95th (6.61 s), because a few very slow visits are outweighed by many fast ones. Averages hide the people the metric exists to protect.

Who pays: low-end phones see a 75th percentile of 4.98 s, against 1.44 s on desktop (3.5× worse). A page tested on a desktop never meets these visitors.

A simulated population, not data from any real site. Lab runs: your laptop = 1× CPU, 100 Mbps, cached; Lighthouse mobile = 4× CPU, 150 ms, 1.6 Mbps, cold.

Things to try

  1. “The developer’s world”. Mostly desktops on fast connections with a warm cache. The page looks fine, the laptop line sits well inside “good”, and you’d ship it.
  2. “A mobile-first, emerging-market audience”. Same page, same code. The distribution slides right, and the 75th percentile fails. The lab lines didn’t move at all.
  3. Compare the median to the 95th percentile. The worst one visit in twenty is many times slower than the typical one.
  4. Switch the metric to INP, then tick “Ship half the JavaScript”. Look at which row of the audience table gained the most.
  5. Switch to CLS. Most visits have no shift at all, so the median is zero. The problem is entirely in the tail.
  6. Apply every fix to the emerging-market audience, and see how far it gets. Then compare it with the “developer’s world” with no fixes.

How the field is judged

  • The 75th percentile. Three in four visits are at least this good. Core Web Vitals check the 75th percentile of each metric, not the average, because the average hides the slow visits the metric is there to protect.
  • Three thresholds. LCP: good up to 2.5 s, poor above 4 s. INP: good up to 200 ms, poor above 500 ms. CLS: good up to 0.1, poor above 0.25. A page passes when each metric’s 75th percentile is good.
  • Where the data comes from. The Chrome UX Report collects it from real Chrome users who’ve opted in, over a rolling 28 days. PageSpeed Insights shows that field data at the top, and a Lighthouse lab run below it.
  • Your own measurement. The web-vitals library reports each metric from the visitor’s browser, so you can segment by device, country and page:
import { onLCP, onINP, onCLS } from 'web-vitals';

function send(metric) {
  navigator.sendBeacon('/vitals', JSON.stringify({ name: metric.name, value: metric.value, id: metric.id }));
}
onLCP(send);
onINP(send);
onCLS(send);

sendBeacon is used because the final values are often only known when the visitor leaves the page.

Why the lab and the field differ

Who is being measured
Lighthouse is one synthetic visitor on a fixed device profile. The field is everyone, including the low-end phones, slow connections and cold caches you didn’t think of.
What happens after load
A lab run loads the page and stops. Real people scroll, tap and stay. INP and CLS are measured across the whole visit, and a lab run sees only the start. This is why Lighthouse reports Total Blocking Time as a lab stand-in for INP.
Variation
The same lab run, repeated, gives slightly different numbers. The field’s variation is enormous, and that spread is the information.
State
Lab runs are usually logged-out, uncached and without ads, consent banners or personalisation. Real visits are not.

Use both, for different jobs. Field data tells you whether you have a problem and for whom. Lab data lets you reproduce it and test a fix, because it’s repeatable and you can control it. A good workflow: find the worst segment in the field, recreate that device and network in the lab, fix it, then confirm in the field, which takes a few weeks to show up.

Common misconceptions

“A high Lighthouse score means my website is fast.”

It means one simulated visit was fast. The score depends on the profile Lighthouse chose, and on whether it saw your real tags, banners and logged-in state. It says nothing about the 75th percentile of your visitors, or about what happens after load. A good score is a useful signal, and the field is the verdict.

“Our average load time is 1.8 seconds.”

The average describes nobody. The visits are lopsided: many fast ones and a tail of very slow ones. The median and the 75th and 95th percentiles describe what people experience.

“Field data is noisy, so lab data is more reliable.”

Lab data is repeatable, which isn’t the same as representative. It reliably measures the wrong thing if the conditions don’t match your visitors. Field noise is real people.

“Optimise for the typical user.”

The typical user is the one you’re already serving well. Look at the table beneath the histogram: the same fix does far more for the low-end phone than for the desktop. Performance work pays off at the tail.

Don’t memorise rules. Understand the system.

Twelve experiments, one idea. Each technique here exists because of something the browser does: it can only request what it has discovered, it has one main thread, it can’t paint without the CSS, a cache only helps if the files stay unchanged. Once you know the mechanism, the advice stops being a checklist. You can tell when a technique applies, what it costs, and how to check it worked.

Measure before optimising. Then measure the people who actually use your site.

See it on a real site

  1. Enter your URL in PageSpeed Insights. Read the field-data section first. Does the page pass? Which metric fails? Then compare it with the lab score underneath.
  2. Add the web-vitals library to the site and send the values to your analytics, with the device type and connection.
  3. In DevTools, set CPU throttling and network throttling to match your worst common visitor, not your best, and reproduce the problem there.
  4. After a fix, wait for the 28-day field window to turn over before declaring victory.

Model assumptions

What this simulation simplifies
  • The population is generated, not measured: visits get a device (desktop 1×, mid-range phone 4×, low-end phone 6× CPU), a connection (fast, average or slow) and a cache state, plus random variation. Slow connections are more likely on phones.
  • LCP is server and connection time, plus download time (a quarter of the bytes on a warm cache), plus startup JavaScript and rendering, scaled by the CPU. INP follows startup JavaScript and the CPU, with a heavy tail. CLS is zero for most visits and log-normal for the rest.
  • The same random numbers are used for every setting, so changing a slider or a fix moves the same visits instead of drawing a new sample.
  • Lab runs: “your laptop” is a 1× CPU, 100 Mbps, warm cache. “Lighthouse mobile” is 4× CPU, 150 ms round trip, 1.6 Mbps, cold cache. Neither has random variation.
  • Thresholds are the real Core Web Vitals ones. Everything else is illustrative and not calibrated to any real site or to the Chrome UX Report.