Synthetic Industry

Troubleshooting guide · updated 2026-10-11

One slow endpoint or page: measure the distribution before you change any code

Build a measurement two people would agree on, read percentiles instead of averages, separate lab from field figures and decide what a believable before and after comparison has to show.

Agree what slow means in numbers

Slow is a feeling until it has numbers, and two people measuring differently will never agree on whether a change helped. Name the target exactly: one endpoint with a fixed request, or one page with a fixed starting state. Decide where it will run: a non-production copy with data of a realistic size, because many slow queries are slow only when the tables are large. Decide how many times it will run and what is discarded: the first runs often include cold caches and one-off setup, so warm it up and throw those results away. Write the procedure down so that a second person can follow it and get the same kind of result.

  • Fix the request, the data, the machine and the number of runs, and write them down.
  • Keep the procedure in a script, not in someone's memory.

Read the distribution, not the average

Google's site reliability book makes the point clearly: averages hide outliers. A service averaging 100 milliseconds can still take 5 seconds for one request in a hundred. When a page waits on several backends, the 99th percentile of one backend can become the median response of the page. So record each run, report the median and a high percentile such as the 95th, and look at the spread. The same book recommends keeping the latency of successful requests apart from the latency of failed ones, because a failure that returns at once can make the overall figure look better than it is. A slow error is worse than a fast one, so track it, do not drop it.

  • Report the median and the 95th percentile together, with the number of runs.
  • Keep failed requests in a separate column rather than averaging them in.

Know how much noise a single run carries

Even with nothing changed, repeated measurements move. Lighthouse's own documentation lists the sources: nondeterminism in the page itself, the local network, client hardware and other work on the same machine. Its advice is to use aggregates such as the median instead of single results, and it states that the median of five runs is twice as stable as one run. It also says not to run several instances at once on one machine, because they compete. Before you change any code, take the baseline twice on separate occasions. If the two disagree by more than the improvement you hope to see, the measurement cannot yet tell you anything and needs more runs or a quieter machine.

  • Take the baseline twice, on separate occasions, before changing anything.
  • Do not run two measurements at the same time on one machine.

Lab and field answer different questions, and how the paid job is accepted

web.dev explains that lab data comes from a preset device and network and is meant to be repeatable, while field data comes from real visits and is a distribution, usually reported at the 75th percentile. They differ because of device, network, cache state and what real users do, and the article says to use field data to prioritise and lab data to debug and check before release. For web pages, the Web Vitals good thresholds are 2.5 seconds for Largest Contentful Paint, 200 milliseconds for Interaction to Next Paint and 0.1 for Cumulative Layout Shift, at the 75th percentile of page loads. The measurement job here is a lab measurement of the server's response time for one endpoint or server-rendered page on a non-production copy; improving one of those browser-side Web Vitals on a page is a separate job. Acceptance is two agreeing baselines, an after measurement taken the same way, unchanged responses for the request set and, if a target was agreed in writing, a plain statement of whether it was met. A speed-up is not promised. Start an enquiry with the path and the stack in general terms; send no code or data.

Sources and limits

  • Google SRE book: monitoring distributed systems Checked 2026-10-11.
    • Averages hide outliers: a service averaging 100 ms can still take 5 seconds for 1 percent of requests.
    • Recording request counts bucketed by latency shows the distribution, and the latency of successful requests should be kept apart from the latency of failed requests.
    • When a page depends on several backends, the 99th percentile of one backend can become the median response of the frontend.
  • Lighthouse: variability Checked 2026-10-11.
    • Scores can change when the code has not, from page nondeterminism, network, hardware and resource contention.
    • Aggregates such as the median should be used instead of single results, and the median of five runs is described as twice as stable as one run.
    • Several Lighthouse instances should not run at once on one machine.
  • web.dev: why lab and field data can be different Checked 2026-10-11.
    • Lab data comes from a preset device and network and aims to be repeatable; field data comes from real visits and is a distribution, reported as the 75th percentile.
    • Field data is what to use to prioritise effort, while lab data suits debugging and checks before release.
  • web.dev: Web Vitals Checked 2026-10-11.
    • The good thresholds are 2.5 seconds for Largest Contentful Paint, 200 milliseconds for Interaction to Next Paint and 0.1 for Cumulative Layout Shift, measured at the 75th percentile of page loads.