Experiments / 07

Service Workers

What happens when you put a programmable proxy between the page and the network?

  • Advanced
  • 12 min
  • Impact ●●○

The problem

The HTTP cache (Experiment 06) is a set of rules you hand to the browser. A service worker is the opposite: a small program you write that sits between your pages and the network and decides, request by request, what to do. Answer from a cache. Go to the network. Do both. Show an offline page.

That makes it the most powerful caching tool on the web and one of the easiest to get wrong. It doesn’t make your site faster by itself. It lets you define what “faster” and “works offline” mean, and then it holds you to that choice, including after you deploy something new.

1 · Choosing a strategy

Six visits: a normal one, a second, a flight with no network, a terrible connection, a visit after a deploy, and the day after. Start on “Cache-first everything”, then walk through the other setups.

Setups

Fast and works offline once cached. And never updates.

Strategy for each kind of request
Worker options
3.00 s
Offline visit
works
Visits out of date
2
Visits broken
0
Bad-connection load
0 ms
Six visits
  • network
  • cache
  • stale-while-revalidate
  • timed out → cache
  • failed
  • out of date
For each visit and file: how the service worker answered, with the page state, load time and data
Visitindex.htmlstyles.cssapp.jshero.jpg/api/feedResult
First visitonlinenetnetnetnetnetworks1.00 s · 538 KB
⚙ Worker registered and installed
Second visitonlinenetnetnetnetnetworks1.00 s · 538 KB
On a planeofflinecachecachecachecachecacheworksinstant · 0 KB
Bad connectionvery slowcachecachecachecachecacheworksinstant · 0 KB
▲ deploy: index.html, styles.css and app.js change
Back online, after a deployonlinecachecachecachecachecacheout of dateinstant · 0 KB
⚙ sw.js unchanged: no update, old cache stays
The next dayonlinecachecachecachecachecacheout of dateinstant · 0 KB

The second visit still uses the network. The first page load isn’t controlled by the worker, so nothing was cached. Without a precache the worker only learns from requests it sees after it takes over.

Stale on “back online, after a deploy”. A cache-first worker never looks for a newer copy, and the worker file itself didn’t change, so the cache is never replaced: this visitor is stuck on the old version.

On the bad connection the page is instant, because the cache answers without waiting for the network.

The page still works offline.

The page’s first load is never controlled by its worker. The HTTP cache is set aside; “very slow” is about 0.3 Mbps with 1.5 s latency.

Things to try

  1. “No service worker”. Offline is broken and the bad connection takes tens of seconds. This is your baseline.
  2. “Cache-first everything”. Offline works and the bad connection is instant. After the deploy it’s stuck on the old version.
  3. “Network-first everything”. Always fresh when you’re online, but look at the bad-connection visit. Drag the timeout to 0 and watch it get worse.
  4. Mixed versions. Set documents to network-first and assets to cache-first with fingerprinting off. The new HTML talks to the old JavaScript.
  5. “Sensible mix”. Tick fingerprinting off and on to see what it buys.
  6. “Stale-while-revalidate”. Instant everywhere, always one visit behind, and the bytes add up.

What a service worker is

  • A JavaScript file you register from a page. It runs off the main thread, has no access to the DOM, and wakes up to handle events such as install, activate and fetch.
  • It works only over HTTPS (and on localhost for development), and only controls pages inside its scope, by default the folder it is served from.
  • The page that first registers it isn’t controlled by it. The worker takes over from the next navigation, unless it calls clients.claim(). That’s why the second visit in the widget still uses the network when nothing was precached.
  • Its fetch handler sees every request the controlled page makes, and can answer from anywhere.
// in the page
if ('serviceWorker' in navigator) {
  navigator.serviceWorker.register('/sw.js');
}

The strategies

Network only
Don’t touch it. Right for requests that must never be cached, such as payments or anything personalised and security-sensitive.
Cache first
Instant and works offline. Right for files that never change at their URL, such as fingerprinted assets. Wrong for anything that changes in place: it will never ask again.
Network first
Fresh when you’re online, with the cache as a safety net. Right for documents and data. The cost is waiting on the network before you fall back, so add a timeout.
Stale-while-revalidate
Answer from the cache immediately and refresh it in the background. Right when slightly old is fine. The visitor is always one version behind, and every visit downloads again.

Use a different strategy for different content. Documents and API data want freshness. Fingerprinted assets are immutable, so serve them from the cache. That combination is the “sensible mix” preset.

Precache and runtime cache

Precaching downloads the app shell at install, so the second visit already works offline. Runtime caching stores things as they’re requested. Precache what the whole site needs, and runtime-cache the rest.

One more cost to know about: a worker that isn’t running has to start before it can answer a navigation, which can add delay. Navigation preload lets the browser start the network request while the worker boots.

A sketch of the whole thing (real projects usually use a library such as Workbox rather than hand-rolling it):

const VERSION = 'v5';
const SHELL = ['/', '/styles.3f9a1c.css', '/app.8c2e77.js', '/offline.html'];

self.addEventListener('install', (event) => {
  event.waitUntil(caches.open(VERSION).then((c) => c.addAll(SHELL)));
});

self.addEventListener('activate', (event) => {
  event.waitUntil(
    caches.keys().then((names) =>
      Promise.all(names.filter((n) => n !== VERSION).map((n) => caches.delete(n)))
    )
  );
});

self.addEventListener('fetch', (event) => {
  const req = event.request;
  if (req.method !== 'GET') return;
  if (req.mode === 'navigate') {
    // network first for documents, with an offline fallback
    event.respondWith(fetch(req).catch(() => caches.match('/offline.html')));
    return;
  }
  // cache first for fingerprinted assets
  event.respondWith(caches.match(req).then((hit) => hit ?? fetch(req)));
});

2 · Why won’t my update arrive?

You deployed a new sw.js. Start with two tabs open, pick “Why won’t it update?”, and see why reloading doesn’t do it.

Tabs open on v1
In the new sw.js
Then the visitor
Everyone on the new version?
not yet
Worker v1
active
Worker v2
waiting
What happens, in order
  1. 1
    Deploy

    A new sw.js is on the server. Open tabs are still running version 1.

  2. 2
    Update check

    On the next navigation (or at most daily) the browser fetches sw.js and finds it has changed byte for byte.

  3. 3
    Install

    The new worker (v2) runs its install event: it precaches the new assets. The old worker (v1) keeps serving every tab.

  4. 4
    Waiting

    v2 is installed but waits: it cannot activate while any page is still controlled by v1.

Each tab
  • Tab 1page v1worker v1Running the old version
  • Tab 2page v1worker v1Running the old version

The new worker is waiting. It can’t activate while any page is still controlled by the old one. Reloading a tab doesn’t help: the reloaded page is itself controlled by the old worker, which is still active.

Follows the service worker spec: a waiting worker activates when no client is controlled by the active one.

The update lifecycle

  1. The browser re-fetches sw.js on navigations (and at least daily). If it differs by even one byte, it installs it as a new worker.
  2. The new worker runs install, usually precaching new files. The old worker keeps serving.
  3. The new worker waits. It becomes active only when no page is controlled by the old worker, which means every tab closed. A reload isn’t enough, because the reloaded page is itself controlled by the old worker.
  4. On activate the new worker deletes caches the old version left behind.

skipWaiting() and clients.claim() shorten this, at a cost: you can end up with an old page running under a new worker, and a new worker may have deleted files the old page still needs. They’re safe only when each release is backwards compatible with the one before.

Two rules of thumb: version your caches so activation can clean up, and serve sw.js with no-cache. Modern browsers already skip the HTTP cache when they check the main worker file, but scripts it imports can still be cached, so don’t rely on the default.

Common misconceptions

“A service worker makes my site faster.”

It makes repeat visits and bad connections better, if you configure it well. A first visit gets nothing from it, and a worker that has to start before it can answer a navigation can add a little delay. It’s a tool, not an optimisation on its own.

“Cache-first is the fast option, so use it everywhere.”

Cache-first is the fast option for things that never change at their URL. Use it on a document or API response and you’ve built the stale-forever trap from the first widget.

“I’ll just call skipWaiting() so updates always work.”

It fixes the waiting, and creates version mismatches. Look at the tabs in the second widget with skipWaiting and claim on: old pages under a new worker.

Other things that bite

  • Only cache what you mean to. Cache successful (200) responses, not errors, and be careful with personalised pages and anything behind a login.
  • Caches grow. Browsers cap storage and may evict it under pressure. Don’t assume something you cached is still there.
  • Third-party responses can be “opaque”: you can store them but not inspect them, including whether they succeeded.
  • It adds a layer to debug. A bug that “fixes itself in a private window” is often a worker.

See it on a real page

  1. In DevTools open Application, then Service workers. You’ll see its status (activated, waiting), and the checkboxes Update on reload, Bypass for network and Offline. “Update on reload” is for development: it hides the waiting problem that your users will have.
  2. Open Cache storage to see exactly what the worker has stored, and delete it to reset.
  3. In Network, the Size column shows (ServiceWorker) for responses the worker answered.
  4. Test the offline and bad-connection cases for real, using the throttling and Offline controls, because that’s the situation you built this for.

Model assumptions

What these simulations simplify
  • Five resources (20 KB HTML, 60 KB CSS, 150 KB JS, 300 KB image, 8 KB API data that changes every visit). One deploy before visit 5 changes the HTML, CSS and JS.
  • The HTTP cache is set aside so only the worker’s decisions show. Server responses are always correct.
  • Visit conditions: online is 100 ms and 10 Mbps. “Very slow” is 1.5 s and 0.3 Mbps. Offline has no network.
  • The page’s first load isn’t controlled. A changed sw.js is detected on the first online visit after the deploy. The new worker installs then, and takes over the next visit.
  • Network-first gives up after its timeout only if the transfer would take longer; the request then continues in the background and refreshes the cache.
  • The update lifecycle follows the specification, and doesn’t model browser-specific details such as the 24-hour update check or navigation preload.