We have just open-sourced the engine NorthDuty uses to answer that question. It is on GitHub as northduty-health-check under the MIT licence, and you can run it against any URL from your terminal. This post explains what it checks and the three design decisions that shaped it. Each one came from a mistake that is easy to make when you build monitoring.

What a health check measures

A single check produces one JSON result covering:

  • Availability: the HTTP status code, and whether the site redirected somewhere it shouldn't, such as a parked domain or a spam site.
  • Security basics: SSL certificate validity and days to expiry, domain registration expiry, and security headers.
  • DNS: whether the hostname resolves, and to which address.
  • Speed: DNS lookup, TCP connect, TLS handshake, time to first byte, full page load, and the Core Web Vitals FCP, LCP and CLS.
  • Rendering: whether the page is visibly blank, which scripts, styles and images failed to load, which JavaScript errors were thrown, and every XHR/fetch call the page made, with its status.
  • Bot protection: whether Cloudflare, Akamai, DataDome, PerimeterX or Imperva served a challenge instead of the real page.

That list is the easy part. The hard part is deciding when to collect each piece, and what to report when you didn't.

Decision 1: don't launch a browser for every check

Half of that list needs a real browser. You can't measure LCP, catch a JavaScript error or see a blank page without rendering the page in Chromium. So the obvious design is to open a browser for every check.

It doesn't hold up. A full render takes 10–20 seconds of a CPU core and downloads the whole page, often 2–3 MB. At a one-minute interval that is 1,440 page loads a day for a single site. That is enough to get your monitor blocked by the site's own firewall, and enough compute that the bill outgrows the plan price.

Most of what you need every minute doesn't involve rendering at all. "Is the site up, what did it answer, is it still on its own domain, is the certificate valid" is one HTTPS request. So the engine has two tiers:

Uptime tierFull tier
BrowserNo, one HTTPS requestYes, Chromium via Playwright
Typical cost200–500 ms, mostly waiting on the network10–20 s of one CPU core
Data pulled from the siteHeaders and the first 64 KB of HTMLThe whole page
MeasuresStatus, headers, redirects, TTFB, SSL, DNS, challenge detectionAll of that, plus load time, Web Vitals, blank pages, failed resources, JS errors, API calls, accessibility and SEO

In NorthDuty, the uptime tier runs on every scheduled check, and the full tier runs on a project's first check, at most once an hour after that, and immediately after a deploy is detected. A deploy is exactly when SSL, redirects and Core Web Vitals are most likely to break, and the uptime tier sees none of the rendering problems.

One detail matters here. The uptime tier still classifies firewall and bot challenges, using the same detection functions as the browser, fed markers parsed from the raw HTML. Without that, a Cloudflare "checking your browser" page would be recorded as a healthy 200 every minute.

Decision 2: "not measured" must never look like "no errors"

This is the bug we were most careful about, because it is so easy to ship.

An uptime-tier check doesn't render the page, so it can't know whether the page threw JavaScript errors. If it reports jsErrors: [], an empty list, every dashboard downstream will show a green "No errors". That's a false all-clear, shown on the exact checks where nobody looked.

So the uptime tier reports null, never 0 and never an empty list, for everything only a render can produce: load time, paint metrics, blank-page detection, failed resources, JavaScript errors and API calls. Each result also records which tier ran (checksTier). In the NorthDuty app those fields show as Not measured, and the health score leaves out what an uptime-only check never tested, instead of quietly scoring it as perfect.

If you build your own monitoring, this is the one rule worth copying: keep "we looked and found nothing" and "we never looked" as different values all the way to the screen.

Decision 3: one browser at a time is faster than many

Page timings are only meaningful if the machine measuring them isn't the bottleneck. After the first byte arrives, a browser check is CPU-bound: parsing, layout, script execution. Run five renders at once on a four-core machine and you don't get five fast checks. You get five slow ones, and the slowness is your monitor's, not the site's.

So each browser check takes a host-wide slot before launching, and by default only one renders at a time. If no slot frees up within 90 seconds, the check runs anyway and flags it in the result: a noisy timing is better than a gap in the uptime record. Uptime-tier checks take no slot, because they spend their time waiting on the network.

The engine also benchmarks the machine it runs on and adjusts CPU-bound timings to a reference desktop device. If the host is too slow for its timings to describe the site, the result marks them as untrustworthy instead of reporting them as the site's problem.

Safety: a monitor shouldn't fetch whatever it's told

A health checker fetches URLs that users type in, which makes it a classic server-side request forgery risk. Point it at 169.254.169.254 and a naive checker will happily read a cloud server's metadata. The engine refuses private, loopback, link-local and cloud-metadata addresses. It also checks where a public hostname actually resolves, so a domain pointed at an internal IP is blocked too.

Try it on your own site

You need Node.js 20 or newer:

  • Clone the repository, then run npm install and npm run install:browsers (this downloads Chromium once).
  • node index.js https://your-site.com runs a full check and prints the JSON result.
  • node index.js https://your-site.com --uptime runs the fast, browserless tier.
  • node index.js https://your-site.com --watch re-checks every 60 seconds until you stop it.

The repository also includes an SQS queue worker and a Dockerfile if you want to run checks at scale, and a field-by-field guide to what each result means and when it should trigger an alert.

What a health check can't tell you

A health check answers "does this page load and render properly?" It can't answer "can a customer still buy something?" A checkout can fail at the payment step while every page involved loads in under a second with no errors. A contact form can render perfectly and silently stop sending.

That's the part NorthDuty adds on top of this engine: real-browser journeys that add to cart, check out, log in and submit forms the way a customer would, on a schedule and after every deploy. The health check tells you the site is up. The journeys tell you it still works.