Skip to content
WanderNerd

Methodology

How we calculate this stuff.

Every number on WanderNerd comes from a documented, deterministic calculation over official records. Here is the whole recipe, minus the jargon.

Where the data comes from

We use the U.S. Department of Transportation's on-time performance records, published by the Bureau of Transportation Statistics. Larger U.S. airlines are required to report these records, so the data covers major carriers and their regional partners on domestic routes. Small commuter airlines, charters, and international-only flights aren't in it.

BTS publishes each month with a lag of about two to three months. Every page shows the months behind its numbers. Right now that's June 2025 – May 2026, so you'll see “data through May 2026” around the site.

We use marketing flight numbers where BTS provides them. That's the number on your ticket, even when a regional partner operates the plane.

What "on time" means

A flight counts as on time when it arrives within 15 minutes of schedule. That's the U.S. DOT convention, and it's why our percentages say “arrive within 15 minutes.” Departure punctuality uses the same 15-minute rule against scheduled departure.

Cancelled and diverted flights are not hidden inside the on-time percentage. They're reported separately on every page, because a cancellation is not a delay, it's a much worse day.

The WanderNerd Reliability Score

The score is one deterministic formula, 0–100, applied identically to every flight, route, and airline:

score = on-time% + 8 − 4×cancelled% − 4×diverted% − 0.5×max(0, avg delay − 10)

In words: the on-time arrival rate is the backbone. We add a small constant because “on time” already includes a 15-minute grace period. Cancellations and diversions are penalized four-to-one because they ruin plans in a way a 20-minute delay doesn't. Chronic average delays beyond 10 minutes shave off the rest. The result is clamped to 0–100 and rounded. No decimals, because a decimal would imply precision the data doesn't have.

Score bands and their labels
ScoreLabelTranslation
90–100Nerd ApprovedSet your watch by it.
80–89Pretty ReliableHonestly, not bad.
65–79Usually FineFine. Usually.
50–64A Little DiceyBuild in a buffer.
0–49Bring SnacksAnd maybe a hotel option.

Sample sizes (or: why some pages have no score)

We publish a score once a page has at least 30 scheduled flights in the data. Below that we still show the raw history, clearly labeled, because eight flights of history is an anecdote, not a statistic.

Ranking has a higher bar. Lists like “who flies it best” or “best historical flights” only include entries with at least 100 flights, so eleven lucky departures can never outrank 250 solid ones. Anything under the bar is shown separately, unranked, with its sample size. Ties break by sample size, then alphabetically, so the order is stable.

Observations like “Friday is rough” have their own rules: each weekday or time-of-day bucket needs a minimum sample, and the wording is matched to the size of the gap. Differences under 3 percentage points get called small, because they are. Stronger phrasing is reserved for gaps above 7 and 15 points.

Flight numbers that switch routes

Airlines reuse flight numbers, and a number can fly Seattle to Chicago in June and Los Angeles to Boise in November. Blending those into one “route” would be nonsense, so every flight number gets a consistency check:

  • If at least 80% of its flights flew one route, we present that route and disclose the rest, which stay included in the totals.
  • Below that, the page says plainly that the number covered several routes, lists them with their share of flights, and links each route's own page for clean numbers. No score, no comparisons, and search engines are told not to index it.

What this can and can't tell you

Historical reliability is probability, not prophecy. A flight that landed on time 90% of the time is a genuinely good bet, and can still be three hours late tomorrow. In particular:

  • Schedules change. A flight number can move to a different time, plane, or route.
  • Seasons matter. Last winter's storms are in the data; next winter's aren't.
  • Averages hide variance. Two flights can share an average delay while one is metronomic and the other swings wildly. That's why we show medians and distributions.
  • We show history, not live status. For today's gate and delay, check your airline.

Every observation on a page is computed from that page's numbers with the sample and effect-size rules above. Nothing is generated filler.

Questions about the numbers? Tell us. Nerds love mail about statistics.