The 2026 World Cup in air traffic · Observed State Index
Summary
In July 2026 this system published a clean result: during the World Cup, all eight host cities raised their air traffic and all five control cities lowered theirs, without a single exception, at a probability of 1 in 1,287 of happening by chance.
This document is the audit of that result, carried out six weeks later. The conclusion is that the published table cannot be reproduced from the data, because of two defects that piled up on exactly the period of the tournament: the source lost two thirds of its captures, and the clock those captures were classified by was wrong.
Recalculated with the real capture hour, something does hold: all five controls fall and five of the eight hosts rise. But the perfect separation disappears — Dallas and Atlanta land below three controls — and the p-value moves from 0.00078 to a range of 0.011 to 0.045 depending on how the aggregation is done.
The finding, then, was not false. It was far less certain than was claimed, and it was claimed with a figure that was never computed on these data.
What was claimed
The original study compared the tournament period (11 June – 19 July 2026) against the six preceding weeks, across thirteen air corridors: eight in host cities and five chosen for sharing neither region nor time zone with them.
It reported Vancouver +26 points, New York +15, Los Angeles +15, Chicago +14, and the other four hosts positive; Sydney −30, Bangkok −20, Singapore −20, and the other two controls negative. From that, a Mann-Whitney test at p = 0.00078.
Why it was looked at again
Because of a note in the system's own documentation: between 6 June and 26 August 2026, the air traffic capture process was timing out halfway through the list of 30 nodes, dying, retrying and starting again from the beginning, dying at the same point.
The tournament period falls entirely inside that window. The comparison period falls almost entirely outside it:
| Window | Dates | Days affected |
|---|---|---|
| Six preceding weeks | 30 Apr → 10 Jun | 5 of 42 |
| Tournament | 11 Jun → 19 Jul | 39 of 39 |
Comparing those two windows is comparing two different data qualities. That alone justified the review; what turned up on doing it was worse.
First defect: coverage collapsed
Each node's identifier carries a numeric prefix (001_ams, 002_atl, … 030_yvr), so alphabetical order is the order in which the process went through them. The loss follows that position to the letter:
| Node | Pos. | Captures before | Captures during | Days (of 39) |
|---|---|---|---|---|
| atl_atlanta | 002 | 247 | 217 | 39 |
| bkk_bangkok | 003 | 235 | 213 | 39 |
| dfw_dallas | 008 | 203 | 131 | 39 |
| gru_sao_paulo | 012 | 178 | 69 | 18 |
| jfk_new_york | 016 | 160 | 48 | 14 |
| lax_los_angeles | 018 | 151 | 39 | 14 |
| ord_chicago | 023 | 140 | 30 | 13 |
| syd_sydney | 028 | 127 | 30 | 13 |
| yvr_vancouver | 030 | 124 | 37 | 14 |
Nine of the thirteen nodes have data for between 13 and 18 days of the 39 the tournament lasted. And which days survived was decided by nothing to do with the World Cup: it was decided by how far the process got before it died.
A tournament measured on a third of its days is not badly measured. It is unmeasured on two thirds.
Second defect: the clock was wrong
The system compares each capture against the history of that same time band, and at the time the identity of the band was derived from the run's sequence number, not from the real hour of the capture. That defect is documented separately; what had not been measured is how much it affected this particular period.
It can be measured, by crossing the run number against the real hour carried in the raw data. During the tournament:
| Run | band 1 | band 2 | band 3 | band 4 |
|---|---|---|---|---|
| 01 | 39 | 0 | 0 | 0 |
| 02 | 24 | 15 | 0 | 0 |
| 03 | 24 | 0 | 15 | 0 |
| 04 | 0 | 24 | 3 | 12 |
| 05 | 0 | 24 | 2 | 1 |
| 06 | 0 | 24 | 0 | 2 |
| 07 | 0 | 0 | 22 | 3 |
| 08 | 0 | 0 | 22 | 2 |
| 09 | 0 | 0 | 22 | 1 |
| 10 | 0 | 0 | 0 | 22 |
| 11 | 0 | 0 | 0 | 22 |
| 12 | 0 | 0 | 0 | 22 |
The real pattern is one of triples per band: runs 1 to 3 are band 1, runs 4 to 6 are band 2, runs 7 to 9 band 3 and runs 10 to 12 band 4 — because each band was retried three times before giving up.
A calculation assuming "run 3 = band 3" is right 15 times out of 39. "Run 4 = band 4", 12 out of 39. And runs 5 through 12 have no band at all that corresponds to them under that assumption.
In total, in the tournament window 23% of captures would have received the correct label. In the comparison window, 45%. Two badly labelled periods, with different error patterns, compared against each other.
The attempt to reproduce the table
The original study says it compares "the average of daily counts per node". That phrase admits several readings, so four were tried, all on the data already rebuilt with the real hour:
| Method | Does it reproduce the table? |
|---|---|
| Mean per capture, % change | No — the sign agrees on 6 of 13, worse than chance |
| Daily sum, % change | No |
| Mean per capture, absolute difference | Partly — Vancouver gives +25.2 against the +26 published, but Bangkok comes out with the opposite sign |
| Band against band | The closest: the sign agrees on 11 of 13, but the magnitudes are nothing alike |
None reproduces the published table. And one detail rules out the most convenient explanation: the naive methods put three host cities at the bottom — Atlanta −45.6%, Dallas −51.4%, Denver −35.5% — because with bands 3 and 4 lost above all (14:30 and 19:30 UTC, the busiest and therefore the heaviest to download), the tournament window ended up dominated by the small-hours bands, which in North America are the emptiest.
That is to say: the coverage artefact could not have manufactured the published result. It would have done the opposite: destroyed it. What remained to explain, then, was where that table came from, and the wrong clock is the only cause identified that is capable of producing an arbitrary ordering across these nodes.
The recalculation, with the real hour
Anchoring each capture to its real hour and comparing band against band — which is how the system operates today — with two ways of aggregating each node's four bands:
| Node | Group | Mean of % | Balanced panel |
|---|---|---|---|
| yvr_vancouver | host | +9.9% | +10.6% |
| den_denver | host | +2.2% | +2.6% |
| mex_mexico_city | host | −0.6% | +1.3% |
| lax_los_angeles | host | +2.9% | +1.2% |
| ord_chicago | host | +3.2% | 0.0% |
| jfk_new_york | host | +0.7% | −0.8% |
| sin_singapore | control | −3.7% | −5.2% |
| bkk_bangkok | control | −4.1% | −6.0% |
| gru_sao_paulo | control | −3.1% | −6.7% |
| dfw_dallas | host | −4.8% | −7.0% |
| atl_atlanta | host | +2.2% | −7.7% |
| jnb_johannesburg | control | −11.3% | −8.8% |
| syd_sydney | control | −7.2% | −12.8% |
The two columns are not cosmetic alternatives. The first averages the percentage changes of the four bands, which gives Atlanta's band 2 — a mean of 30 aircraft — the same weight as its band 3, which has 832. The second sums each window's four band means and compares them, which amounts to asking how much a full day's cycle changed if sampled once per band. The second is the defensible one, and it is what moves Atlanta from +2.2% to −7.7%.
The p-values, computed by exact enumeration of the 1,287 possible partitions rather than by normal approximation:
| Method | U | p (1-tail) | p (2-tail) | Perfect separation |
|---|---|---|---|---|
| Mean of % | 3 | 0.0054 | 0.011 | No |
| Balanced panel | 6 | 0.0225 | 0.045 | No |
| Published | 0 | — | 0.00078 | Yes |
And an observation about that last row: 0.00078 is exactly 1/1,287, the probability of the single most extreme arrangement out of the 1,287. It is not a statistic computed on data: it is the p-value corresponding to perfect separation, whatever the magnitude of the numbers. Once the perfect separation is gone, that figure has nothing left to refer to.
What holds and what does not
It holds: all five controls fell, under both ways of aggregating. Five of the eight hosts rose. The difference between the two groups is still statistically significant at 5% under both estimators, and the largest magnitude is still Vancouver's.
There is one point in the finding's favour worth stating, because it cuts against this document's conclusion: if the process died more readily on the heaviest captures, the days that survived are the quietest ones, and that bias pushes every node downward. The real gap between hosts and controls could be understated rather than inflated.
It does not hold: the perfect separation, the "without a single exception", the 1 in 1,287 figure and the published magnitudes. Dallas and Atlanta end up below three controls.
It cannot be known: where exactly the original table came from. The wrong clock is the most likely cause and the only one identified, but the calculation that produced it has not been reconstructed.
And there is a limit no correction repairs: nine of the thirteen nodes have a third of the tournament measured. That data does not exist and cannot be recovered; the tournament ended and the defect was fixed afterwards. Any future version of this analysis starts from there.
What this case shows
That a clean result deserves more suspicion than a messy one, not less. The perfect separation across thirteen nodes was the reason this finding looked solid, and it was in fact the signal that it was worth looking at twice.
That two defects already documented separately — the lost captures and the band derived from the run number — had not been crossed against the analyses that depended on them. Each had its note; no note said which published conclusions fell inside its period.
That auditing is cheap when a system keeps its raw data. Everything in this document came from four queries against the same storage that feeds the site, with no new instrumentation and without asking the source for a single datum again.
That correcting downward is part of the method. A system that only publishes what confirms its findings is not a measuring instrument: it is an argument.
Methodological note
The Gulf airspace closure finding is not affected by any of this, and that was checked rather than assumed. In its period (1 January – 5 March 2026) the run number did correspond to the real band: run 1 was right 64 times out of 64, and runs 2, 3 and 4 were right 59 out of 64 each. Retries existed, but they were occasional — 28 extra runs in 64 days, against 191 in the tournament window.
What is affected is the analysis of the recovery after that event, which spans March to August 2026 and therefore falls inside the degraded period. That is why it is not published: it is pending recalculation.
Reproducibility
The data are the same that feed the site: adsb.lol captures stored in a columnar format, queried with standard SQL. The four queries behind this audit are the coverage table by node and window, the band distribution by node and window, the crossing of run number against real hour, and the means by node and band. The aggregation afterwards and the exact test were done on those means.
← Back to the index