CAT researchOverviewEURUSDGBPUSDXAUUSDGBPCADEURCADAUDJPYGBPJPYResearch notes

Result — the GBP-leg test FAILED. Closure stands.

Criteria pre-registered in zms_gbp_leg_prereg_2026-08-27.md, committed in 2f0d0c3 before any run. Raw output: zms_gbp_leg_results_2026-08-27.json. Nothing in the criteria was edited after seeing these numbers.

Verdict

criterion outcome
C5 GBPUSD still passes at its real 11-tick spread HOLDS — PF 1.848, exPF 1.801. Test is live, not void.
C1 GBPCAD passes all eight HOLDS — PF 1.614, exPF 1.580, median month +79.5, 55.0% profitable, ex-top-3 +5,500, best month 16.1%.
C2 same carrying cell as GBPUSD, n ≥ 30, PF > 1 FAILS — GBPCAD carries on H4P1 (n=162, PF 1.924, +5,034); GBPUSD carries on H1P1 (n=205, PF 2.029, +2,823). Different cells.
C3 neither control passes holds by the letter — see below, it should not comfort anyone
C4 GBPCAD passes at all four H4 offsets HOLDS — passes at 0 / 3,600 / 7,200 / 10,800s.

Decision rule as written: "C2 or C4 fails → No consistent mechanism carries it. Closure stands." C2 failed. The test is a fail and ZMS remains closed.

The finding that matters more than C2

C3 was written as "neither EURCAD nor AUDJPY passes all eight", evaluated at the canonical 3,600s offset. At that offset neither does, so C3 technically holds. But the offset sweep run for C4 was applied to the controls too, and:

H4 offset GBPCAD (treatment) AUDJPY (control) EURCAD (control)
0s PASS PASS fail
3,600s PASS fail fail
7,200s PASS PASS fail
10,800s PASS PASS fail

AUDJPY — no GBP leg, no USD leg — passes all eight falsifiers at three of the four H4 grid offsets. It misses the canonical one on a single criterion (both halves PF > 1; everything else passes there too). Reading C3 as satisfied because the control happened to fail at the one offset we designated canonical would be exactly the post-hoc rescue the pre-registration exists to prevent. Reported as a failure of C3 in substance.

This is the substantive kill, and it is bigger than the GBP question. It means "passes all eight falsifiers" is not rare in this family — it is close to a coin flip across instruments and nuisance parameters. The eight criteria were built to be hard, and on this evidence they are not hard enough to distinguish a real effect from an arbitrary instrument.

The carrying cell is not stable even within one instrument

C2 compared carrying cells across instruments. The offset sweep shows they are not stable within one:

0s 3,600s 7,200s 10,800s
GBPUSD H1P1 H1P1 H4P2 H4P2
GBPCAD H4P1 H4P1 H1P1 H4P1

Moving an arbitrary 4-hour boundary by two hours moves which sub-strategy earns the money. That is not a mechanism; it is reshuffling. It also retroactively weakens every "cell X carries instrument Y" statement made in this project, including the ones I made about EURUSD and GBPUSD earlier today.

Related: EURCAD fails overall (median month −31.3, 45.0% of months profitable, halves below 1) while its H4P1 cell reads PF 1.972 on n=116 — a strong cell inside a failing instrument. Cell-level numbers should not be quoted as evidence on their own.

All six instruments, canonical config

R:R 1.0, defaults, 5m, H4 offset 3,600s, each at its own measured median spread.

symbol role spread n PF exPF median mo % prof ex-top-3 best mo carry verdict
GBPUSD reference 11 582 1.848 1.801 +127.9 63.3% +5,154 2025-04, 16.2% H1P1 PASS
EURUSD reference 5 596 1.100 1.076 −22.0 46.7% −847 H1P1 FAIL
GBPCAD treatment 37 611 1.614 1.580 +79.5 55.0% +5,500 2023-07, 16.1% H4P1 PASS
GBPJPY treatment (underpowered) 26 134 0.894 0.822 +77.8 53.8% −3,670 2024-09 H4P2 UNTESTABLE
EURCAD control 24 553 1.294 1.254 −31.3 45.0% +312 2022-02, 40.7% H4P1 FAIL
AUDJPY control 10 562 1.264 1.239 +41.0 53.4% +763 2026-03, 33.4% H4P1 FAIL (passes at 3 of 4 offsets)

GBPJPY landed exactly where it was declared in advance: n=134 against the 200 minimum, so UNTESTABLE and counting neither way. For information, it was negative anyway (PF 0.894, exPF 0.822, ex-top-3 −3,670).

Cost model note

This was the first run in the project at realistic spreads — each instrument charged its own median from the Tickstory spread column instead of the flat 2 ticks used everywhere before. Charging GBPUSD its real 11 ticks instead of 2 cost it 0.097 of profit factor (1.945 → 1.848) and changed no verdict. Costs were never the explanation for anything here, in either direction. Every earlier number in this project is mildly optimistic by a similar margin; none of the conclusions turn on it.

What is now closed, and what is not

Closed: the GBP-leg hypothesis. GBPUSD's pass does not follow the GBP leg — its partner instrument passes on a different cell, and a control with neither leg passes at three of four offsets. GBPUSD is a favourable member of a family in which passing is common, and it is not re-examined without new data.

Also closed, and this is the wider casualty: the eight-criterion falsifier set as a sufficient standard. It was built from what killed ZMS on 2026-08-26 and it does its job on obviously bad runs, but three of six instruments here pass it at one offset or another. It is a floor, not a filter.

Not closed: EURGBP is still the control that would settle the USD question, and is still a broken 14 KB stub. It remains the single highest-value data task.

Standing rule added: any future ZMS claim must be reported across the full H4 offset sweep, not at one offset. A single-offset result is now known to be capable of flipping both the verdict and the carrying cell.