Written and committed BEFORE running. Criteria in this file are not edited after the commit that adds it. Any change requires a new file and a new run.
2026-08-27.
On 5 years of Tickstory 5m data, ZMS defaults cleared all eight falsifiers on GBPUSD (PF 1.945, ex-outlier 1.896, median month +131.9, ex-top-3 +5,745) and failed on EURUSD (PF 1.129, ex-top-3 −595), at every H4 grid offset tested. The two pairs are ~0.9 correlated and share the USD leg. The only thing GBPUSD has that EURUSD does not is the GBP leg.
H1 (treatment): the pass follows GBP. Then a GBP cross with no USD leg should pass. H0 (control): the pass is the favourable tail of a family of runs. Then GBP crosses should behave like anything else, and non-GBP non-USD crosses should sometimes pass too.
| role | symbol | span | status |
|---|---|---|---|
| treatment, primary | GBPCAD | 2021-03-25 → 2026-03-24 | full 5y |
| treatment, secondary | GBPJPY | 2024-03-21 → 2025-05-09 | ~14 months, file truncated mid-line — declared UNDERPOWERED in advance |
| control | EURCAD | 2021-03-26 → 2026-03-25 | full 5y |
| control | AUDJPY | 2021-03-26 → 2026-03-25 | full 5y |
| reference | GBPUSD, EURUSD | 5y | re-run under the new cost model |
GBPJPY carries ~14 months against the others' 60. It is reported for information and cannot satisfy any criterion in either direction. The treatment arm is therefore a single instrument, which is a real weakness of this test and is stated here rather than discovered afterwards. EURGBP — still a broken 14 KB stub — remains the control that would settle this best and is still unavailable.
Every previous run in this project used slippage_ticks = 2. The Tickstory files carry
the broker's own spread column, and the measured medians are:
| EURUSD | GBPUSD | GBPCAD | GBPJPY | EURCAD | AUDJPY | |
|---|---|---|---|---|---|---|
| median spread (ticks) | 5 | 11 | 37 | 26 | 24 | 10 |
| p90 | 8 | 17 | 53 | 37 | 35 | 16 |
So GBPUSD's pass was computed at roughly a fifth of its real spread, and crosses are 2–4× wider again. Charging the crosses realistically while leaving GBPUSD optimistic would rig the comparison, so:
slippage_ticks is set to each instrument's own measured median spread, applied
adversely on market and stop fills only (limit fills unaffected, per
execution.py:_slip). This is deliberately conservative — it charges the full spread
on entry and again on a stop exit — and, more importantly, it is constructed
identically for all six instruments, which is what a treatment-vs-control comparison
requires.
ZMS defaults, unchanged: H1+H4 bias on, T1 split 50%, R:R gate 1.0, max 3 entries,
no session flat, divergence off, Rolling-24 anchor, percentile thresholds. 5m chart,
H4 grid offset 3,600s. pyakao.zms.run_zms, conservative fill model. No parameter is
tuned in this test.
An instrument is UNTESTABLE — and can satisfy no criterion in either direction — if
it produces fewer than 200 closed trades, or if run_zms emits any warning (HTF leg
warm-up, chart percentile warm-up, H4 grid advisory). GBPJPY is expected to land here.
All five must hold for H1 to survive.
Report.verdict(halves=...)["passes"]
is True): PF > 1, ex-outlier PF > 1, best month ≤ 60% of net, median month > 0,
50% of months profitable, net ex-top-3-months > 0, both halves PF > 1.
| outcome | conclusion |
|---|---|
| C5 fails | Test VOID. The GBPUSD pass was a cost artifact. Closure stands, strengthened. |
| C1–C5 all hold | H1 survives its first real test. ZMS is not reopened and takes no capital; it becomes one named forward-observation candidate (GBP crosses, defaults, real spreads) with a stated review date. |
| C1 fails | The pass does not follow the GBP leg. Closure stands, strengthened. |
| C3 fails | The pass is not GBP-specific. Closure stands. |
| C2 or C4 fails | No consistent mechanism carries it. Closure stands. |
A mixed result is a fail. "GBPCAD passed but a control also passed", "GBPCAD passed at three offsets of four", "GBPJPY looked good" — each of these is reported as a failure of this test, not as encouragement.
Report.verdict(), not argued.Prior probability I would state now, before running: low. Six modifications, nine instruments and two timeframes have failed this project's tests, every apparent positive so far has been instrument-specific, outlier-carried, single-month, or directional beta, and this hypothesis was generated post-hoc from the very run it seeks to explain.