PatchOracle

← Research verdicts

Forecasting the spread, honestly

negative-resultforecast

Question

Whether a fitted model forecasts the spread better than a naive persistence null.

Verdict

Pooled R2 beats the harder null (+0.0729 vs -0.0666) but per-item (equal-weighted and notional-weighted) R2 LOSES to it (-0.4959 vs -0.3538; -1.6372 vs -0.4366). Reporting only the pooled figure 'would have been selection dressed as a result.'

In plain language (3 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

pooled
One number computed over every item's rows thrown together. A pooled figure can look strong while the per-item figures are weak, which is why both are always reported here.
null
The dull baseline a model has to beat, such as assuming tomorrow looks like today. It is chosen and published before any model runs, so it cannot be picked afterwards to make a result look better.
notional-weighted
Two ways of averaging across items. Equal-weighted counts every item the same. Notional-weighted counts the items with more money moving through them for more.
Population
381 items, 434,492 rows.
Resolution
Daily resolution (2020-08-01 -> 2026-01-31).
Cutoff
Development span only (2020-08-01 -> 2026-01-31).
Baseline
The harder naive/persistence null from forecast-baselines.md.

Limitations

Pooled R2 beats the null but per-item (equal-weighted and notional-weighted) R2 LOSES to it -- reported on the population that matters operationally (per-item), not smoothed into the pooled figure.

Notes

REQUIRED NEGATIVE RESULT for issue #130's index acceptance criterion -- reported as negative on the population that matters operationally (per-item), not smoothed into a qualified positive via the pooled figure. docs/writeup/technical-narrative.md §5 cites this document under 'issue #48'; #48 does not itself appear inside spread-forecast.md's own text (only #42, #46, #57, #64, #66 do) -- flagged here rather than silently resolved, per issue #130's accuracy rule. Also contains the two leakage demonstrations retold in docs/writeup/failure-stories.md story 5.

Source

Citation
po-research:spread-forecast:2988c9ed4e68311d
Write-up
docs/research/spread-forecast.md
Data
docs/research/spread-forecast.json
Artifact
research-index.json v1 — sha256:cb3d976779a5cbcb4188039a8d26d5e558a2d844030a638c2f16fe5a8eb0a2fe
Build data cutoff
2026-08-12
Issues
#42, #46, #57, #64, #66