PatchOracle

Methods & provenance

What this project holds back from itself, how it counts its own attempts, what a number on this site actually measures, and what it may publish. Every claim below traces to a specific decision record in patchoracle-seed.

The seal

Every measurement on this site's research route takes its data through a sealed boundary: everything from 2026-02-01T00:00:00Z onward is withheld from development. The boundary is a one-way door — moving it earlier (withholding more) is always allowed; moving it later, after any experiment has run, would need its own decision and would be an admission, not a tuning step. One row past the boundary is treated as a breach, not a tolerance: there is no percentage of leakage this project accepts as fine.

Two named exceptions exist, and only two. `gate_b_calibration` (issue #63) and `gate_b_self_built_bars` (issue #65) each read a narrow slice of post-boundary order-book depth to calibrate how much bar-derived fill detection overstates real fills — every recorded order-book snapshot postdates the boundary, so this measurement has no development-span alternative. Each grant is held by an exact experiment name, not a category: "calibration may read the sealed span" was considered and rejected, because a category is exactly the kind of rule that stretches to cover whatever comes next. A second Gate B sample over a different set of bars needed, and got, its own separate grant rather than inheriting the first.

The chance ledger

This project was graded FAIL once for overfitting: a grid of 24 configurations was searched over roughly 524 rows and reported on 42 trades, with the size of that search never recorded anywhere. The mechanism was not fraud — it was an uncounted denominator, and an uncounted denominator always flatters whichever configuration a search happens to land on.

Every configuration this project has scored since is now counted before it is scored, by a single ledger that mints a receipt for each attempt. The scoring function refuses to run without a genuine receipt traced back to that ledger on disk — not merely an object that claims to be one. A deflation computed against an undercounted denominator is the same failure with extra steps, so the count itself is a measurement this project makes, rather than a number anyone attests to after the fact.

Weighted quote, not touch

Every historical price this project reports is a volume-weighted average across the resting order book, not the best bid or ask. In one measured snapshot, Hypixel's own summary endpoint reported a buy price of 741.76 while the best ask in the same book was 731.80 — about a percentage point apart, and about 1.8 percentage points at the median across the full item set. Every spread computed from the historical panel is therefore a proxy for the real crossing cost, measured wider than the touch, not the touch itself. Only the live feed — the true best bid and best ask, read directly from the order book — is a touch price.

Execution bounds, and what a filled cycle earned

Every fill this project reports comes from simulating orders against recorded book state, never from a real trade — the bazaar API has no order-placement endpoint. Two fill rules bound how optimistic that simulation is allowed to be: `touch_only` asks only whether a bar's price extreme reached the resting order's limit, and is a proven UPPER bound on what could have filled, not a positive result on its own. `volume_confirmed` also checks that enough volume traded to plausibly account for the fill. A "real candidate" claim requires clearing costs under `volume_confirmed` together with a pessimistic resolution of every ambiguous case — a claim resting only on `touch_only` or an optimistic reading is bound-dependent, not a result.

The completed-cycle residual (return minus drift) was this project's first measure of what a filled cycle earned beyond the market's own drift. It turned out to be dominated by the width of the weighted quote itself and by which cycles the fill rule selects into the completed population — the residual is not earned edge going forward, and expected value per attempted cycle (`capture` in this project's own vocabulary) now carries that claim instead, because no-fill attempts remain in its denominator where the residual's did not. This project also runs two versions of its resting-order strategy forward, `v1` (immutable, pre-registered before any backtest existed) and `v2` (informed by a parameter search). As of this build, `v2` is WITHHELD — its sweep found no interior optimum on every tuning axis, so no configuration was frozen — not abandoned and not unattempted; `v1` alone carries this project's primary forward verdict while that holds.

Gate B's calibration window

Gate B measures how much bar-derived fill detection overstates real fills, by resolving identical simulated orders against recorded order-book depth rather than bar extremes alone. Every hour of that recorded depth postdates the sealed boundary, so this measurement runs under the named exemptions above — the only two the seal grants.

This is a FIRST calibration over roughly 52 hours of recorded depth across 400-413 items and about 2,770 resolvable attempted cycles — a narrow window that covers only 2, 6 and 12-hour patience, not the 1, 3 or 7-day horizons this project's longer-patience results would need. It is not a settled constant: it is one measurement over one short recording window, reported so every bar-derived result elsewhere on this site can be read against it, and it will be superseded by a wider calibration rather than trusted indefinitely as this recorder's history grows.

What this project may publish, and from which source

Two upstream CoflNet endpoints back this project's history, under two different licences that moved at different times. `/export` (a paid, higher-resolution feed) narrowed its publication rule under issue #102: whether something built from it may publish depends on what the resulting artifact IS, not on whether export data was ever consulted while building it. A derived statistic — a fill-rate curve, a calibration result, a spread distribution, a verdict — publishes, because it is this project's own measurement and cannot be inverted back into the observations that produced it. A raw series — price, depth or volume at the granularity observed — does not, because shipping it is redistribution of CoflNet's data however it is reshaped. `/history`, the older free endpoint this project's deep archive was built from, is unchanged and more restrictive still: those price series remain fine for local work and are not redistributable at all without separately revisiting CoflNet's terms. Every artifact this site publishes declares which kind it is before it can be built at all, and a route showing a number always says which side of that line the number came from.

UTC is canonical

Every timestamp on this site is UTC. The sealed boundary itself is defined in real UTC time rather than wall-clock time precisely so its meaning cannot shift depending on which build produced it, and the daily panel's rows are calendar-day aggregates anchored at 02:00 UTC — a "day" here always means that same fixed UTC window, not a local midnight that would shift by timezone or by daylight saving.

This build

This page and every research verdict on this site are rendered from one promoted publication manifest, authenticated by hash before anything here trusts its bytes.

Manifest
d1e115dba980e81869a82afb759728b5bb2ce40fc58a6359c39413384a48f59a
Producer commit
215f756130e18c8dae1226fbcdc6d026792b8ee2
Data cutoff
2026-10-08T09:59:42.750000+00:00

See research verdicts for the terminal findings these methods govern, and about & method for how the price panel itself is collected and served.