PatchOracle

Research verdicts

Every terminal canonical finding this project has published, in one list. Negative results and inconclusive measurements are included, not only the ones that came out flattering. A finding that has been superseded by a later, corrected result is not deleted. It stays reachable at its own permanent address, linked from the result that replaced it, but it is not listed by default here.

Verdicts are published exactly as the research index records them, machine vocabulary and all. Where a verdict uses a word that means something specific here, open the plain-language note under it.

Canonical verdicts
33
Superseded / archival
0
Source artifact
research-index.json
Artifact version
1

Research verdicts, as a bazaar menu

Lore

Balance-aware retrospective portfolio diagnostic

Non-result

Cutoff
Final 90-day holdout unread.
Baseline
Frozen portfolio protocol conditional on an eligible R2 survivor.

Typed non-result: no eligible candidate stream and no three-balance retrospective result.

No balance result or live-profitability claim can be computed without an eligible stream.

COMMON

Read the full write-up

Every canonical verdict, in full

MIXED

Bid-ask bounce, or real mean reversion?

Whether the same-day return-reversal signal is mechanical bid-ask bounce or a real mean-reversion effect.

1d — GENUINE — the signal survives measured on a single side (mid -0.29, bid -0.39, ask -0.27; touch retention 113%), so it is not bid-ask bounce. 2h — INCONCLUSIVE — the signal partly survives measured on a single side (mid -0.36, bid -0.30, ask -0.20; touch retention 71%). Some is bounce and some may be real; this test cannot apportion it. Return-based models stay blocked [at 2h].

In plain language (2 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

touch retention
How much of a signal survives when it is measured on one side of the book alone instead of the midpoint. A signal that mostly vanishes was mostly bounce.
bid-ask bounce
A fake price wiggle caused by trades alternating between the buy price and the sell price. The item's value has not moved; only which side traded last has.
NEGATIVE RESULT

CoflNet export gap-recovery feasibility

Determine whether narrow export re-requests can deterministically repair historical full-book gaps.

Narrow re-requests are not deterministic historical-gap recovery; prospective recording is the valid execution-grade route.

OPERATIONAL STATUS

Export coverage

Report present/absent/failed chunk accounting for the /export acquisition.

Operational acquisition-status snapshot (present/absent/failed chunk accounting), not a fixed research verdict. Figures move as the /export acquisition proceeds; treat as current-as-of-generation only.

SUPPORTING MEASUREMENT

Fixed-cohort patch-event relative-change aggregates

Compute fixed-cohort, event-indexed median relative-price-change aggregates for admitted patch-notice events.

287 cohort-event rows cleared the minimum-sample and promotion gate out of 182 admitted events x their matched cohorts (9 cohorts, partition tradeable-v1-b3bfe7781e). Each row is one event's cohort-wide MEDIAN relative price change; no per-item value, extreme or cumulative statistic is ever computed or published. A supporting dataset for the optional patch-events route, not a hypothesis-test verdict.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

cohort
A group of items fixed in advance, so the grouping cannot be adjusted after the result is seen.
NEGATIVE RESULT

Forecasting the spread, honestly

Whether a fitted model forecasts the spread better than a naive persistence null.

Pooled R2 beats the harder null (+0.0729 vs -0.0666) but per-item (equal-weighted and notional-weighted) R2 LOSES to it (-0.4959 vs -0.3538; -1.6372 vs -0.4366). Reporting only the pooled figure 'would have been selection dressed as a result.'

In plain language (3 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

pooled
One number computed over every item's rows thrown together. A pooled figure can look strong while the per-item figures are weak, which is why both are always reported here.
null
The dull baseline a model has to beat, such as assuming tomorrow looks like today. It is chosen and published before any model runs, so it cannot be picked afterwards to make a result look better.
notional-weighted
Two ways of averaging across items. Equal-weighted counts every item the same. Notional-weighted counts the items with more money moving through them for more.
MIXED

Gate A: does any resting-order cycle clear its costs?

Whether any resting-order cycle clears its costs.

CONCLUSIVE NEGATIVE — under a touch-only fill rule, which is a proven upper bound on what could have filled, the median resting cycle does not clear its costs (median net capture per attempted cycle -0.530%, 40.7% of items above zero). Real fills are a subset of these, so the true result can only be worse. No depth data can rescue this configuration.

In plain language (2 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

touch_only
A deliberately generous rule for deciding an order filled: it counts a fill whenever the market price reached the order's price at all. Real fills are a subset of these, so a result measured this way is an upper bound on what could have happened, not an estimate of what did.
attempted cycle
One full attempt at the strategy: place a buy order, wait, and sell if it fills. Attempts that never filled are counted too, so an average per attempted cycle cannot be flattered by dropping the ones that did not work.
RESULT

Gate A's long-patience capture: drift, or earned spread?

Decompose Gate A's long-patience positive capture into drift share vs. genuine per-item residual, graded against ADR-0004's pre-registered criteria.

GENUINELY EARNED — 1d: drift share 0.084 is below DRIFT_SHARE_THRESHOLD = 0.5 and the median per-item residual +7.170% is above zero ... ; 2h: drift share -0.062 is below DRIFT_SHARE_THRESHOLD = 0.5 and the median per-item residual +10.819% is above zero ... . Graded against ADR-0004's table, at 7-day patience under touch_only with the artifact guard on, 0bps from the weighted quote reference.

In plain language (6 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

drift share
The fraction of a result explained by the whole market moving rather than by the strategy doing anything.
DRIFT_SHARE_THRESHOLD
The cutoff, fixed in advance at 0.5, for how much of a result may come from the whole market drifting rather than from the strategy earning it. Above the cutoff the result is not credited.
residual
What is left of a return once the known costs and effects are subtracted. It describes how the quote was priced and is not earned edge, which is why it no longer decides whether a strategy passes.
touch_only
A deliberately generous rule for deciding an order filled: it counts a fill whenever the market price reached the order's price at all. Real fills are a subset of these, so a result measured this way is an upper bound on what could have happened, not an estimate of what did.
bps
Basis points. One basis point is a hundredth of a percent, so 100 bps is 1 percent.
weighted quote
The headline price Hypixel reports, averaged across the orders on one side of the book rather than taken from the best one. It sits wider than the touch, about 1.8 percentage points at the median, so a spread measured against it overstates the real cost of crossing.
OPERATIONAL STATUS

Gate B at 3- and 7-day horizons

Determine whether recorded depth can resolve 3- and 7-day calibration attempts.

Every 3- and 7-day cell is not yet measurable; the result is immature coverage, not a failed or extrapolated calibration.

RESULTSEALING EXEMPTION

Gate B on self-built bars: a sample that grows, and whether it agrees

Whether Gate B's calibration verdict holds on a self-built, growing bar sample instead of the static panel.

[sample: self-built bars (recorder snapshots, 2h grid)] NO OVERSTATEMENT MEASURED — the median item's bar-derived fill count differs from recorded depth by +0.00 percentage points of attempted cycles, and the two agree at the median; 48% of items overstate. ... This is a FIRST calibration over 51.9 hours of recorded depth, 413 items, 7,787 resolvable attempted cycles, not a settled constant.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

attempted cycle
One full attempt at the strategy: place a buy order, wait, and sell if it fills. Attempts that never filled are counted too, so an average per attempted cycle cannot be flattered by dropping the ones that did not work.
RESULTSEALING EXEMPTION

Gate B: how much does bar-derived fill detection overstate real fills?

How much bar-derived fill detection overstates real fills, calibrated against recorded order-book depth.

NO OVERSTATEMENT MEASURED — the median item's bar-derived fill count differs from recorded depth by -14.29 percentage points of attempted cycles, and the bar panel MISSED fills the recorded book shows, which is a defect in the opposite direction and is not a reason for confidence; 28% of items overstate. Under the touch-only rule ... bar data is fit for purpose for fill detection at this configuration. This is a FIRST calibration over 51.9 hours of recorded depth, 400 items, 2,770 resolvable attempted cycles ... not a settled constant.

In plain language (2 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

attempted cycle
One full attempt at the strategy: place a buy order, wait, and sell if it fills. Attempts that never filled are counted too, so an average per attempted cycle cannot be flattered by dropping the ones that did not work.
touch_only
A deliberately generous rule for deciding an order filled: it counts a fill whenever the market price reached the order's price at all. Real fills are a subset of these, so a result measured this way is an upper bound on what could have happened, not an estimate of what did.
OPERATIONAL STATUS

Gate C: observed vs. elapsed cadence windows (not an expected-value verdict)

Report v1's observed vs. elapsed forward-run cadence coverage -- explicitly not an expected-value verdict.

v1 spoke in 3 of 3 elapsed windows (100.0%; 2 with a complete cohort, 1 torn, 0 deliberately empty, 0 with no record) at an assumed 24h cadence anchored 2:00 UTC, over 2026-08-07 00:45 -> 2026-08-09 02:00 UTC. No outages: every one of the 3 elapsed windows has a record. (measured live at 2026-08-09T22:59:17Z)

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

cohort
A group of items fixed in advance, so the grouping cannot be adjusted after the result is seen.
RESULT

Is the 2h drift residual's excess selection on spread width, or a defect? (#84)

Whether the 2h drift residual's excess capture is explained by selection on wider fill spreads, or is a defect.

Completed cycles select NARROWER, not wider, BUY-fill spreads (median fill-vs-item spread delta -0.403pp at 1d, -1.126pp at 2h); no evidence of an inverted book side, mismatched item, or entry/exit join defect.

NEGATIVE RESULT

Item-selection interior real-candidate search

Determine whether a preregistered item-selection grid produces an eligible real-candidate strategy with an interior optimum.

Terminal outcome: `grid_edge_unidentified`. The winning vol_7d / 30-day / top-40 cell had +0.9155% EV per attempted cycle over 6,174 attempts but t=1.552 < 3.33 and sat at both continuous grid boundaries, so it is not a real candidate and does not trigger another search.

In plain language (3 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

grid_edge_unidentified
The best setting the search found sat on the outer edge of the range that was searched. That means the edge of the search chose it, not the data. Widen the range and the answer would move, so this is not an answer yet.
vol_7d
A ranking of items by how much they traded over the previous seven days.
attempted cycle
One full attempt at the strategy: place a buy order, wait, and sell if it fills. Attempts that never filled are counted too, so an average per attempted cycle cannot be flattered by dropping the ones that did not work.
SUPPORTING MEASUREMENT

Naive forecast baselines

Establish naive/persistence forecast baselines before any model is built.

Published before any model exists, per the project's standing rule that a null chosen after seeing model results cannot be distinguished from a rationalization. Supplies the persistence/naive nulls spread-forecast.md is graded against.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

null
The dull baseline a model has to beat, such as assuming tomorrow looks like today. It is chosen and published before any model runs, so it cannot be picked afterwards to make a result look better.
OPERATIONAL STATUS

Recorded depth: coverage and provenance

Report recorder-era coverage and third-copy backup currency for recorded order-book depth.

Recorder era coverage and third-copy backup currency, regenerated per run. Not a fixed research verdict.

NON-RESULT

Strategy: the tuned sweep

Find the interior expected-value optimum over the declared execution-parameter grid.

THE OPTIMUM IS NOT IDENTIFIED. Expected value rises monotonically to the edge of the grid on depth_bps, sell_depth_bps, ttl_days, and the best configuration sits at that edge. The grid bound chose this configuration, not the data. (grid_edge_unidentified: true; canonical digest 1eeb1ad85faa5d700330d0180334d38a8ea859596f3d140ffe2213e12e525887)

In plain language (5 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

depth_bps
How far below the going rate a buy order is placed, in basis points. Deeper means a better price if it fills, and a smaller chance it fills at all.
sell_depth_bps
The same setting on the way out: how far above the going rate the sell order is placed once the buy has filled.
ttl_days
How many days an order is left sitting before it is given up on. This is the patience setting.
grid_edge_unidentified
The best setting the search found sat on the outer edge of the range that was searched. That means the edge of the search chose it, not the data. Widen the range and the answer would move, so this is not an answer yet.
canonical digest
A fingerprint of the exact numbers a verdict was written against. If the underlying result is ever regenerated and changes, the fingerprint stops matching and the claim no longer applies.
SUPPORTING MEASUREMENT

The calendar-window migration: what moved

What changed in the calendar-window panel migration, and whether Gate A's headline result moved.

Daily panel: 37 columns identical, 8 moved, 4 added; 2h panel: 42 identical, 0 moved, 4 added (must be zero, and is). Gate A's headline re-runs to -0.530% with 40.7% of items above zero -- unchanged. Movement is tilted outside the tradeable tier on all 8 moved columns but NOT confined there.

SUPPORTING MEASUREMENT

Tier-A lifecycle watchability and frozen three-day bar baseline

Measure fixed-horizon lifecycle watchability for the frozen recorder cohort and reproduce one frozen legacy bar baseline.

3d is first measurable for aggregate watchability (419 complete-lifecycle-eligible placements); the frozen legacy comparator has no economic drift, only receipt-bound touch-only capacity-counter instrumentation drift.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

touch_only
A deliberately generous rule for deciding an order filled: it counts a fill whenever the market price reached the order's price at all. Real fills are a subset of these, so a result measured this way is an upper bound on what could have happened, not an estimate of what did.
OPERATIONAL STATUS

Toxicity-MM corpus admission

Admit the private full-book development corpus and account for temporal gaps.

COVERAGE_GAPPED with 1,253 temporal gaps; downstream replay must fail closed across them.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

COVERAGE_GAPPED
A status meaning the recording has holes in it. Anything built on top has to stop at each hole rather than guess across it.
INCONCLUSIVE

Toxicity-MM executable-markout veto

Test whether an executable-markout veto improves the unchanged frozen replay baseline.

Economically inconclusive despite formal selection; it cannot open the holdout or support profitability.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

holdout
A slice of the data set aside and deliberately never looked at, saved for one final test. Looking at it early spends it, and it cannot be put back.
NEGATIVE RESULT

Toxicity-MM fill-hazard policy slice

Test whether a causal fill-risk gate improves the frozen pessimistic strategy economically.

Negative stop: no fill-hazard gate passed the strict economic-improvement rule.

In plain language (2 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

negative stop
The measurement came back negative and the work stopped there, by a rule agreed before the measurement ran. It was not retried under different settings until it came out differently.
fill-hazard
An attempt to predict how likely a resting order is to fill in the next stretch of time, so that orders unlikely to fill could be skipped.
SUPPORTING MEASUREMENT

Toxicity-MM gap-aware replay

Replay frozen strategy policies without bridging corpus gaps or hiding queue ambiguity.

Gap-aware aggregate replay exists; a temporal gap is a no-trade coverage outcome and is never interpolated.

OPERATIONAL STATUS

Toxicity-MM Gate 0

Freeze the cohort and test full-book cache fidelity before replay.

REQUIRES_ACQUISITION: zero of 60 items were usable in the original corpus account.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

REQUIRES_ACQUISITION
A status meaning the data needed does not exist yet. The question cannot be answered until it is recorded.
NEGATIVE

Toxicity-MM latency20 strategy negative stop

Score the latency-corrected pessimistic market-making strategy reprice5-pessimistic-latency20-v1 from bounded aggregate slices.

The slice intersection did not pass; terminal negative stop with the holdout unaccessed.

In plain language (2 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

negative stop
The measurement came back negative and the work stopped there, by a rule agreed before the measurement ran. It was not retried under different settings until it came out differently.
holdout
A slice of the data set aside and deliberately never looked at, saved for one final test. Looking at it early spends it, and it cannot be put back.
NEGATIVE RESULT

Toxicity-MM R1 terminal account

Bind R1 to the preregistered fill-hazard result.

R1 is a terminal non-candidate because the fill-hazard measurement is a negative stop.

In plain language (3 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

non-candidate
Ruled out as something to build on. Later stages that depended on it do not get to run on a weaker substitute.
fill-hazard
An attempt to predict how likely a resting order is to fill in the next stretch of time, so that orders unlikely to fill could be skipped.
negative stop
The measurement came back negative and the work stopped there, by a rule agreed before the measurement ran. It was not retried under different settings until it came out differently.
NON-RESULT

Toxicity-MM R2 terminal non-result

Run the one-shot R2 holdout only if the exact frozen R1 account survives.

R2 is a one-shot non-result because the exact frozen R1 account is a non-candidate.

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

non-candidate
Ruled out as something to build on. Later stages that depended on it do not get to run on a weaker substitute.
SUPPORTING MEASUREMENT

Tradeable cohorts and the patch-event admission rule

Freeze the disjoint cohort partition of the tradeable tier and the fixed per-release list of admitted patch-notice events.

Freezes the disjoint cohort partition of the tradeable tier and the fixed per-release list of official-notice events admitted into it. Computes no price aggregate itself (by design -- see analysis/cohorts.py docstring).

In plain language (1 term)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

cohort
A group of items fixed in advance, so the grouping cannot be adjusted after the result is seen.
NON-RESULT

Twenty-second resolution gate

Create a future twenty-second protocol only if R2 has a survivor.

No twenty-second protocol exists because R2 produced no survivor.

NON-RESULT

v2: the digest-bound NotFrozen non-result

Record v2's frozen/not-frozen status as an explicit non-result.

v2 is WITHHELD, not frozen and not abandoned. Reason (re-recorded 2026-08-09): weighted-quote geometry is attributed, but the per-attempt frontier still ends at the grid wall under a fill model whose trade-off did not identify an interior optimum. Bound to docs/research/strategy.json's canonical digest 1eeb1ad85faa5d700330d0180334d38a8ea859596f3d140ffe2213e12e525887.

In plain language (4 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

withheld
The work was done and the result was not published. That is a different thing from abandoned and a different thing from never attempted.
weighted quote
The headline price Hypixel reports, averaged across the orders on one side of the book rather than taken from the best one. It sits wider than the touch, about 1.8 percentage points at the median, so a spread measured against it overstates the real cost of crossing.
interior optimum
A best setting that sits inside the range searched rather than on its edge. Finding one is what shows the data picked the setting. Not finding one is why the search above is not treated as an answer.
canonical digest
A fingerprint of the exact numbers a verdict was written against. If the underlying result is ever regenerated and changes, the fingerprint stops matching and the claim no longer applies.
SUPPORTING MEASUREMENT

What the price-bound guard changes about the measured fill rate

How much of the touch-only fill rate the price-bound guard artifact itself accounts for, versus a real fill.

At the touch the correction is ~1%. At 25% placement depth on the BUY side, 17.83% of unguarded fills came from the price-bound artifact alone under the primary VOLUME_CONFIRMED rule, and VOLUME_CONFIRMED carries a HIGHER false-fill share than TOUCH_ONLY at every depth >= 1% ('confirmation concentrates the contamination rather than diluting it').

In plain language (4 terms)

This site's own explanation of the words above. The verdict itself is published exactly as the research index records it and is not rewritten here.

the touch
The best price sitting on the order book right now. It is not the same as the headline price the game shows, which is an average across the orders on that side.
price-bound artifact
A fill the simulation counted only because the recorded price range touched the order's price, with nothing else showing a trade actually happened. It is a measurement error, not a fill.
VOLUME_CONFIRMED
A stricter fill rule that also requires traded volume to back up the price move. Stricter did not turn out to mean more accurate: it was measured carrying a higher share of false fills than the generous rule at every depth tested.
touch_only
A deliberately generous rule for deciding an order filled: it counts a fill whenever the market price reached the order's price at all. Real fills are a subset of these, so a result measured this way is an upper bound on what could have happened, not an estimate of what did.
RESULT

Why a 100/min ceiling is reached at one request every three minutes

Determine why the 100/min export rate ceiling is effectively reached at one request every three minutes.

Supported hypothesis: FAN-OUT. 429 risk is 3.9x higher after a large chunk (18% of 73) than after a small one (5% of 22); inconsistent with a shared-address or competing-process explanation.

OPERATIONAL STATUS

Widening historical depth to the tradeable tier

Track acquisition progress of /export historical depth widened to the tradeable tier.

Operational scope/progress snapshot for the /export widening acquisition (413 items, 2,891 chunks, ordered by notional traded per week). Not a fixed research verdict; figures move as acquisition proceeds.

See methods & provenance for the seal, the chance ledger, execution bounds, and how the weighted quote differs from the touch.