This project's own rule library, scored on the same terms as the books it implements.
A library of rules, each one implemented from a specific page of a specific book, each emitting falsifiable claims: a level, a direction or a date, with a tolerance, priced before the outcome is known. Claims are graded forward, never re-scored after the fact.
A 63,990-claim walk-forward. The comparison that matters is
pass_rate against
null_all (like-for-like) — a like-for-like pair — and the
engine comes in at -2.9pp.
accuracy - null_rate -- and the report that owns it carries the field accuracy_is_not_comparable_to_null: true and a reading note saying in terms: do not read that difference as skill. The like-for-like pair is pass_rate vs null_all, about -2.91pp. The house was publishing the one quantity its own source forbids, twice, as its headline. The sign and rough size do not change; the label does. Awaiting reports/walkforward/, which is not in the restored state.
A second defect, found on 2026-09-08 and worth more than the headline. The project prices pooled results by an effective sample size — how many genuinely independent series the universe is worth, given that SPY and QQQ are not two facts about the world.
That statistic was computed as measured + 1.0 × (declared −
measured): it charged one effective series for every symbol file that
was absent. With 6 of 21 files present, the headline 17.8 was
2.8 measured and 15.0 charged by fiat. With all 21 present it is
7.455, entirely measured.
The part that makes it a finding rather than a bug: the constant had two consumers with opposite signs. One made pooled results harder to clear; the other, added three days later and never declared, made them easier. So a placeholder introduced to be conservative was, in net, anti-conservative — pooled rules were priced on up to 2.39× more observations than they had.
The structural fix that follows: a constant declared conservative should enumerate its consumers and assert the count, so that adding one fails a test rather than silently reversing what the constant means. That is proposed here, not yet implemented, and saying so is the difference between a plan and a claim.
The forward record is short — a single market episode, graded once — and instruments observed over one overlapping fortnight share the same market-wide move, so the independent observations number far fewer than the rows. A count of rows is not a sample size. Nothing here is evidence in either direction, and it is published so that the next measurement has something to disagree with.