15 Validation & Anti-Overfitting
Author: Todd B. Adams Reinforces: Proposal §6 — Model Proofing · Reading order: doc 08 of the research-platform package
The claim of this doc. GNN-in-finance research has a reputational problem: it overfits, and it is rarely held to the multiple-testing and out-of-sample standards the asset-pricing literature demands. SBFoundation brings a production-grade, literature-calibrated validation gauntlet that any new producer — including a network one — must pass. This is the platform’s single biggest contribution to the credibility of the proposal’s results.
15.1 The gauntlet already in production — BUILT
The platform’s “mathematics of trust” is not aspirational; it gates real promotions today. Each check is adopted from the literature, not invented. Full narrative: whitepaper §5.
| Check | Question it answers | Grounded in |
|---|---|---|
| Information Coefficient hurdle | Does the signal actually predict? | Grinold & Kahn, Fundamental Law |
| Multiple-testing / false-discovery | Did I get lucky across many tries? | Harvey, Liu & Zhu (2016); Benjamini–Yekutieli (2001) |
| Permutation / Monte-Carlo (MCPT) | Could randomness produce this? | White (2000) Reality Check; the “Masters” suite |
| Survivorship haircut | Are dead companies flattering it? | Shumway (1997); Brown et al. (1992) |
| Net-of-cost + capacity | Does the edge survive trading? | Corwin–Schultz (2012); Novy-Marx & Velikov (2016) |
| Deflated Sharpe + PBO + lockbox | Am I overfitting by tuning? | Bailey & López de Prado (2014); López de Prado (2018) |
| Purged + embargoed CV | Is my out-of-sample actually clean? | López de Prado (2018) |
| Honest family-wise trial count | Deflated against how many trials, really? | platform’s cross-axis trial ledger (F-321) |
15.2 Why this is decisive for the proposal
- The synthetic simulator graduates into the gauntlet. The proposal’s §6 simulator proofs the physics on 100 synthetic stocks by injecting a known shock and confirming the brittleness index moves. That is a necessary first step — but a synthetic positive is not evidence of a real edge. On this platform, the network signal’s empirical results would be run through MCPT, PBO, deflated-Sharpe and a lockbox, so a “phase-transition detector” claim is tested for luck, tuning, and leakage before it is believed.
- Multiple-testing rigor a GNN paper rarely applies. A spatial-temporal GNN has enormous architectural degrees of freedom (layers, heads, windows, couplings). The platform’s honest trial-count deflation (F-321) and PBO are exactly the instruments for “best-of-many-architectures” — the dominant failure mode of deep-learning finance.
- Report-then-enforce keeps unproven signals harmless. A network early-warning signal ships report-only first (doc 06); it cannot move a book until it has survived attributed evidence — the same discipline that keeps every experimental factor out of real capital.
- The platform is willing to publish a negative. The DY-connectedness spike (findings) is a documented, honest falsification. That culture — disclosure over avoidance — is what makes a positive result from the same platform worth trusting.
15.3 The standard the network claim must clear
For a Supra-Laplacian early-warning signal to be reported as real, it would need to show — on the platform’s empirical universe, not a synthetic panel — a signal that survives permutation, deflates under honest trial counts, holds out-of-sample through a lockbox, and does so net of the composition-drift and coincidence artifacts the spike already surfaced. That is a high bar by design, and clearing it is precisely what would make the proposal’s contribution defensible.
Cross-links: whitepaper §5 — mathematics of trust · reference papers · proposal §6