16  Validation & Anti-Overfitting

Author: Todd B. Adams Reinforces: Proposal §6 — Model Proofing · Reading order: chapter 08 of this part; follows the gap analysis & research roadmap, precedes reproducibility & experiment infrastructure.

What this chapter establishes. Graph-model research in finance has a reputational problem: it overfits, and it is rarely held to the multiple-testing and out-of-sample standards the asset-pricing literature demands. We argue that a production-grade, literature-calibrated validation gauntlet — which any new producer, a network one included, must pass — is the platform’s single biggest contribution to the credibility of the proposal’s results.


16.1 The gauntlet already in production

The platform’s “mathematics of trust” is not aspirational; it gates real factor promotions today, and each check is adopted from the literature rather than invented. The full narrative is in the platform whitepaper.

Check Question it answers Grounded in
Information-coefficient hurdle Does the signal actually predict? Grinold & Kahn, Fundamental Law
Multiple-testing / false-discovery control Did we get lucky across many tries? Harvey, Liu & Zhu (2016); Benjamini–Yekutieli (2001)
Permutation / Monte-Carlo tests Could randomness produce this? White (2000), Reality Check
Survivorship haircut Are dead companies flattering it? Shumway (1997); Brown et al. (1992)
Net-of-cost and capacity Does the edge survive trading? Corwin–Schultz (2012); Novy-Marx & Velikov (2016)
Deflated Sharpe ratio, probability-of-backtest-overfitting, lockbox Are we overfitting by tuning? Bailey & López de Prado (2014); López de Prado (2018)
Purged and embargoed cross-validation Is the out-of-sample actually clean? López de Prado (2018)
Honest family-wise trial count Deflated against how many trials, really? the platform’s cross-axis trial ledger

16.2 Why this is decisive for the proposal

  1. The synthetic simulator graduates into the gauntlet. The proposal’s simulator proves the physics on 100 synthetic stocks by injecting a known shock and confirming that a brittleness index moves. That is a necessary first step, but a synthetic positive is not evidence of a real edge. On this platform the network signal’s empirical results would run through permutation tests, the probability-of-backtest-overfitting estimate, the deflated Sharpe ratio, and a lockbox, so a phase-transition-detector claim is tested for luck, tuning, and leakage before it is believed.
  2. Multiple-testing rigour a graph-model paper rarely applies. A spatial-temporal graph model has enormous architectural degrees of freedom — layers, attention heads, windows, couplings. Honest trial-count deflation and the overfitting estimate are exactly the instruments for the “best-of-many-architectures” search that is the dominant failure mode of deep-learning finance.
  3. Report-then-enforce keeps unproven signals harmless. A network early-warning signal would ship report-only first (chapter 06); it could not move a book until it had survived attributed evidence — the same discipline that keeps every experimental factor out of real capital.
  4. The platform is willing to publish a negative. The connectedness spike (findings) is a documented, honest falsification. That disposition — disclosure over avoidance — is what makes a positive result from the same platform worth trusting.

16.3 The standard the network claim must clear

For a Supra-Laplacian early-warning signal to be reported as real, it would need to show — on the platform’s empirical universe, not a synthetic panel — a signal that survives permutation, deflates under honest trial counts, holds out-of-sample through a lockbox, and does so net of the composition-drift and coincidence artefacts the spike already surfaced. That is a high bar by design, and clearing it is precisely what would make the proposal’s contribution defensible.


Cross-links: platform whitepaper — mathematics of trust · reference library · proposal §6