16 Validation & Anti-Overfitting
Author: Todd B. Adams Reinforces: Proposal §6 — Model Proofing · Reading order: chapter 08 of this part; follows the gap analysis & research roadmap, precedes reproducibility & experiment infrastructure.
What this chapter establishes. Graph-model research in finance has a reputational problem: it overfits, and it is rarely held to the multiple-testing and out-of-sample standards the asset-pricing literature demands. We argue that a production-grade, literature-calibrated validation gauntlet — which any new producer, a network one included, must pass — is the platform’s single biggest contribution to the credibility of the proposal’s results.
16.1 The gauntlet already in production
The platform’s “mathematics of trust” is not aspirational; it gates real factor promotions today, and each check is adopted from the literature rather than invented. The full narrative is in the platform whitepaper.
| Check | Question it answers | Grounded in |
|---|---|---|
| Information-coefficient hurdle | Does the signal actually predict? | Grinold & Kahn, Fundamental Law |
| Multiple-testing / false-discovery control | Did we get lucky across many tries? | Harvey, Liu & Zhu (2016); Benjamini–Yekutieli (2001) |
| Permutation / Monte-Carlo tests | Could randomness produce this? | White (2000), Reality Check |
| Survivorship haircut | Are dead companies flattering it? | Shumway (1997); Brown et al. (1992) |
| Net-of-cost and capacity | Does the edge survive trading? | Corwin–Schultz (2012); Novy-Marx & Velikov (2016) |
| Deflated Sharpe ratio, probability-of-backtest-overfitting, lockbox | Are we overfitting by tuning? | Bailey & López de Prado (2014); López de Prado (2018) |
| Purged and embargoed cross-validation | Is the out-of-sample actually clean? | López de Prado (2018) |
| Honest family-wise trial count | Deflated against how many trials, really? | the platform’s cross-axis trial ledger |
16.2 Why this is decisive for the proposal
- The synthetic simulator graduates into the gauntlet. The proposal’s simulator proves the physics on 100 synthetic stocks by injecting a known shock and confirming that a brittleness index moves. That is a necessary first step, but a synthetic positive is not evidence of a real edge. On this platform the network signal’s empirical results would run through permutation tests, the probability-of-backtest-overfitting estimate, the deflated Sharpe ratio, and a lockbox, so a phase-transition-detector claim is tested for luck, tuning, and leakage before it is believed.
- Multiple-testing rigour a graph-model paper rarely applies. A spatial-temporal graph model has enormous architectural degrees of freedom — layers, attention heads, windows, couplings. Honest trial-count deflation and the overfitting estimate are exactly the instruments for the “best-of-many-architectures” search that is the dominant failure mode of deep-learning finance.
- Report-then-enforce keeps unproven signals harmless. A network early-warning signal would ship report-only first (chapter 06); it could not move a book until it had survived attributed evidence — the same discipline that keeps every experimental factor out of real capital.
- The platform is willing to publish a negative. The connectedness spike (findings) is a documented, honest falsification. That disposition — disclosure over avoidance — is what makes a positive result from the same platform worth trusting.
16.3 The standard the network claim must clear
For a Supra-Laplacian early-warning signal to be reported as real, it would need to show — on the platform’s empirical universe, not a synthetic panel — a signal that survives permutation, deflates under honest trial counts, holds out-of-sample through a lockbox, and does so net of the composition-drift and coincidence artefacts the spike already surfaced. That is a high bar by design, and clearing it is precisely what would make the proposal’s contribution defensible.
Cross-links: platform whitepaper — mathematics of trust · reference library · proposal §6