15  Gap Analysis & Research Roadmap

Author: Todd B. Adams Reinforces: Proposal §3 and §5 · Reading order: chapter 07 of this part; follows the integrated target architecture, precedes validation & anti-overfitting.

What this chapter establishes. Here is the honest ledger of what is built versus what the research programme proposes, expressed as an increment plan in the platform’s own idiom. Nothing below hides a gap; the point is that each increment is scoped against an existing seam, and the earliest, cheapest increment is the one that most de-risks the research.


15.1 Built versus proposed — the ledger

Capability Proposal role Build status Where
Point-in-time data and analytical model Feeds all tensors implemented chapter 02
Stock–factor bipartite layer (L1) Multiplex Layer 1 implemented chapter 04
Correlation / crowding / factor networks Layer-3 proxy, risk observability implemented chapter 05
Fama-French, AQR, and FRED reference series Benchmarks and regime covariates implemented chapter 03
Validation gauntlet Anti-overfitting for network claims implemented chapter 08
EDGAR and FINRA ingestion machinery Sourcing path for L2/L3 partially built chapter 03
Supply-chain layer (L2) Directed stock–stock graph proposed — not yet built —
13F ownership layer (L3) Bipartite fund–stock graph proposed — not yet built —
Supra-adjacency tensor builder §3 subsystem 2 proposed — not yet built chapter 06
Spatial-temporal graph engine §3 subsystem 3 proposed — not yet built chapter 06
Supra-Laplacian spectral analyser §3 subsystem 4 (early warning) proposed — not yet built chapter 06

15.2 The increment plan

Ordered so the earliest step is the cheapest and most decisive — the platform’s standard “spike before you build” discipline, of which the connectedness spike is the precedent: it killed the weakest hypothesis for the price of a script.

  1. Increment A — spectral observability producer (report-only). A new leaf that assembles the built Layer-1 factor graph (plus the crowding correlation network) into a single-layer Laplacian and publishes its algebraic connectivity and spectral radius as a report-only sidecar. Reuses the crowding-producer template end to end. De-risks the core physics claim on real data before any graph model or new dataset.
  2. Increment B — supply-chain layer (L2). Extend the EDGAR client from fundamentals verification to 10-K customer–supplier edge extraction, yielding a directed stock–stock dataset. The largest data-engineering item.
  3. Increment C — 13F ownership layer (L3). Reuse the EDGAR client for 13F holdings, yielding a bipartite fund–stock dataset that replaces the comomentum crowding proxy with holdings ground truth.
  4. Increment D — multiplex tensor and graph engine. Assemble Layers 1–3 into the supra-adjacency tensor; introduce the isolated graph-learning engine on a research-day cadence.
  5. Increment E — phase-transition evaluation and, eventually, a gate. Validate the multi-layer spectral early-warning signal through the full gauntlet (chapter 08); only after survived evidence does it inform anything downstream (report-then-enforce).

15.3 Honest risks, stated up front

  • The core hypothesis may not hold. Our own spike found single-layer connectedness coincident, not leading (connectedness-spike findings). Increment A is designed to test whether the multi-layer spectral version does better — and to fail cheaply if it does not.
  • Supply-chain data is noisy. 10-K relationship extraction is a natural-language-processing problem with real precision limits; the roadmap treats it as a research task with its own evaluation, not a solved ingestion.
  • The graph-learning dependency is heavy. It is quarantined to one research-day package (chapter 06) so it never touches the nightly production path.

Cross-links: proposal §3/§5 · target architecture · connectedness-spike findings