15 Gap Analysis & Research Roadmap
Author: Todd B. Adams Reinforces: Proposal §3 and §5 · Reading order: chapter 07 of this part; follows the integrated target architecture, precedes validation & anti-overfitting.
What this chapter establishes. Here is the honest ledger of what is built versus what the research programme proposes, expressed as an increment plan in the platform’s own idiom. Nothing below hides a gap; the point is that each increment is scoped against an existing seam, and the earliest, cheapest increment is the one that most de-risks the research.
15.1 Built versus proposed — the ledger
| Capability | Proposal role | Build status | Where |
|---|---|---|---|
| Point-in-time data and analytical model | Feeds all tensors | implemented | chapter 02 |
| Stock–factor bipartite layer (L1) | Multiplex Layer 1 | implemented | chapter 04 |
| Correlation / crowding / factor networks | Layer-3 proxy, risk observability | implemented | chapter 05 |
| Fama-French, AQR, and FRED reference series | Benchmarks and regime covariates | implemented | chapter 03 |
| Validation gauntlet | Anti-overfitting for network claims | implemented | chapter 08 |
| EDGAR and FINRA ingestion machinery | Sourcing path for L2/L3 | partially built | chapter 03 |
| Supply-chain layer (L2) | Directed stock–stock graph | proposed — not yet built | — |
| 13F ownership layer (L3) | Bipartite fund–stock graph | proposed — not yet built | — |
| Supra-adjacency tensor builder | §3 subsystem 2 | proposed — not yet built | chapter 06 |
| Spatial-temporal graph engine | §3 subsystem 3 | proposed — not yet built | chapter 06 |
| Supra-Laplacian spectral analyser | §3 subsystem 4 (early warning) | proposed — not yet built | chapter 06 |
15.2 The increment plan
Ordered so the earliest step is the cheapest and most decisive — the platform’s standard “spike before you build” discipline, of which the connectedness spike is the precedent: it killed the weakest hypothesis for the price of a script.
- Increment A — spectral observability producer (report-only). A new leaf that assembles the built Layer-1 factor graph (plus the crowding correlation network) into a single-layer Laplacian and publishes its algebraic connectivity and spectral radius as a report-only sidecar. Reuses the crowding-producer template end to end. De-risks the core physics claim on real data before any graph model or new dataset.
- Increment B — supply-chain layer (L2). Extend the EDGAR client from fundamentals verification to 10-K customer–supplier edge extraction, yielding a directed stock–stock dataset. The largest data-engineering item.
- Increment C — 13F ownership layer (L3). Reuse the EDGAR client for 13F holdings, yielding a bipartite fund–stock dataset that replaces the comomentum crowding proxy with holdings ground truth.
- Increment D — multiplex tensor and graph engine. Assemble Layers 1–3 into the supra-adjacency tensor; introduce the isolated graph-learning engine on a research-day cadence.
- Increment E — phase-transition evaluation and, eventually, a gate. Validate the multi-layer spectral early-warning signal through the full gauntlet (chapter 08); only after survived evidence does it inform anything downstream (report-then-enforce).
15.3 Honest risks, stated up front
- The core hypothesis may not hold. Our own spike found single-layer connectedness coincident, not leading (connectedness-spike findings). Increment A is designed to test whether the multi-layer spectral version does better — and to fail cheaply if it does not.
- Supply-chain data is noisy. 10-K relationship extraction is a natural-language-processing problem with real precision limits; the roadmap treats it as a research task with its own evaluation, not a solved ingestion.
- The graph-learning dependency is heavy. It is quarantined to one research-day package (chapter 06) so it never touches the nightly production path.
Cross-links: proposal §3/§5 · target architecture · connectedness-spike findings