11 The Multiplex Data Substrate
Author: Todd B. Adams Reinforces: Proposal §5 — Data Requirements & Sourcing · Reading order: chapter 03 of this part; follows the end-to-end pipeline, precedes the factor layer as multiplex Layer 1.
What this chapter establishes. The proposal names three data layers. One runs in production here today; the other two are not yet built — but they are not blank slates, because adjacent sourcing machinery for each already exists in the codebase. This chapter is the honest map of which is which.
11.1 The three proposal layers, mapped to platform reality
| Proposal layer (§5) | Target variables | Platform reality today | Build status |
|---|---|---|---|
| Asset pricing & factors (Layer 1) | OHLCV, Fama-French 3/5, momentum, volatility | Fully sourced: point-in-time end-of-day prices, a governed factor library, Fama-French-5 plus momentum return series, and an AQR factor-return reference shelf | implemented |
| Supply-chain links (Layer 2) | Customer–supplier directionality, revenue share | Not ingested. But an SEC EDGAR fetch-and-parse harness already runs — the 10-K extraction path the proposal names — built for a point-in-time integrity audit | proposed — not yet built (adjacent EDGAR machinery implemented) |
| Institutional ownership (Layer 3) | 13F holdings, share counts, assets under management | Not ingested. But FINRA short-interest is sourced for positioning factors, and a comomentum producer computes a return-correlation crowding proxy — both approximations of the holdings-based ground truth | proposed — not yet built (adjacent short-interest and crowding machinery implemented) |
11.2 Layer 1 — the built layer
This is the load-bearing point of the whole part: the proposal’s hardest data layer to get right — clean, point-in-time, survivorship-corrected factor loadings — is the one the platform already delivers, and delivers empirically rather than by simulation. The proposal generates this layer synthetically (§6.1, structural stochastic differential equations); here it is measured and audited. Detail is in chapter 04.
Beyond raw factors, the platform carries the academic reference series the proposal’s benchmarking needs:
- Fama-French 5 factors plus the Carhart momentum factor, daily and monthly — the same Kenneth French Data Library source the proposal cites, already ingested — for benchmarking constructed factors against the canonical series.
- AQR factor returns (quality-minus-junk, betting-against-beta, time-series momentum) on a committed reference shelf, for cross-checking constructed factors against a second published source.
- Federal Reserve (FRED) macro series — term, credit, and volatility spreads and the risk-free curve — as regime covariates for a phase-transition model.
11.3 Layer 2 — supply chain (proposed)
The proposal’s free-sourcing path is SEC EDGAR parsing — Form 10-K extraction via natural-language processing. The platform already has an EDGAR client: a standard-library fetcher with company-identifier resolution, filings-index parsing, and XBRL company-facts extraction, built for a point-in-time integrity audit (EDGAR findings). Extending it from fundamentals verification to customer–supplier edge extraction is a scoped ingestion task, not new infrastructure. It is the least-mature of the three layers and, honestly, the largest single data-engineering item in the roadmap (chapter 07).
11.4 Layer 3 — institutional ownership (proposed)
The proposal wants 13F holdings to model ownership crowding. The platform approaches the same construct from two directions today, each an approximation rather than the ground truth:
- FINRA short interest is ingested and drives positioning-style factors (days-to-cover, short interest as a percentage of float) — the short side of ownership positioning, point-in-time-snapped to settlement plus dissemination lag.
- A comomentum producer infers crowding from return correlations rather than holdings — a proxy that a 13F layer would replace with ground truth (chapter 05).
A 13F ingestion path would reuse the same EDGAR client as Layer 2, so the two proposed layers share a sourcing route.
11.5 Why the gaps are credible, not hand-waving
Every proposed layer above attaches to a mechanism that is already built: EDGAR ingestion, FINRA sourcing, crowding computation, and a declarative dataset framework (dataset reference) that makes adding a new dataset a configuration change rather than a rewrite. The research adds datasets and a model; it does not add a data platform.
Cross-links: dataset reference · EDGAR findings · proposal §5 · factor layer