17 Reproducibility & Experiment Infrastructure
Author: Todd B. Adams Reinforces: Proposal §3 — pipeline & evaluation · Reading order: chapter 09 of this part; follows validation & anti-overfitting, precedes the literature bridge.
What this chapter establishes. Doctoral-grade research is judged partly on reproducibility. On this platform every run is an immutable, identity-stamped, fully audited experiment — not a notebook re-run whose state is lost. A network producer added to the platform inherits this audit trail automatically, so its results are reproducible by construction.
17.1 Every run is an experiment record
| Property | What it gives the research |
|---|---|
| One run identity per run (a dated identifier stamped end-to-end) | A single identity threaded through the whole night; every artefact traces to one run |
| Immutable inputs and outputs | A run’s data and results do not change after the fact |
| Per-stage structured sidecars | Each producer emits a schema-versioned result blob |
| Harvested run bundle | All sidecars rolled into one bundle the dashboard reads |
| Reharvest | Re-derive a past run’s bundle against the current analytical layer — reproduce a result on demand |
| Operator dashboard | Every run, stage, factor, and signal browsable |
| Run modes | A monitoring nightly cadence versus deeper, high-permutation evidence runs |
17.2 Provenance, not just results
The platform records what it received and when, not only what it concluded:
- Audit-first ingestion. Every vendor fetch — including failures — is recorded, so “what did we actually have on date X?” is always answerable. For a phase-transition study, being able to reconstruct the exact information set available on a historical date is essential and non-negotiable.
- Point-in-time knowledge dates. Fundamentals and slow-cadence factors are snapped to the market day their information became knowable, with the leak audited against SEC EDGAR ground truth (EDGAR findings).
- Schema-versioned sidecars. Producers version their output, so a downstream reader degrades gracefully across runs of different vintage — the network producer would do the same.
17.3 A disciplined change process behind every producer
The platform’s own development lifecycle is itself part of the reproducibility story: capabilities are specified as design briefs, decomposed into tracked tasks, built test-first, and validated through tiered gates before merge. A research increment (chapter 07) enters through the same process, so the code that produces a published result is reviewed, tested, and traceable to a specification rather than a one-off script.
17.4 What this means for the proposal
The proposal’s evaluation engine needs to compare a network signal against realised stress across many historical dates and to re-run cleanly as the model changes. On this platform that comparison is a producer whose every run is identity-stamped, sidecar-recorded, and reharvestable — so a reviewer, or a future author, can reproduce any reported number, and the model’s evolution is a legible history rather than a lost sequence of notebook states.
Cross-links: EDGAR findings · platform capability map · proposal §3