17  Reproducibility & Experiment Infrastructure

Author: Todd B. Adams Reinforces: Proposal §3 — pipeline & evaluation · Reading order: chapter 09 of this part; follows validation & anti-overfitting, precedes the literature bridge.

What this chapter establishes. Doctoral-grade research is judged partly on reproducibility. On this platform every run is an immutable, identity-stamped, fully audited experiment — not a notebook re-run whose state is lost. A network producer added to the platform inherits this audit trail automatically, so its results are reproducible by construction.


17.1 Every run is an experiment record

Property What it gives the research
One run identity per run (a dated identifier stamped end-to-end) A single identity threaded through the whole night; every artefact traces to one run
Immutable inputs and outputs A run’s data and results do not change after the fact
Per-stage structured sidecars Each producer emits a schema-versioned result blob
Harvested run bundle All sidecars rolled into one bundle the dashboard reads
Reharvest Re-derive a past run’s bundle against the current analytical layer — reproduce a result on demand
Operator dashboard Every run, stage, factor, and signal browsable
Run modes A monitoring nightly cadence versus deeper, high-permutation evidence runs

17.2 Provenance, not just results

The platform records what it received and when, not only what it concluded:

  • Audit-first ingestion. Every vendor fetch — including failures — is recorded, so “what did we actually have on date X?” is always answerable. For a phase-transition study, being able to reconstruct the exact information set available on a historical date is essential and non-negotiable.
  • Point-in-time knowledge dates. Fundamentals and slow-cadence factors are snapped to the market day their information became knowable, with the leak audited against SEC EDGAR ground truth (EDGAR findings).
  • Schema-versioned sidecars. Producers version their output, so a downstream reader degrades gracefully across runs of different vintage — the network producer would do the same.

17.3 A disciplined change process behind every producer

The platform’s own development lifecycle is itself part of the reproducibility story: capabilities are specified as design briefs, decomposed into tracked tasks, built test-first, and validated through tiered gates before merge. A research increment (chapter 07) enters through the same process, so the code that produces a published result is reviewed, tested, and traceable to a specification rather than a one-off script.

17.4 What this means for the proposal

The proposal’s evaluation engine needs to compare a network signal against realised stress across many historical dates and to re-run cleanly as the model changes. On this platform that comparison is a producer whose every run is identity-stamped, sidecar-recorded, and reharvestable — so a reviewer, or a future author, can reproduce any reported number, and the model’s evolution is a legible history rather than a lost sequence of notebook states.


Cross-links: EDGAR findings · platform capability map · proposal §3