23  Experiment: Empirical Results for the Canonical Factor Catalogue

Date: 2026-08-03 Status: partially run — the theory and configuration fields are populated; two significance fields (a naive and a Newey–West t-statistic) are specified but not yet computed. First entry under the experiments/ convention (see Experiments)

What this establishes. Re-grounding the catalogue against the live platform configuration overturned its most eye-catching prior result: the four factors once reported as the “cleanest” permutation-test signals in the library were, two months later, shown to rest on a look-ahead defect and deprecated for cause. The lesson is methodological — a stale narrative would have kept citing an artefact — and it is the one firm finding this entry records. Everything else here is provenance: what was run, what was looked up, and the exact method for the two fields still pending.

2026-08-03 note. The per-factor results this experiment produced now live in the per-factor, per-family factor catalogue as a single, self-contained record — economic justification, style, platform wiring, and empirical results in one place per factor — rather than split between the factor-models chapter and this file. This entry now holds the provenance record. Read the catalogue for current numbers; read this for how they were, or will be, produced.

23.1 Question this tests

The factor catalogue lists the platform’s canonical factors with their economic justification and style, but that alone is no evidence the platform’s own implementation of each one actually carries a signal. This experiment supplies that evidence.

It also bears directly on the factor layer as multiplex Layer 1: these factors are the proposal’s Layer 1 (the factor-exposure bipartite layer of the multiplex network). A network built on factors with no real predictive content is a network built on noise — so this is a prerequisite check for the proposal’s first layer, not just a platform-hygiene exercise.

23.2 What was run

The empirical fields in the catalogue are drawn from the existing five-sidecar diagnostic run of 26 May 2026, synthesised in the factor platform appendix — factor diagnostics (augmented Dickey–Fuller and entropy), per-factor information coefficient, factor contribution, quantile/turnover tearsheets, and the cross-sectional permutation test. No new run was executed for this entry; it re-projects that run’s results onto the catalogue, plus fresh configuration lookups against the per-factor definitions for each factor’s style, hypothesis class, expected sign, source table and column, and current lifecycle status.

That configuration lookup surfaced a correction, folded into the catalogue. Four factors — the lower Bollinger band (bb_lower_20), the 50- and 200-day moving averages (ma_50d, ma_200d), and the 20-day volume-weighted average price (vwap_20d) — were the ones the 26 May synthesis called out as having “real” pre-fix permutation-test significance. But all four were deprecated for cause on 20 July 2026 because of a look-ahead defect in their raw price-level source columns. The “cleanest permutation result in the library” turned out to be a look-ahead artefact two months later — see the bb_lower_20 entry in the catalogue for the detail. That is exactly the kind of thing re-grounding against live configuration catches and a stale narrative would not.

Caveat carried forward. The 26 May run predates three fixes that materially affect these numbers: an increase in the nightly permutation count from 100 to 1,000 with a corrected null construction, forward-return winsorisation in the tearsheet stage, and a rescaling of composite factor contributions. The catalogue’s permutation-test fields are therefore provisional until regenerated against a post-fix deep research run.

23.3 Method for the two pending fields

Neither t-statistic exists in any current sidecar — the information-coefficient, permutation, and diagnostic producers do not compute a mean-return significance test, and the platform’s heteroskedasticity-and-autocorrelation-consistent OLS routine is currently wired only to strategy net-asset-value streams, not to the per-factor return panel. The exact method (so it is reproducible once run) lives in the catalogue’s Method section:

  • Naive t-statistic: mean(F) / (std(F) / sqrt(T)) on each factor’s realised-return series — the classic Fama–MacBeth / Fama–French second-pass report.
  • Newey–West (HAC) t-statistic: the same series, intercept-only OLS via the platform’s existing HAC OLS routine (ols_hac) — the same kernel already used for the CAPM and attribution alpha estimates.
  • Why both: comparing naive against HAC on the same series is itself diagnostic — it shows how much of the naive “significance” is just serial correlation from overlapping-window, slow-moving characteristics.
  • Neither is a promotion gate. As set out in Validation & anti-overfitting, the platform’s actual gate is the information-ratio-plus-permutation test against an honest trial count.

23.4 Next steps

  1. Write the query computing the naive and HAC t-statistics per factor and fill in the two pending significance fields in the catalogue.
  2. Resolve the two duplicate pairs (beta / market_beta_252d, momentum_12_1 / momentum_12m_1m) before republishing.
  3. Regenerate against the first deep research run to post-date the three fixes above.
  4. Once stable, promote the catalogue’s empirical layer from provisional to publication-quality confidence, noting that in the catalogue’s provenance section rather than moving the content again.

23.5 Journal

Motivated by and recorded in the journal entry of 2026-08-03.