29 2026-08-03
I restructured the research repository from a single business-whitepaper README into a research-notebook shape: Vision → Literature Review → Research Questions → Methodology → Experiments → Publications → Research Journal → Code → Data Catalog → Academic References.
What moved and why:
- The original whitepaper content moved to whitepaper unchanged — it’s still the right document for a potential user of the platform, but it isn’t a research vision statement, so it no longer owns the top-level README.
phd/,research-platform/,platform/, andlibrary/were left in place — they’re already solid and heavily cross-linked, so the new top-level folders (methodology/,experiments/,data-catalog/,code/,references/,journal/) are thin indexes into them rather than copies.- Genuinely new: research-questions.md (extracted from the research proposal so the current hypothesis doesn’t require reading the full proposal), this journal, and references/bibliography.md (DOI-normalised, distinct from the library’s prose summaries).
Open follow-ups, tracked so they don’t get lost:
library/research-map.mdis a stub — the actual paper-relationship graph hasn’t been generated yet (candidate: the knowledge-graph skill).data-catalog/licensing.mdanddata-catalog/refresh-schedules.mdare scaffolds — licensing terms per data vendor and refresh cadence aren’t documented centrally yet, even though the platform enforces them operationally.experiments/has no entries yet under the new per-experiment convention; existing results still live inresearch-platform/findings/.
29.1 Later the same day: first experiments/ entry — factor catalogue results
I added the factor-catalogue empirical-results experiment, the first entry under this morning’s new convention. Motivation: the canonical 13-factor catalogue in factor-models-and-benchmarks §10 states economic justification and style but no empirical evidence — and these factors are literally the proposal’s multiplex Layer 1, so “does the platform’s own implementation of each one actually show a signal” isn’t optional context, it’s a prerequisite.
What happened:
- I re-projected the existing five-sidecar run of 26 May 2026 (already synthesised in
platform/factors.md) onto the 13 catalogue rows, and pulled style, hypothesis class, and expected sign from the factor configuration directly rather than guessing. - I added two columns at the operator’s request: a naive mean-return t-statistic and a Newey–West (heteroskedasticity-and-autocorrelation-consistent) t-statistic, computed the same way the factor-models note already computes its attribution alpha — reusing the existing HAC ordinary-least-squares estimator on the factor-return panel instead of the strategy net-asset-value streams it is currently wired to.
- Both t-statistic columns are specified but not computed — no sidecar currently produces them, and I didn’t want to hand-wave numbers into a repo that otherwise polices this hard. I marked them pending with the exact method so a future pass (or a script) can fill them in mechanically.
- I surfaced two things worth acting on before this table is trusted:
- the market-beta pair and the 12-minus-1 momentum pair are known-identical duplicates per the run synthesis — counting both sides would inflate the honest trial count the platform’s overfitting audit otherwise tracks; (2) the lower-Bollinger-band factor and the atomic residual-momentum factor have real evidence (the cleanest permutation-test pass, and the highest information ratio in the run respectively) but aren’t in the §10 catalogue at all, while the deprecated residual-momentum composite occupies conceptual space that could be confused with its atomic sibling.
Open follow-up: write the actual naive-and-HAC t-statistic computation and rerun this once the three pending factor fixes have all landed in a weekly deep run (full permutation counts), so the permutation-test column stops being pre-fix evidence.
29.2 Later still: merged the catalogue into one self-contained, per-family document
The operator asked to restructure again: list each factor in its own section, grouped by family, list-style rather than wide tables, each factor self-contained (justification, source table/column, empirical results together) — and to stop splitting the catalogue across the factor-models note §10 and the experiments/ entry.
- I created the factor catalogue: 13 family sections, each factor as its own subsection with style, source
table.column, hypothesis class and expected sign, lifecycle status, information ratio, permutation-test result, and the two pending t-statistics — all pulled from the same config lookups and run synthesis as the earlier pass, just reshaped and consolidated. - The factor-models note §10 is now a short pointer to the new file instead of the table.
- I slimmed the factor-catalogue experiment entry down to a provenance record (what was run, the config-lookup correction, the pending-field method) rather than a duplicate of the data — the numbers live in one place now.
- I added two factors the original 13-row table didn’t cover but the run data called for: the atomic residual-momentum factor (highest information ratio in the run) under Momentum, and the lower-Bollinger-band factor under short-horizon reversal — kept specifically as a cautionary entry, since re-checking current config revealed it (and the 50- and 200-day moving averages and the 20-day volume-weighted average price) had been deprecated for cause in July for a look-ahead defect, after being the run’s cleanest permutation-test result. A good reminder that a clean permutation p-value is necessary, not sufficient.