5  Factor Catalogue

In factor-based investing, a factor is a characteristic or attribute of securities that has historically been associated with differences in returns and risk across a broad set of investments. Common equity factors include:

A strategy is the investment process used to gain exposure to one or more factors: the factor is the source of a return or risk premium, and the strategy is the implementation method used to capture it. A single factor can be captured by several distinct strategies — a value manager, for instance, might define “cheap” using price-to-book, price-to-earnings, enterprise-value-to-EBITDA, or a composite of several valuation metrics; each is attempting to capture the same value factor, but each is a different strategy.

Factor Illustrative strategy
Value Buy the cheapest 20% of stocks ranked by price-to-book.
Momentum Buy stocks with the strongest past 12-month returns.
Quality Construct a portfolio of highly profitable companies with low debt.
Low volatility Hold stocks with the lowest historical volatility.

The research literature further distinguishes three related concepts:

In practice, a factor is the underlying return driver, while a strategy is the set of portfolio-construction rules used to obtain exposure to it. A strategy may target one factor or combine several, and multiple strategies can seek exposure to the same factor.

For every factor in this catalogue we state the economic hypothesis it encodes, the sign and horizon of its measured predictive signal, and where that signal currently sits in the promotion funnel — with the finding leading and the construction demoted to a method note. The measurements reported here are gross and in-sample unless explicitly labelled otherwise; none should be read as having cleared the corrected, net-of-cost, multiplicity-adjusted promotion gate.

5.1 Current run snapshot

Auto-generated snapshot from the nightly deep run of 2026-09-25, refreshed 2026-09-27. Only the content between the LEMONADE:RUN-SNAPSHOT markers is regenerated each deep run; every hand-authored section outside them is left untouched.

The 30‑factor deep run on 25 September 2026 produced several in‑sample IC‑IR signals, with value_composite leading at +0.366 and cash_based_operating_profitability close behind at +0.349. Other notable signals include piotroski (+0.305), sue (+0.228), share_turnover (+0.226) and earnings_yield (+0.212). All of these factors were flagged as “watch” or “experimental” and none were deemed ready for transition, with permutation p‑values ranging from 0.038 to 0.792. The only factor that passed validation was earnings_growth_qoq.

  • Value and quality factors show the strongest IC‑IR, but none are ready for deployment.
  • Share turnover has the lowest permutation p‑value, indicating the most statistically robust signal among the watch list.

5.1.1 Value

  • value_composite — lifecycle: watch · IC-IR @21d: +0.366 · corrected IC-IR: +0.411 · permutation p: 0.214 · gate verdict: no transition
  • earnings_yield — lifecycle: watch · IC-IR @21d: +0.212 · corrected IC-IR: +0.311 · permutation p: 0.703 · gate verdict: no transition
  • fcf_yield — lifecycle: watch · IC-IR @21d: +0.211 · corrected IC-IR: +0.273 · permutation p: 0.435 · gate verdict: no transition
  • pe_ratio — lifecycle: experimental · IC-IR @21d: -0.135 · corrected IC-IR: -0.177 · permutation p: 0.457 · gate verdict: no transition
  • pb_ratio — lifecycle: experimental · IC-IR @21d: -0.302 · corrected IC-IR: -0.321 · permutation p: 0.976 · gate verdict: no transition
  • ps_ratio — lifecycle: watch · IC-IR @21d: -0.460 · corrected IC-IR: -0.477 · permutation p: 0.970 · gate verdict: no transition

5.1.2 Quality

  • cash_based_operating_profitability — lifecycle: watch · IC-IR @21d: +0.349 · corrected IC-IR: +0.415 · permutation p: 0.399 · gate verdict: no transition
  • piotroski — lifecycle: watch · IC-IR @21d: +0.305 · corrected IC-IR: +0.399 · permutation p: 0.679 · gate verdict: no transition
  • gross_profitability — lifecycle: experimental · IC-IR @21d: +0.198 · corrected IC-IR: +0.189 · permutation p: 0.265 · gate verdict: no transition
  • roic_spread — lifecycle: experimental · IC-IR @21d: +0.087 · corrected IC-IR: +0.134 · permutation p: 0.856 · gate verdict: no transition
  • roic — lifecycle: experimental · IC-IR @21d: +0.023 · corrected IC-IR: +0.061 · permutation p: 0.816 · gate verdict: no transition
  • fundamental_quality_composite — lifecycle: experimental · IC-IR @21d: -0.093 · corrected IC-IR: -0.055 · permutation p: 0.727 · gate verdict: no transition
  • accruals — lifecycle: experimental · IC-IR @21d: -0.134 · corrected IC-IR: -0.039 · permutation p: 1.000 · gate verdict: no transition

5.1.3 Growth

  • sue — lifecycle: watch · IC-IR @21d: +0.228 · corrected IC-IR: +0.303 · permutation p: 0.792 · gate verdict: no transition
  • earnings_growth_qoq — lifecycle: validated · IC-IR @21d: +0.080 · corrected IC-IR: +0.095 · permutation p: 0.002 · gate verdict: no transition

5.1.4 Liquidity

  • share_turnover — lifecycle: experimental · IC-IR @21d: +0.226 · corrected IC-IR: +0.188 · permutation p: 0.038 · gate verdict: no transition

5.1.5 Momentum

  • momentum_12_1_vol_scaled — lifecycle: watch · IC-IR @21d: +0.199 · corrected IC-IR: +0.236 · permutation p: 0.655 · gate verdict: no transition
  • momentum_12m_1m — lifecycle: watch · IC-IR @21d: +0.168 · corrected IC-IR: +0.191 · permutation p: 0.952 · gate verdict: no transition
  • industry_momentum — lifecycle: experimental · IC-IR @21d: +0.149 · corrected IC-IR: +0.159 · permutation p: 0.002 · gate verdict: no transition
  • px_vs_ma_200d — lifecycle: experimental · IC-IR @21d: +0.074 · corrected IC-IR: +0.101 · permutation p: 0.824 · gate verdict: no transition
  • seasonality_same_month — lifecycle: experimental · IC-IR @21d: +0.049 · corrected IC-IR: +0.061 · permutation p: 1.000 · gate verdict: no transition
  • px_vs_vwap_20d — lifecycle: deprecated · IC-IR @21d: -0.037 · corrected IC-IR: -0.021 · permutation p: — · gate verdict: no transition
  • px_vs_ma_50d — lifecycle: deprecated · IC-IR @21d: -0.069 · corrected IC-IR: -0.055 · permutation p: — · gate verdict: no transition

5.1.6 Volatility

  • bb_pct — lifecycle: experimental · IC-IR @21d: -0.079 · corrected IC-IR: -0.059 · permutation p: 1.000 · gate verdict: no transition
  • max_ret_21d — lifecycle: experimental · IC-IR @21d: -0.160 · corrected IC-IR: -0.277 · permutation p: 0.016 · gate verdict: no transition
  • volatility_30d — lifecycle: deprecated · IC-IR @21d: -0.165 · corrected IC-IR: -0.288 · permutation p: — · gate verdict: no transition

5.1.7 Positioning

  • days_to_cover — lifecycle: experimental · IC-IR @21d: -0.081 · corrected IC-IR: -0.095 · permutation p: 0.856 · gate verdict: no transition
  • short_pct_float — lifecycle: experimental · IC-IR @21d: -0.322 · corrected IC-IR: -0.362 · permutation p: 0.982 · gate verdict: no transition

5.1.8 Investment

  • asset_growth — lifecycle: experimental · IC-IR @21d: -0.097 · corrected IC-IR: -0.094 · permutation p: 0.369 · gate verdict: no transition
  • net_share_issuance — lifecycle: watch · IC-IR @21d: -0.394 · corrected IC-IR: -0.486 · permutation p: 0.657 · gate verdict: no transition

The information ratio at the 21-day horizon and the permutation p-value are the gross, pre-haircut, in-sample figures from this single run’s cross-sectional information-coefficient and permutation-test producers; the corrected information ratio applies the survivorship/overlap haircut. Theory, provenance, and per-family narratives are in the hand-authored sections below.

5.2 Provenance

  • Theory fields (economic justification, style, anchor citation, hypothesis class, expected sign, source table/column, lifecycle status) are drawn directly from the platform’s per-factor configuration as of early August 2026.
  • Universe and horizon. Unless otherwise noted, every information-coefficient figure in the per-family sections below is estimated on the platform’s investable US equity cross-section at a 21-trading-day forward horizon. Per-factor sample sizes — the number of cross-sectional dates entering each estimate — are not itemised in the run synthesis we work from; we flag them as not-yet-itemised rather than guess a value.
  • Empirical fields in the per-family sections come from a diagnostic run of May 2026, synthesised in the factor library appendix. That run predates three subsequent corrections that materially change these numbers: raising the nightly permutation count from 100 to 1000 with a corrected null construction, winsorising forward returns in the quantile/tearsheet stage, and rescaling composite factor contributions. We therefore treat every gross, in-sample permutation figure below as provisional pending re-estimation.
  • Two diagnostic columns — a naive t-statistic and a heteroskedasticity-and-autocorrelation-consistent (Newey–West) t-statistic — are specified but not yet computed. No producer currently emits them; the Method section states exactly how they are to be filled in. They are marked not yet computed throughout rather than guessed.
  • Re-grounding against the live configuration surfaced a correction to the run narrative. The May 2026 synthesis described the 20-day lower Bollinger band, the 50- and 200-day moving averages, and the 20-day volume-weighted average price as the four factors with “real” pre-correction permutation significance. All four were subsequently deprecated for cause on 20 July 2026: they are raw price-level source columns carrying a look-ahead defect in the price series they are built on, not a validated edge. The lower-band factor’s entry below records this reversal; the scale-free replacement, the Bollinger percent-b factor (§13), supersedes it.
  • Known duplicates — resolved 5 August 2026. The trailing-beta factor and its legacy alias, and the vol-scaled 12-1 momentum factor and its legacy alias, are numerically identical to fourteen digits; neither legacy alias ever had its own configuration file. The legacy aliases are retired; market_beta_252d and momentum_12m_1m are the canonical survivors and are the only ones counted in the platform’s honest family-wise trial ledger.

5.3 Method: the pending t-stat columns

Every factor entry below carries two not yet computed diagnostic fields. Both are specified the same way, stated once here rather than repeated for each factor:

  • Naive t-statistic — the classic Fama–MacBeth / Fama–French second-pass report. Take the daily factor-return series \(F_t\) for the factor from the platform’s factor-return panel (Factor Models and Benchmarks §2.2) and compute \(t_{\text{naive}} = \mathrm{mean}(F) / (\mathrm{std}(F) / \sqrt{T})\).
  • HAC (Newey–West) t-statistic — regress the same series on a constant only using the platform’s existing Newey–West heteroskedasticity-and-autocorrelation-consistent ordinary-least-squares estimator (Bartlett weights, lag \(L = \lfloor 4\,(n/100)^{2/9}\rfloor\), Newey–West 1994) — the identical kernel already used for the capital-asset-pricing-model and attribution alphas. The intercept is the mean factor return; its HAC standard error replaces the naive \(\mathrm{std}(F)/\sqrt{T}\) denominator.
  • Why both. Daily factor returns built on slow-moving or overlapping-window characteristics (12-1 momentum, vol-scaling, time-series momentum) are serially correlated by construction, so the naive t-statistic overstates significance. The gap between the two is itself diagnostic.
  • Neither is a promotion gate. As established in Validation and Anti-Overfitting, a t-statistic — naive or HAC — is exactly the single-test statistic the Harvey–Liu–Zhu multiple-testing critique targets. The platform’s actual promotion gate combines the information ratio with a cross-sectional permutation p-value assessed against an honest trial count; these two columns are diagnostic context, not a second gate.
  • Not yet wired. The HAC estimator today runs only against strategy net-asset-value streams, not the per-factor return panel; computing these columns requires a small new script, not a configuration change.

5.4 1. Market

Compensation for bearing undiversifiable market risk — the one factor everyone must hold; the capital-asset-pricing-model premium. Sharpe 1964 [#116], Lintner 1965 [#117].

5.4.1 market_beta_252d

  • Economic hypothesis / expected sign: not recorded in configuration.
  • Measured signal (gross, in-sample): not itemised in the run synthesis.
  • Cross-sectional permutation test (pre-correction): p ≈ 1.000 (at the ceiling), with an unbiased information ratio of roughly −6.8 — no evidence of a standalone edge.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental — the canonical survivor of the trailing-beta duplicate pair (resolved 5 August 2026); the legacy alias is retired.
  • Style: beta.
  • Method note: trailing 252-day rolling ordinary-least-squares slope against the market proxy (fact_eod.market_beta_252d_f); daily cadence; the diagnostics place it in the “healthy” quadrant (stationary, entropy ≥ 3). \[\beta_{i,t} = \dfrac{\mathrm{Cov}_{252d}(r_i,\, r_{SPY})}{\mathrm{Var}_{252d}(r_{SPY})}\] The factor and its legacy alias are identical to fourteen digits; the alias carried no independent configuration and is now retired.

5.5 2. Size

Small caps earn a premium for illiquidity, distress, and limited analyst coverage; partly compensation, partly a limits-to-arbitrage effect. Banz 1981 [#125], Fama–French 1992 [#55].

5.5.1 market_cap

  • Economic hypothesis / expected sign: not recorded in configuration.
  • Measured signal (gross, in-sample): information ratio at 21 days ≈ −0.5 (inverted — part of the low-volatility/size cluster; the premium is on the short side).
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: size.
  • Method note: year-end close times shares outstanding (fact_valuation_annual.market_cap_f); annual cadence — the short-sample caveat applies, since stationarity and entropy diagnostics have little power at roughly a dozen annual observations. \[\text{MktCap}_{i,t} = P^{YE}_{i,t} \times \text{SharesOut}_{i,t}\]

5.6 3. Value

Cheap (high earnings-, book-, and free-cash-flow yield) stocks out-earn expensive ones — a risk premium for distressed/low-growth firms and mispricing from extrapolation. Fama–French 1992/1993 [#55][#44].

5.6.1 value_composite

  • Economic hypothesis / expected sign: risk premium, positive.
  • Measured signal (gross, in-sample): information ratio at 21 days = +0.75 — the highest in the May 2026 diagnostic run. The sign flips at a one-day horizon (do not trade intraday) — the classic value-investor pattern, correctly signed only at a month-plus horizon. This gross figure is provisional pending re-estimation.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental, promotion candidate.
  • Style: value.
  • Method note: an equal-weighted z-score blend of four value yields — earnings, book, sales, and free-cash-flow (fact_fundamental_annual.value_composite_f); annual cadence. The configuration names the four component yields and calls it a “blend”; the equal weighting shown is the simplest reading and is not independently confirmed by a documented weight vector. \[\text{ValueComposite}_{i,t} = \dfrac{1}{4}\displaystyle\sum_{k \,\in\, \{EY,\; B/M,\; S/P,\; FCF/P\}} z\!\left(x_{k,i,t}\right)\]

5.6.2 earnings_yield

  • Economic hypothesis / expected sign: risk premium, positive.
  • Measured signal (gross, in-sample): sign-flip — negative at a one-day horizon, positive at 21 days and beyond.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: value.
  • Method note: diluted earnings per share over year-end close (fact_valuation_annual.earnings_yield_f), the inverse of the price-to-earnings ratio; annual cadence. \[EY_{i,t} = \dfrac{EPS^{diluted}_{i,t}}{P^{YE}_{i,t}}\]

5.6.3 pe_ratio

  • Economic hypothesis / expected sign: risk premium, negative.
  • Measured signal (gross, in-sample): information ratio at 21 days between −0.5 and −0.7 (the inverted group, alongside the volatility factors, market cap, and Amihud illiquidity).
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: value.
  • Method note: year-end close over diluted earnings per share (fact_valuation_annual.pe_ratio_f); annual cadence. \[PE_{i,t} = \dfrac{P^{YE}_{i,t}}{EPS^{diluted}_{i,t}}\]

5.6.4 pb_ratio

  • Economic hypothesis / expected sign: risk premium, negative.
  • Measured signal (gross, in-sample): information ratio not itemised. The raw top-minus-bottom-quantile spread of +37.40 at the 21-day horizon reported in the May 2026 synthesis is a pre-winsorisation artefact — un-winsorised cross-section outlier dominance, not a real spread — and is bounded once forward returns are winsorised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: value.
  • Method note: market capitalisation over total stockholders’ equity (fact_valuation_annual.pb_ratio_f); annual cadence. \[PB_{i,t} = \dfrac{P^{YE}_{i,t} \times \overline{Shs}^{diluted}_{i,t}}{\text{StockholdersEquity}_{i,t}}\]

5.6.5 ps_ratio

  • Economic hypothesis / expected sign: risk premium, negative.
  • Measured signal (gross, in-sample): information ratio not itemised. The raw quantile spread of +44.78 at 21 days is, like the book-to-price case, a pre-winsorisation artefact, bounded once forward returns are winsorised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: value.
  • Method note: market capitalisation over revenue (fact_valuation_annual.ps_ratio_f); annual cadence. \[PS_{i,t} = \dfrac{P^{YE}_{i,t} \times \overline{Shs}^{diluted}_{i,t}}{\text{Revenue}_{i,t}}\]

5.6.6 pfcf_ratio

  • Economic hypothesis / expected sign: risk premium, negative.
  • Measured signal (gross, in-sample): information ratio not itemised. The raw quantile spread of −36.78 at 21 days is again a pre-winsorisation artefact, bounded once forward returns are winsorised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: value.
  • Method note: market capitalisation over free cash flow (fact_valuation_annual.pfcf_ratio_f), the inverse of free-cash-flow yield; annual cadence. \[PFCF_{i,t} = \dfrac{P^{YE}_{i,t} \times \overline{Shs}^{diluted}_{i,t}}{\text{FCF}_{i,t}}\]

5.7 4. Momentum

Past 12-1 winners keep winning — under-reaction to news and delayed information diffusion (behavioural). Jegadeesh–Titman 1993 [#49], Carhart 1997 [#57].

5.7.1 momentum_12m_1m

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): information ratio at 21 days = +0.586; provisional pending re-estimation.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental — the canonical survivor of the 12-1 momentum duplicate pair (resolved 5 August 2026); the legacy alias is retired.
  • Style: momentum.
  • Method note: twelve-month cumulative return skipping the most recent month (fact_eod.momentum_12_1_f); daily cadence. The factor and its legacy alias are identical to fourteen digits; the alias survives only as a raw feature-column name inside the backtest engine and is retired as a factor. \[\text{Mom}_{i,t} = \dfrac{P_{i,t-21}}{P_{i,t-252}} - 1\]

5.7.2 momentum_12_1_vol_scaled

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): information ratio at 21 days = +0.629; provisional pending re-estimation. Vol-scaling is expected to help it survive the overlap correction once the corrected permutation null is applied.
  • Cross-sectional permutation test: at the ceiling pre-correction.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental, promotion candidate.
  • Style: momentum.
  • Method note: the 12-1 trend divided by trailing realised volatility (fact_eod.momentum_12_1_vol_scaled_f; worked example in Factor Models and Benchmarks §6); daily cadence. The same source value also backs the distinct time-series-momentum entry (§12) — identical value, different evaluation path: this factor is scored cross-sectionally, the §12 entry is exempt from the funnel. The strategies that trade each are documented in the Strategy Catalogue. \[\text{Mom}^{vol}_{i,t} = \dfrac{\text{Mom}_{i,t}}{\sigma^{126d}_{i,t}\sqrt{252}}\] where \(\text{Mom}_{i,t}\) is the 12-1 return above and \(\sigma^{126d}_{i,t}\) is the trailing 126-day daily-return standard deviation.

5.7.3 residual_momentum_252d_eod (not in the original 13-row citation table — included here because it carries the strongest result in the run)

  • Economic hypothesis / expected sign: not recorded in configuration (market-neutral momentum).
  • Measured signal (gross, in-sample): information ratio at 21 days = +0.724 — the strongest in the entire May 2026 diagnostic run; provisional pending re-estimation.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental — the atomic sibling of a deprecated composite of the same name, which carries zero rows because no service computes its composite recipe. Consumers should use this atomic factor, not the composite.
  • Style: momentum.
  • Method note: the residual of a 252-day rolling ordinary-least-squares of return on an equal-weighted benchmark, summed over the window \([-252:-21]\) and z-scaled (fact_eod.residual_momentum_252d_f); daily cadence; the diagnostics place it in the “healthy” quadrant. It is a candidate for a fourteenth citation row given this result; the anchor citation is Blitz–Huij–Martens on residual momentum [#50]. Traded by the residual-momentum strategy (Strategy Catalogue §4). \[\varepsilon_{i,\tau} = r_{i,\tau} - \big(\hat\alpha_i + \hat\beta_i\, r_{EW,\tau}\big)\] from a 252-day rolling regression, then \[\text{ResMom}_{i,t} = z\!\left(\displaystyle\sum_{\tau=t-252}^{t-21}\varepsilon_{i,\tau}\right)\]

5.8 5. Profitability / Quality

Profitable, safe, well-managed firms (high gross-profitability, return on invested capital, quality-minus-junk) out-earn junk — a premium for quality that markets under-price. Novy-Marx 2013 [#92], Fama–French 2015 [#56], Asness–Frazzini–Pedersen 2019 [#113].

5.8.1 qmj_composite

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: in the floor bucket — a null-mean information ratio near zero and a degenerate permutation null, typical of slow-moving annual fundamentals.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental. Despite its name it is atomic by design — a single pre-blended z-score column computed upstream, not a true composite, so it always registers as a single input to the factor-contribution decomposition rather than as a composite.
  • Style: quality.
  • Method note: the Asness–Frazzini–Pedersen Quality-Minus-Junk blend across profitability, growth, safety, and payout (fact_fundamental_annual.qmj_composite_f); annual cadence — the short-sample caveat applies. \[QMJ_{i,t} = z(\text{Profitability}_{i,t}) + z(\text{Growth}_{i,t}) + z(\text{Safety}_{i,t}) + z(\text{Payout}_{i,t})\]

5.8.2 roic_spread

  • Economic hypothesis / expected sign: risk premium, positive.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: quality.
  • Method note: current-year return on invested capital minus the weighted average cost of capital, winsorised and industry-z-scored to \([0,1]\) (fact_moat_annual.roic_spread_f); annual cadence — the short-sample caveat applies. \[\text{ROICSpread}_{i,t} = \text{clip}_{[0,1]}\Big(z_{industry}\big(\text{ROIC}_{i,t} - \text{WACC}_{i,t}\big)\Big)\] where \(\text{ROIC} = \text{NOPAT}/\text{InvestedCapital}\).

5.8.3 gross_profitability

  • Economic hypothesis / expected sign: risk premium, positive.
  • Measured signal (gross, in-sample): not itemised in the May 2026 synthesis.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: quality.
  • Method note: gross profit over total assets (fact_moat_annual.gross_profitability_f, Novy-Marx 2013); annual cadence — the short-sample caveat applies. \[GP_{i,t} = \dfrac{\text{GrossProfit}_{i,t}}{\text{TotalAssets}_{i,t}}\]

5.8.4 piotroski

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): sign-flip — negative at a one-day horizon, positive at 21 days and beyond.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: quality.
  • Method note: a nine-criterion financial-strength score (fact_fundamental_annual.piotroski_f); annual cadence. \[F_{i,t} = \displaystyle\sum_{k=1}^{9} \mathbb{1}[\text{criterion}_k \text{ met}] \in [0,9]\]

5.9 6. Investment / Accruals

Firms that invest conservatively and have low accruals out-earn aggressive investors — over-investment and earnings-quality mispricing. Sloan 1996 [#83], Cooper–Gulen–Schill 2008 [#103], Fama–French 2015 [#56].

5.9.1 accruals

  • Economic hypothesis / expected sign: behavioural, negative.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: quality.
  • Method note: Sloan’s (1996) total-accruals-to-total-assets (fact_moat_annual.accruals_f); annual cadence — the short-sample caveat applies. The original family mapping also names a quarterly earnings-growth factor for this family, but no configuration file was found under that exact name; the discrepancy remains to be reconciled against whatever earnings-growth factor is actually registered. \[\text{Accruals}_{i,t} = \dfrac{\text{NetIncome}_{i,t} - \text{CFO}_{i,t}}{\text{TotalAssets}_{i,t}}\]

5.10 7. Low volatility / Defensive

Low-risk stocks earn higher risk-adjusted returns — leverage constraints and lottery-preference bid up high-vol names (the low-volatility anomaly / betting-against-beta). Ang et al. 2006 [#108], Frazzini–Pedersen 2014 [#93], Novy-Marx 2014 [#10].

5.10.1 volatility_30d

  • Economic hypothesis / expected sign: behavioural, negative.
  • Measured signal (gross, in-sample): information ratio at 21 days between −0.5 and −0.7, alongside the 60-, 126-, and 252-day volatility factors — part of the low-volatility/size inverted cluster; the premium is on the short leg, so the sign inverts on promotion.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: the May 2026 diagnostic entry reads experimental with a strong signal; the most recent deep run marks it deprecated, a divergence awaiting operator reconciliation.
  • Style: volatility.
  • Method note: the standard deviation of daily log returns over 30 sessions, annualised by \(\sqrt{252}\) (fact_eod.volatility_30d_f); daily cadence. \[\sigma^{30d}_{i,t} = \text{STDDEV}_{30d}\big(\ln(P_{i,\tau}/P_{i,\tau-1})\big) \times \sqrt{252}\]

5.11 8. Liquidity / Illiquidity

Illiquid stocks (high Amihud price-impact, low dollar volume) require a return premium for the cost and risk of trading them. Amihud 2002 [#48].

5.11.1 amihud

  • Economic hypothesis / expected sign: not recorded in configuration.
  • Measured signal (gross, in-sample): in the inverted group (with the price-to-earnings ratio, the volatility factors, and market cap).
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: liquidity.
  • Method note: the mean of absolute daily return over dollar volume across a trailing window (fact_eod.amihud_f); daily cadence; the diagnostics place it in the “healthy” quadrant. \[\text{Amihud}_{i,t} = \dfrac{1}{N}\displaystyle\sum_{\tau=t-N+1}^{t} \dfrac{|r_{i,\tau}|}{\$\text{Vol}_{i,\tau}}\]

5.11.2 adv_dollar_20d

  • Economic hypothesis / expected sign: not recorded in configuration.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: liquidity.
  • Method note: the mean of adjusted close times volume over the trailing 20 sessions (fact_eod.adv_dollar_20d_f); daily cadence. Not to be confused with the undiscounted 20-day average volume (no dollar scaling), a distinct factor that sits in the “drift” diagnostics quadrant (non-stationary, high entropy). \[\text{ADV\$}^{20d}_{i,t} = \dfrac{1}{20}\displaystyle\sum_{\tau=t-19}^{t} P_{i,\tau} \times V_{i,\tau}\]

5.12 9. Short interest / Positioning

Heavily-shorted / high-days-to-cover names underperform — informed short sellers and short-sale-constraint overvaluation (divergence of opinion). Boehmer–Jones–Zhang 2008 [#101], Rapach–Ringgenberg–Zhou 2016 [#100], Miller 1977 [#111].

5.12.1 days_to_cover

  • Economic hypothesis / expected sign: behavioural, negative.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: positioning.
  • Method note: the vendor’s days-to-cover quantity, trusted verbatim (no platform recompute from average daily volume) (fact_short_interest.days_to_cover_f); a settlement disclosure lag of 18 calendar days applies to the input. \[\text{DTC}_{i,t} = \dfrac{\text{ShortInterest}_{i,t}}{\text{ADV}_{i,t}}\]

5.12.2 short_pct_float

  • Economic hypothesis / expected sign: behavioural, negative.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: positioning.
  • Method note: short interest over diluted shares outstanding as a float proxy (fact_short_interest.short_pct_float_f); the 18-calendar-day disclosure lag applies. \[\text{ShortPctFloat}_{i,t} = \dfrac{\text{ShortInterest}_{i,t}}{\text{DilutedSharesOut}_{i,t}}\]

5.13 10. Seasonality

Same-calendar-month historical returns recur — persistent cross-sectional seasonalities from mood, liquidity, and information cycles. Heston–Sadka 2008 [#37], Hirshleifer et al. 2020 [#40].

5.13.1 seasonality_same_month

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: momentum.
  • Method note: the mean, over the prior twenty years at most (five required), of the instrument’s monthly return in the same calendar month (fact_eod.seasonality_same_month_f, Heston–Sadka 2008); daily cadence. \[\text{Seas}_{i,t} = \dfrac{1}{K}\displaystyle\sum_{y=1}^{K} r^{month}_{i,\,\text{same-month}(t),\,y}, \quad 5 \le K \le 20\]

5.14 11. Earnings momentum / PEAD

Prices under-react to earnings surprises (standardised unexpected earnings), drifting for weeks after the announcement — the post-earnings-announcement drift. Bernard–Thomas 1989 [#75].

5.14.1 sue

  • Economic hypothesis / expected sign: behavioural, positive.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: growth.
  • Method note: earnings-surprise percentage divided by its own standard deviation over the strictly-prior eight announcements (fact_earnings_momentum.sue_f). Traded by the earnings-momentum strategy (Strategy Catalogue §5). \[\text{SUE}_{i,q} = \dfrac{\Delta EPS\%_{i,q}}{\text{STDDEV}_8\big(\Delta EPS\%_{i,\,q-1:q-8}\big)}\]

5.15 12. Time-series momentum / Trend

An asset’s own past 12-1 return predicts its next-month return across asset classes — a trend-following / slow-diffusion premium. Moskowitz–Ooi–Pedersen 2012 [#115], Hurst–Ooi–Pedersen 2017 [#95].

5.15.1 tsmom_12_1_vol_scaled

  • Economic hypothesis / expected sign: risk premium, positive.
  • Measured signal (gross, in-sample): not itemised — this factor is exempt from the cross-sectional research funnel, so it carries no cross-sectional information coefficient. Its only evidence comes from the strategy that trades it.
  • Cross-sectional permutation test: not itemised (funnel-exempt).
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental.
  • Style: momentum.
  • Method note: the same source value as the vol-scaled 12-1 momentum factor (§4), registered as its own factor because it has a distinct evaluation path — it is never scored by the cross-sectional information-coefficient, permutation, or tearsheet producers. Its trend-sign mechanic consumes \(\text{sign}(\text{TSMom}_{i,t})\), not the cross-sectional rank; the mechanic and its walk-forward backtest are documented in the Strategy Catalogue §3. \[\text{TSMom}_{i,t} = \text{Mom}^{vol}_{i,t}\]

5.16 13. Short-horizon reversal

Very-short-horizon losers bounce (and winners fade) — liquidity provision / over-reaction correction (contrarian). Lehmann 1990 [#96].

5.16.1 bb_pct

  • Economic hypothesis / expected sign: behavioural, negative.
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test: not itemised.
  • Naive & HAC t-statistics: not yet computed (see Method).
  • Lifecycle: experimental — the scale-free replacement for the deprecated lower-band factor below.
  • Style: volatility.
  • Method note: the position of the adjusted close within the 20-day Bollinger band (fact_eod.bb_pct_f); daily cadence; the diagnostics place it in the “healthy” quadrant. Because it is a ratio of price differences, a future split scales numerator and denominator identically and cancels — it is scale-free and point-in-time clean with respect to corporate actions, which is precisely the property its deprecated predecessor lacked. \[\%B_{i,t} = \dfrac{P_{i,t} - L^{20}_{i,t}}{U^{20}_{i,t} - L^{20}_{i,t}}\] where \(U^{20} = \text{SMA}^{20d} + 2\sigma^{20d}\) and \(L^{20} = \text{SMA}^{20d} - 2\sigma^{20d}\) (0 at the lower band, 1 at the upper band).

5.16.2 bb_lower_20 (superseded predecessor — not in the original 13-row table, retained for the correction it forces)

  • Economic hypothesis / expected sign: behavioural, positive — short-horizon mean-reversion (Lehmann 1990 contrarian reversal).
  • Measured signal (gross, in-sample): not itemised.
  • Cross-sectional permutation test (May 2026, since retracted): at the time, one of only four factors with a non-collapsed permutation null and a positive unbiased information ratio (in the range 0.21–1.35). This result is now known to be contaminated by a look-ahead defect in the raw price-level source, not evidence of a real edge. The other three factors in that same “real significance” group — the 50- and 200-day moving averages and the 20-day volume-weighted average price — were deprecated for the identical reason on the same date and are not part of this catalogue’s thirteen families.
  • Naive & HAC t-statistics: not yet computed, and moot — the factor is deprecated.
  • Lifecycle: deprecated for cause (20 July 2026) — a raw price-level source column carrying a look-ahead defect. Replaced by the scale-free percent-b factor above.
  • Style: volatility.
  • Method note: the 20-day simple moving average of adjusted close minus twice the 20-day standard deviation (fact_eod.bb_lower_20_f) — a raw price level, not scale-free, which is exactly the look-ahead defect that got it deprecated; daily cadence. \[L^{20}_{i,t} = \text{SMA}^{20d}_{i,t} - 2\,\sigma^{20d}_{i,t}\] We retain this entry as a cautionary record: the May 2026 synthesis called this factor’s permutation result the cleanest in the library, and two months later we found it to be a look-ahead artefact. A clean permutation p-value is necessary evidence, not sufficient.

5.17 Next steps

  1. Write the query computing the naive and HAC t-statistics per factor against the factor-return panel, and fill in every not-yet-computed field above.
  2. Resolve the two duplicate pairs — done 5 August 2026: market_beta_252d and momentum_12m_1m are the canonical survivors; the legacy aliases are retired. Remaining follow-up: if either legacy alias still has rows in the factor-status registry, flip their status there too (a database change, out of scope for this documentation pass).
  3. Backfill the hypothesis class and expected sign for the factors missing them in configuration (market_beta_252d, market_cap, amihud, adv_dollar_20d).
  4. Reconcile the Investment/Accruals mapping — the quarterly earnings-growth factor named there has no matching configuration file under that name.
  5. Re-estimate the empirical fields against the first deep run following the three corrections (corrected permutation null, forward-return winsorisation, composite-contribution rescaling).
  6. Consider formally adding residual_momentum_252d_eod (§4) as a fourteenth citation row, given its result is the strongest in the run.