The Study Population

Every empirical claim in this book is a claim about a population. This appendix defines that population — what we call the universe — states its inclusion rules as sampling criteria, and sets out the limitations that bound what the sample can support.

What this contributes. A cross-sectional result is only as good as the cross-section it was measured on. Two of the failure modes that most often invalidate published factor research — survivorship inflation and look-ahead — are properties of population construction rather than of estimation, and a third, selection on liquidity, is smuggled in by inclusion rules that go unstated. This chapter states ours, with measured magnitudes where we have them, so that a reader can judge which of our results the sample supports and which it does not.

1. Four populations, not one

The single most common error available here is to measure on one population and claim for another. We use four, and they are not interchangeable.

The admitted population is every instrument that passes security-type admission: ordinary common equity, listed on a major US exchange, alive on the date in question. It is the widest population we consider and the parent of the other three.

The research cross-section is what factor evidence is measured over. It is the admitted population further restricted, at every date, by a price floor and a trailing dollar-volume floor (§2.2). This is the population behind every information coefficient, permutation test, and diagnostic in this book. Where a chapter reports factor evidence without naming a population, this is the one.

The deployment population is what a strategy may actually hold. It applies its own, stricter executability floors. It is deliberately a subset of the research cross-section: we permit ourselves to study names we would not trade, because excluding them from study would bias the estimate of the phenomenon, while including them in a portfolio would overstate what is collectable.

Index membership is exact historical constituent membership for the standard benchmarks, used only where a claim is specifically about an index. It has its own coverage limits (§5.5).

The research-and-deployment relationship is the one to keep in view: research is a superset of deployment by design. A factor premium that exists only in names below the deployment floors is a real phenomenon and an uncollectable one, and the book must say both.

2. Admission stated as sampling criteria

2.1 Security-type and venue

We admit ordinary common equity and exclude, by construction: funds and pooled vehicles, warrants, rights, preferred shares, and cross-listing duplicates of an instrument already admitted under its primary line. Each exclusion removes a security whose return series is not the object of study — a warrant’s return is a derivative of the underlying, not an independent observation of it, and admitting both would double-count the same firm.

Venue admission is US major exchanges only. This is a deliberate narrowing with a measured history behind it, recorded in §5.1, and it matters more than a venue rule usually would: our market-day calendar, our currency units, and our liquidity floors all assume a US listing.

Foreign-domiciled companies whose primary listing is on a US exchange are admitted — they file with the US regulator and trade on the same calendar in the same currency. Domicile and venue are independent axes, and conflating them is a mistake we made and corrected. A small number of domiciles are excluded on audit-access grounds.

2.2 Liquidity and price floors

At every date, an instrument enters the research cross-section only if, on that date, it satisfies:

  • a price floor on the adjusted close;
  • a trailing dollar-volume floor, computed as an average of daily close-times-volume over a twenty-market-day trailing window;
  • a data-coverage requirement on that window — at least fifteen of the twenty days must carry data, or the instrument is excluded for that date.

These are point-in-time: an instrument can enter and leave the cross-section as its price and turnover change, and its membership on a historical date reflects only what was knowable then.

Two properties of these floors matter for interpretation, and both are selection effects rather than mere thresholds. The coverage requirement selects against thinly-traded names, which are disproportionately small — so the cross-section is not simply “the market minus the cheapest names”, it is the market minus the names that trade rarely, and rare trading correlates with the very characteristics several factors are about. And the price floor is applied to a back-adjusted series (§5.3), which makes it bind harder in the deep past than its nominal value suggests.

A point-in-time market-capitalisation floor exists in the membership rule but is currently disabled, so no capitalisation criterion is applied. When it is enabled it will be fail-open: point-in-time share counts are available for roughly 78% of the cross-section, and a name without one is admitted rather than dropped, because silently deleting a fifth of the population to enforce a floor would be a worse error than not enforcing it.

3. Point-in-time construction and survivorship

Point-in-time. Membership is evaluated against instrument lifecycle dates, so a historical date sees only instruments that were listed and not yet delisted on that date. Combined with the point-in-time floors above, this means a cross-section reconstructed for a past date is the cross-section that was actually available then — the property that makes a backtested result admissible as evidence rather than as an artefact of hindsight.

Survivorship. Delisted instruments are retained and their price history spliced back, so the panel is not restricted to companies that still exist. This is the difference between a sample and a survivor list, and it is not a small correction: the survivorship haircut we apply to reported information ratios is measured, not assumed, and is of the order of a few hundredths of an information ratio at a quarterly horizon.

The residual survivorship risk is in the reasons for delisting rather than the fact of it. A benign delisting (acquisition at a premium) and a performance delisting are opposite events, and treating both as a uniform loss biases in the other direction.

4. Characteristics

The cross-section is not stationary in size, and this bounds early-period inference more than any single inclusion rule. The admitted population runs from roughly 1,550 instruments in the mid-1990s to roughly 6,900 by the mid-2020s. After the liquidity and price floors, the research cross-section on a recent date numbers around 1,500 names.

Two consequences follow directly. Early-period cross-sectional statistics rest on a materially smaller sample, so their standard errors are wider even where the point estimates look comparable — a period-by-period result that appears to strengthen over time may be reporting nothing but the growth of the denominator. And any statistic that is sensitive to cross-sectional breadth is not comparable across decades without saying so.

Roughly a quarter of admitted instruments report fundamentals in a currency other than the US dollar. Level-denominated quantities and mixed-currency ratios are therefore mis-scaled for those names unless a currency normalisation is applied, and that normalisation is currently disabled. Ratio and margin quantities are currency-neutral and unaffected either way. Any claim resting on a level or on a price-to-fundamental ratio must state which side of this it falls on.

5. Limitations

Stated with magnitudes where we have measured them. These are the reasons a result on this population might be wrong.

5.1 Historical foreign-venue contamination — the largest known defect

Until August 2026 the venue rule admitted eight non-US exchanges, and nothing downstream re-filtered them. The measured contamination was 2,468 foreign listings — 20.4% of the admitted population, 20.5% of the daily research cross-section, and 33% of all stored factor values.

The defect was not merely one of scope. Those instruments were scored on a US market-day calendar in non-US currency units — London quotes are denominated in pence, a hundred-fold level inflation that propagated into both the level-based factors and the dollar-volume floor that was supposed to exclude illiquid names.

Any factor result in this book computed before August 2026 rests on a contaminated cross-section. We have not re-run the full historical evidence base against the corrected population. Results from that period should be read as provisional pending recomputation, and a result whose sign or magnitude depends on the contaminated fifth is not established.

5.2 Delete semantics in the admission list

Admission is refreshed by upsert, which adds and overwrites but does not remove. An instrument that disappears from the vendor’s current snapshot — delisted, acquired, renamed — was therefore never re-evaluated and retained its last admission decision indefinitely. Measured immediately after the venue correction: 18 orphaned records, 15 of them still marked admitted, mostly long-dead names plus two pence-denominated foreign funds. A sweep now closes this inside the refresh. The historical exposure was bounded in practice because most such names carry no recent price data, but it was not zero.

5.3 The price floor is applied to a back-adjusted series

There is no unadjusted close in our price history: the stored close is split- and dividend-adjusted throughout. A nominal price floor is therefore applied to back-adjusted values, which are lower than the actually-traded price in the deep past. The floor bites harder historically than its stated value implies, and it does so in a way that is correlated with time and with dividend history rather than being a clean level cut.

5.4 Missing listing dates treated as always-eligible

Approximately 162 admitted instruments carry no listing date and are treated as eligible throughout. This is conservative for coverage and permissive for look-ahead: an instrument may appear in a cross-section before it actually listed. The count is small relative to the population, but the direction of the error is the unfavourable one and it is not corrected.

5.5 Index constituent coverage

Exact historical constituent membership is reliable from October 2008 for the S&P 500, January 1995 for the Nasdaq-100, and January 1994 for the Dow. Before those dates the change log is incomplete, and an index-membership claim cannot be made.

5.6 The calibration of the floors is not settled

The current price and dollar-volume floors were chosen conservatively rather than derived. A recalibration has been specified and measured — it would widen the research cross-section by roughly 29%, from about 1,470 names to about 1,900 — but it is not in force, and the evidence in this book is measured on the narrower population. A prior comparison of alternative population definitions was run and the incumbent retained on a negative-lift reading; that comparison did not carry a cost-adjusted column, and until it does, a genuinely tradeable illiquid edge and an uncollectable paper edge are observationally identical.

6. Suitability

This population supports cross-sectional claims about liquid US-listed common equity over the period for which the price panel is deep — factor premia, cross-sectional predictability, and their decay — provided the claim is labelled with its period and the pre-August-2026 caveat of §5.1 is honoured.

It does not support, without further work: claims about non-US markets, which are excluded by construction; claims about microcap and illiquid names, which the floors remove and which are precisely the habitat where a small-capital operator might hold a structural advantage; claims resting on level-denominated fundamentals for the quarter of the population reporting in another currency; index-membership claims before the coverage dates in §5.5; and any claim whose sign depends on the contaminated cross-section of §5.1.

It is silent on capitalisation-conditional effects, since no capitalisation criterion is currently applied — a size-conditioned result must construct and state its own partition rather than relying on the population definition to have made one.