Note 01 · 25 August 2026

The breadth illusion in private-market data.

Private-market datasets promise the one thing public factor portfolios cannot deliver: genuinely independent bets. Most of that promise is destroyed before the modelling starts, by an unglamorous data problem that inflates the reported information ratio and leaves no trace in the backtest.

The claim

The attraction of the private record is independence. A succession event in one region is uninformative about a recapitalisation in another, which is exactly the condition the fundamental law assumes and factor portfolios violate. Two hundred public positions driven by one factor is breadth of one; two hundred unrelated private events is breadth of two hundred.

That is the argument. It is correct in principle and routinely wrong in practice, for a reason that has nothing to do with finance.

Where it breaks

The same operating entity appears in registers under several legal names, several identifiers, and several spellings. Holding structures interpose. Officers overlap. A single economic actor can present as four or five apparently distinct records, and unless they are resolved, each is counted as an independent observation.

The consequence is not a rounding error. Under the fundamental law, realised information ratio scales with the square root of breadth, so inflating the count of independent bets by a factor of k inflates the apparent information ratio by √k.

IRapparent = IRtrue ·k ,  k = Nobserved / Ndistinct
k
duplication factor — observed records per distinct economic actor
N
counts before and after entity resolution
A duplication factor of 1.5, which is modest for an unresolved private-market universe, overstates the information ratio by roughly 22 percent. At k = 2 the overstatement is 41 percent. Neither shows up as an anomaly in the equity curve.

Why it is invisible

Every conventional diagnostic passes. The backtest is clean because the duplicated records are real records — the events happened, the dates are right, the outcomes are correct. Nothing is fabricated. The error is in the denominator of a statistic nobody recomputes after the join.

It also survives out-of-sample testing, because the duplication is a property of the data source rather than of the period. A holdout drawn from the same unresolved universe carries the same inflation, so the out-of-sample result confirms the in-sample one and both are wrong by the same factor.

Bounding it

The overstatement is estimable before any strategy is built. Resolve on deterministic identifiers first, then block on name, address and officer overlap, and report k as a property of the dataset rather than as a preprocessing detail. Where resolution is uncertain, the conservative treatment is to merge — over-merging costs breadth and understates the result, while under-merging manufactures it.

We report k alongside any breadth-dependent estimate for the same reason we report a half-life alongside an effect: it is the term most likely to be flattering by accident.

What would falsify this

Research commentary on measurement. It is not an offer, a solicitation, or investment advice, and it recommends no security, strategy or transaction.