Private-market datasets promise the one thing public factor portfolios cannot deliver: genuinely independent bets. Most of that promise is destroyed before the modelling starts, by an unglamorous data problem that inflates the reported information ratio and leaves no trace in the backtest.
The attraction of the private record is independence. A succession event in one region is uninformative about a recapitalisation in another, which is exactly the condition the fundamental law assumes and factor portfolios violate. Two hundred public positions driven by one factor is breadth of one; two hundred unrelated private events is breadth of two hundred.
That is the argument. It is correct in principle and routinely wrong in practice, for a reason that has nothing to do with finance.
The same operating entity appears in registers under several legal names, several identifiers, and several spellings. Holding structures interpose. Officers overlap. A single economic actor can present as four or five apparently distinct records, and unless they are resolved, each is counted as an independent observation.
The consequence is not a rounding error. Under the fundamental law, realised information ratio scales with the square root of breadth, so inflating the count of independent bets by a factor of k inflates the apparent information ratio by √k.
Every conventional diagnostic passes. The backtest is clean because the duplicated records are real records — the events happened, the dates are right, the outcomes are correct. Nothing is fabricated. The error is in the denominator of a statistic nobody recomputes after the join.
It also survives out-of-sample testing, because the duplication is a property of the data source rather than of the period. A holdout drawn from the same unresolved universe carries the same inflation, so the out-of-sample result confirms the in-sample one and both are wrong by the same factor.
The overstatement is estimable before any strategy is built. Resolve on deterministic identifiers first, then block on name, address and officer overlap, and report k as a property of the dataset rather than as a preprocessing detail. Where resolution is uncertain, the conservative treatment is to merge — over-merging costs breadth and understates the result, while under-merging manufactures it.
We report k alongside any breadth-dependent estimate for the same reason we report a half-life alongside an effect: it is the term most likely to be flattering by accident.
Research commentary on measurement. It is not an offer, a solicitation, or investment advice, and it recommends no security, strategy or transaction.