A small number of relationships govern whether an edge exists, whether it survives contact with the market, and whether the evidence for it means anything. Most programmes fail on one of these rather than on the modelling.
The fundamental law of active management states it exactly: realised information ratio scales with forecasting skill and with the root of the number of genuinely independent decisions, degraded by how much of the forecast survives implementation.
Predictive content dissipates as information diffuses. We model it as exponential decay in the information coefficient and report a half-life with every effect, because what a signal is worth is the integral of its edge over the time it survives, not the height of its backtest.
Rates measured on small samples are reported as intervals. We use the Wilson score lower bound, which stays well behaved near zero where the normal approximation does not — and near zero is exactly where a new signal lives.
Families of tests are declared before they are run and controlled by false-discovery rate rather than reported one at a time. On a corpus of this size an uncorrected search will always find something.
Position size follows from the estimate and its uncertainty, not from conviction. We size on a fraction of Kelly, because full Kelly assumes the estimated edge is the true edge — and the whole point of the previous section is that it is not.
The constraint is rarely the model. It is breadth, decay, and honest error bars.
Universes are rebuilt from the record as it stood on the decision date, including entities that later ceased to exist. Assembling a universe from survivors embeds an upward bias no modelling removes.
Identifiers are reconciled across registers before any estimate is computed. One entity appearing under several identifiers inflates apparent breadth and corrupts the root-breadth term directly.
Fields are checked for availability at the decision date, not at the date they were populated. This is the commonest and least visible cause of a beautiful backtest.
Every effect carries its capacity assumption: turnover, participation, and the liquidity of the names it selects. An edge that disappears at size is an artefact of the universe filter.
Every figure resolves to a document, an issuer and a date. Derived fields carry their derivation. Where the record cannot answer, that is the answer given rather than the nearest available number.
The transferable problem is recovering structure from a channel that is mostly noise, and stating precisely how much confidence the data supports.
V. Menon, J. P. Bonilla Ataides, R. Mehta, A. Gu, D. B. Tan, M. D. Lukin
R. Sajith, Z. Song, B. Roberts, V. Menon, Y. Li
V. Menon, N. E. Sherman, M. Dupont, A. O. Scheie, D. A. Tennant, J. E. Moore
V. Menon, A. Gu, R. Movassagh
A. Barbosa, Y. Baum, P. Mundada, G. Hartnett, V. Menon
A. J. Menssen, M. Gullans, T. Manovitz, J. E. Taylor, V. Menon, et al.