Two questions run through this week’s research, and neither is about what markets will do. The first is how much of an impressive result survives the search that found it. The second is what an allocation actually adds to the portfolio someone already holds. Both point the same way for a long-term investor: the interesting number is rarely the headline one.
A backtest can be graded on how it was produced, not only on what it produced. Equity Strategy Backtesting: Luck or Edge? The MinervaScore as a Statistical Robustness Grade starts from a plain observation: trading rules are usually selected after many parameter trials, so a strong historical result can reflect search luck rather than a persistent signal. The standard summaries — return, Sharpe ratio, drawdown — say nothing about how many candidates were tried, whether the chosen rule survived out-of-sample validation, or whether the available history was long enough to support the conclusion at all. The paper proposes a robustness grade that records those things explicitly. The method matters more than the score: it treats the search process as part of the evidence, which is the opposite of how a performance chart is usually read.
The same bias has a name in econometrics, and it is measurable. An NBER working paper on Instrument Hacking studies what happens when researchers evaluate several candidate instruments and report the one with the most favorable statistics. The authors show this selection induces median bias in the resulting estimates, and that the bias grows as more candidates become available. It is the identical mechanism as the backtest problem, in a different discipline: when a result is chosen for looking best among many, its apparent strength is partly an artifact of the choosing. Reading any single reported number without knowing how many were discarded gives an incomplete picture.
A model that looks strong in the sample can decay outside it. Klement on Investing reviews work on the performance decay of LLM trading strategies, following an earlier experiment in which a language model asked to forecast inflation performed poorly once the forecast period fell outside its training window. The concern described is leakage: information about the test period reaches the model through its training data, so backtested results flatter the approach relative to how it behaves on genuinely unseen data. The general lesson is older than the technology. A result is only as trustworthy as the separation between the data that built it and the data that tested it.
Judge an allocation by what the portfolio already carries. A research summary asking whether private equity belongs in pension plans makes a structural point that generalizes well past that asset class. Internal rates of return, cash multiples, and public-market equivalents describe an investment’s own record. None of them answers the question that decides whether it helped: did it improve the portfolio after accounting for the risks that portfolio already carried? A high return may reflect skill, compensation for taking more risk, or access others did not have — and those are different things with different implications for the whole. This is risk-budgeting’s central move, which this blog has described in what a bank treasury knows about risk: size a position by the risk it contributes, not by the story it tells alone.
A holding can carry an exposure its label never mentions. Rate Risk and Rate Insurance, an NBER working paper, decomposes equity returns into a duration-matched government-bond component and a payoff component. Historically, risk and return rose far less with duration for stocks than for their matched government bonds, which the author attributes to what he calls rate insurance: rates fell in bad times, so the bond embedded inside a stock offset part of the equity payoff risk — and the same mechanism worked in reverse in good times, dampening expected returns. Whatever one makes of the decomposition, it is a reminder that exposures do not respect labels. A portfolio’s real risk profile lives in how its holdings move together, which is the same lesson as when a hedge isn’t a hedge.
The throughline. Both halves of this week reward the same discipline: look past the number to the process that produced it, and past the holding to the portfolio around it. A conversation with Kari Vatanen on the Fabozzi series frames this as total portfolio thinking — evaluating decisions against the whole rather than as a collection of independently attractive pieces. That framing is not new, and it is not a technique for finding better investments. It is a way of asking better questions about the ones already on the table: how was this result arrived at, and what does it change about the risk already being carried?
Curious where your portfolio’s risk structure stands? The free ETF Portfolio IQ Score is one way to see.