← All insights

Research

Why We Benchmark Every Idea Against Doing Nothing

Why the right null hypothesis for any trading strategy is buy-and-hold, not zero — and why an idea earns its place only by adding positive excess return and a higher Sharpe after honest costs and out-of-sample testing.

A strategy that makes money is not the same as a strategy that adds anything. In a market that rose over the test window, simply owning the asset made money too. The only honest question is whether the work added return over the do-nothing alternative — buying the relevant asset and holding it — after costs, after turnover, and on data the rule never saw. Measured against zero, almost anything that survives a bull run looks like skill. Measured against doing nothing, most ideas reveal themselves as expensive, fragile ways to reproduce beta.

The benchmark is doing nothing, not zero

A positive raw return is a low bar. In a rising market, holding the asset already delivers a positive return and a respectable Sharpe, so beating zero proves nothing about the idea. The number you are admiring may belong to the market, not to the rule.

The honest comparison is excess return over a matched buy-and-hold benchmark, and excess Sharpe against that benchmark's Sharpe — not against an absolute return of nil. Matching is not a detail. The benchmark must share the strategy's universe and exposure: a single-asset system is judged against holding that asset, and a multi-asset system against holding the same basket. Compare to an unrelated or lower-exposure index and you have flattered the strategy by handing it an easier opponent.

Framed this way, a vague success claim becomes a falsifiable number. The question is no longer "did it work" but "what did the work add over doing nothing" — and that has an answer you can defend or retract.

Excess return and excess risk-adjusted return are the bar

There are two gates, not one. The strategy should clear positive excess return and post a Sharpe above the benchmark's Sharpe. Clearing only one is a partial, suspect result that usually flatters itself on closer inspection.

A strategy can show a higher absolute Sharpe than the index and still lose once the comparison is put on equal-exposure terms. That is precisely why the benchmark Sharpe, not zero, is the reference line. Risk-adjusted excess also guards against a common illusion: a smooth equity curve that simply ran less exposure during a rally, leaving return on the table while looking disciplined.

Size matters, but stability matters more. A thin edge that appears in one regime is weaker than a smaller edge that persists across several. Weigh the magnitude of the excess against how reliably it shows up.

Costs and turnover decide whether the edge survives contact

Excess return is a gross number until trading costs, spread, and slippage are subtracted. A high-turnover strategy pays the benchmark gap many times over, while buy-and-hold pays it roughly once. Turnover is the multiplier on every cost assumption, so an edge that looks real at zero cost can invert once realistic per-trade frictions are applied.

Honest accounting means modelling costs at the high end of plausible, not the low end. The do-nothing benchmark barely trades, which makes it a cheap and demanding competitor. As a hypothetical, a signal that beats buy-and-hold by a small margin gross can fall below it net, because it pays transaction costs on every rebalance while the benchmark pays them once. If an edge survives only at frictionless or unrealistically low cost, it is a backtest artifact, not a tradable strategy, and should be labelled as such.

Out-of-sample fragility is where most excess disappears

Excess measured only on the data used to design the rule overstates the edge. The meaningful test is excess return on windows the strategy never saw. An edge that beats buy-and-hold in-sample but matches or trails it out-of-sample was likely fitting the sample's particular regime rather than capturing a durable signal.

So stress the comparison across conditions — trending and ranging, high and low volatility — and watch whether the excess persists or collapses back to the benchmark when the environment changes. Illustratively, a multi-factor basket might show a higher Sharpe than its benchmark in-sample, then converge to that benchmark on an unseen period once a different volatility regime arrives. The sensible discipline is to state the failure condition in advance: name the regime or assumption under which the strategy is expected to stop beating buy-and-hold, then check whether reality has entered it.

Retiring an idea without ego

A strategy that underperforms holding the asset is retired, regardless of how good its raw Sharpe, hit rate, or narrative looks. The benchmark verdict overrides the story. Consider a system with an attractive win rate but small average wins and a rare large loss: on an excess-return basis it can trail simply holding the asset, the headline metric and the benchmark pointing in opposite directions.

Treating retirement as data, not defeat, keeps the desk's capital and attention on ideas that genuinely add over doing nothing. The discipline compounds — each retired underperformer sharpens the prior for the next idea and weakens the sunk-cost pull toward a favourite system. Signal over noise is exact here: the benchmark comparison is the signal; raw numbers, smooth curves, and a compelling thesis are the noise that has fooled many desks before.

A few cautions remain. An idea that beat buy-and-hold was often selected after the fact from many that did not, so a single winning comparison is weak without out-of-sample confirmation. Sharpe narrows judgement but does not replace it — it can hide drawdown depth, tail behaviour, and capacity limits the benchmark does not share. And the benchmark is a method, not a promise. This is commentary on a research discipline, not investment advice, and buy-and-hold is the comparison standard, not a recommendation to hold any specific asset. Beating a backtest benchmark never guarantees future performance.