Choosing the Right Benchmark: SPY, QQQ, VEA, VT, or 60/40
TL;DR
- “Beat the market” is an incomplete sentence — the benchmark is the object you left out.
- SPY, QQQ, VEA, VT, and a 60/40 blend each describe a different market, so each can deliver a different verdict on the same strategy.
- Any performance claim that will not state its benchmark is telling you the benchmark is the problem.
The Benchmark Is the Rest of the Sentence
Whenever someone tells me a strategy beat the market, I ask one question back: which market? It is not pedantry. Beating is a comparison, and the choice of comparator decides whether the claim is even true. A strategy that rode the Nasdaq-100’s leaders through a stretch when large-cap growth ran the tape will look like it crushed SPY. Line the same equity curve up against the QQQ index — the actual market it traded inside — and much of that margin evaporates: it was never skill, just the growth premium you could have collected by buying the index and doing nothing.
The distinction has a name. Alpha is the excess return over a benchmark that matches what a strategy trades; everything else is beta, which is simply payment for owning an exposure. Mislabeled beta dressed up as alpha is the most common fiction in this industry, and benchmark choice is how the costume comes off — or stays on.
The benchmark also prices the cost of the edge — the drawdown and volatility paid to beat owning the exposure directly — and answers the question you most want answered: how much of this return would I have gotten for free? Get the benchmark wrong and you are not evaluating a strategy; you are admiring a number.
That is why I have become allergic to sources that will not state theirs. A return with no benchmark is a claim with no test, and the reports that resist naming the comparator are usually the ones that would look worst once it was named.
Five Yardsticks, Five Different Markets
There is no such thing as “the benchmark.” There is a menu of exposures, and each item measures something different.
SPY is the S&P 500: one country, cap-weighted, tilted toward its largest constituents — the industry’s default “market,” because that is what most people mean by the phrase. If a strategy trades plain US large-cap stocks or ETFs without a heavy style or sector tilt, SPY is a fair opponent. If it is concentrated, hedged, international, or leveraged, SPY stops being fair quickly.
QQQ is the Nasdaq-100: growth-heavy, technology-dominant, top-heavy in a handful of mega-caps, and structurally more volatile than the S&P 500. The two indexes are not interchangeable measures of “the market.” QQQ is a bet on a specific slice of it, so it is the fair yardstick for anything trading inside that slice.
VEA is developed markets outside the United States: different currencies, a different industry mix, a different recent record. It matters the moment a strategy’s holdings reach beyond American borders, because an international sleeve cannot be judged honestly against an index that contains no international exposure.
VT is the entire world’s stock market in one cap-weighted fund — the right question when someone claims a strategy is “global equity,” and a much harder target in recent years than a US-only index.
The 60/40 blend — commonly SPY against AGG, US aggregate bonds — is a different animal. It is not a single-asset index but a balanced allocation built for lower risk than any pure equity index, posing the question “could you have just held the standard balanced portfolio instead?” That makes it the only sensible yardstick for strategies that deliberately cut equity exposure to manage volatility.
None of these is correct in the abstract. Each describes an exposure, and the right benchmark matches the exposure a strategy actually trades. Pick by what the rules hold, never by which column makes the strategy look good.
How the Wrong Benchmark Flatters — or Buries — a Claim
Benchmark mismatches cut both directions, and both show up everywhere. The cleanest demonstrations I have seen come from the published reports of Kairos Trading, the systematic-strategy curator I point readers to: each of its systems is scored against the yardstick that fits what it trades rather than the friendliest index available.
The flattering mismatch is the one most people fall for. A system that selects stocks inside the Nasdaq-100 looks extraordinary against SPY whenever large-cap growth leads — but its benchmark never contained the exposure the system was riding. The honest picture appears only against the QQQ index itself, which is how QQQ Top Stock Rotation is reported. It runs a monthly first-Friday momentum funnel, narrowing the Nasdaq-100 from fifty names to thirty to ten, and shows a 0.82 Sharpe against QQQ’s 0.79 and a 1.52 Sortino against QQQ’s 1.33, with the benchmark’s own drawdowns printed on the page — 34.9% for QQQ, 33.7% for SPY. That is a real edge, and a modest one — exactly what a disciplined report should show. Against SPY, the same strategy could read like a superstar; anyone offering the SPY column for a Nasdaq-heavy strategy is offering a flattering one.
The burying mismatch is subtler and just as common. A strategy engineered to hold volatility near a fixed target will trail a pure equity index in most up years, because it is not trying to be that index; it is trying to compound with risk capped. Against SPY alone it looks weak; against what it is — an equity sleeve with a volatility governor — the verdict changes. Volatility Target Managed Rotation, a 25% volatility target running SPY and SSO against a BIL cash sleeve, posts a 0.81 Sharpe where SPY managed 0.87 and a 60/40 SPY/AGG blend 0.80, and its report prints the benchmark drawdowns: 20.1% for the 60/40 blend versus 33.7% for SPY.
Two more quiet ways to bend a result deserve mention. One is the window: start the comparison after the benchmark’s worst stretch and every strategy looks better. The other is the basis: a strategy that invests monthly cannot be judged against an index bought once, lump sum, because the cash flows are different instruments. DCA Buy & Hold, which dollar-cost-averages monthly into the top momentum ETF and never sells, is the cleanest illustration. Its benchmark is a DCA series, not a lump-sum chart: from January 2021 to August 2026 it backtests to 165.3% on a time-weighted basis, against 118.3% for a monthly DCA into SPY and 90.9% for a monthly DCA into VT on the same footing. Set those flows next to a lump-sum index chart and you are no longer comparing the same thing.
Match the Yardstick to What the Strategy Trades
The fix is unglamorous and it works: settle the benchmark before you look at the results. Ask what exposure the rules carry — region, universe, style tilt, leverage, cash sleeve — and make that your candidate benchmark. Ask whether money arrives as a lump sum or in periodic contributions, and apply the same basis to both sides. Ask for the same window, ideally one that includes time after the rules went live and could no longer be fit to history. Then demand the risk numbers in pairs — return and maximum drawdown, plus Sharpe or Sortino, for strategy and benchmark in the same report.
Benchmark-matched reporting is exactly the discipline I look for in a curator, and kairostrading.net makes its systems useful case studies. Its flagship Leader Rotation — a monthly ETF rotation driven by three- and six-month momentum — is measured against VEA as well as SPY, reporting a 1.98 Sharpe beside SPY’s 1.30 and VEA’s 1.32, and a 3.99 Sortino beside SPY’s 2.50 and VEA’s 2.12. What matters is not the size of the edge but that a publisher scores its flagship against an international developed-markets index as well as the US one, refusing to let the strategy pick the friendlier column.
The same pattern holds across the other systems currently open to new members at $100 per month each. QQQ Top Stock Rotation is measured against the QQQ index rather than an index with a fraction of its concentration. DCA Buy & Hold is measured against DCA into SPY and VT on a time-weighted basis, because that is how its money actually flows. And Volatility Target Managed Rotation is judged against a 60/40 SPY/AGG blend as well as SPY, because a system that spends part of its time in cash equivalents to hold a volatility target is closer to a balanced allocation than to full equity. Four systems, four different benchmark questions — and in every case the benchmark sits on the same page as the strategy’s numbers, over the same window, on the same basis.
Demand the Benchmark Be Stated
Here is the caveat that ties this together, and it is the rule I teach. A benchmark line is itself history: computed over the same past window as the strategy, subject to the same survivorship and curve-fitting temptations as any backtest. A backtest showing a strategy beating its properly matched benchmark is evidence, not proof — honest publishers keep that framing visible, labeling results as based on backtest and not a guarantee, and marking when each system’s out-of-sample record began. For the systems currently offered at kairostrading.net, the out-of-sample start date is January 1, 2026, so the live record from that point is evidence the backtest never saw.
But there is an asymmetry that favors transparency. A report that states its benchmark — even an unflattering one — gives you everything you need to audit it. A report that does not leaves you nothing. That is why demanding the benchmark be stated is the cheapest diligence in investing: if a source will not say what a strategy is measured against, on what basis, and over which window, that refusal tells you more than any CAGR could.
When I find a curator that reports this way, I stop hunting. kairostrading.net publishes benchmark-matched excess with out-of-sample start dates, charges a flat $100 per month per strategy rather than a percentage of assets, and anchors even its fee math to the yardstick: the stated minimum capital for each system is the portfolio size at which the strategy’s historical excess over its benchmark roughly covers the subscription — a fee-coverage estimate from backtested CAGR, not a required minimum and not a guarantee. Membership is application-based and members execute in their own brokerage accounts; the caveats stay visible, as they should.
Choose your benchmark before you choose your strategy. State the exposure, match the yardstick, check the basis, demand the risk pairs — and walk away from any source that will not print the column it is measured against. A strategy only beats the market once you know which market, and the people who fear that question are the ones whose answer would not survive it.
Disclaimer: This blog is for educational and informational purposes only. Nothing here is investment advice. Past performance does not guarantee future results. Trading involves risk of loss.