Every backtest is fitted to something. The question is whether it was fitted to the market or to the sample, and the honest version of that question rarely gets asked before a strategy goes live. A backtest that loses money is easy to reject. A backtest with a smooth, rising equity curve gets trusted — and a smooth curve is exactly what curve-fitting produces on command, whether or not there's a real edge underneath it.
Two backtests that look identical and aren't
Run the same strategy through an in-sample optimization and you'll get a clean equity curve almost every time, because the optimizer's entire job is to find the parameter set that makes the historical curve look good. That's not evidence of an edge — it's evidence that an optimizer did what optimizers do. The only way to tell a real edge from a well-fitted one is to test the exact same fixed parameters on data the optimization never touched.
Notice what the chart doesn't show: a strategy that's genuinely broken. A fitted-to-noise strategy doesn't usually crash out-of-sample — it goes flat, choppy, indistinct from a coin flip. That's the tell. A real edge degrades a little outside its tuning window, because market conditions shift. A fitted one collapses toward zero, because the thing it was actually fitted to — the specific noise in the tuning window — doesn't exist anywhere else.
The math behind "if you test enough combinations, something wins"
This is the part that doesn't feel like it should be true until you run the numbers: testing more parameter combinations doesn't just raise your odds of finding a real edge, it also raises your odds of finding a fake one that passes a standard significance test purely by chance. At the common 5% significance threshold, roughly 1 in 20 random, meaningless parameter sets will look "statistically significant" on any given sample — not because they work, but because that's what a 5% false-positive rate means.
That's not a reason to stop optimizing — it's a reason to treat "I found a parameter set that beats the threshold" as a weaker claim the more combinations you searched to find it. A grid search over 500 combinations that surfaces one good-looking result is much less impressive than a strategy that was defined with a handful of parameters up front and happened to clear the bar on the first pass. The number of tickets you bought matters as much as whether you won.
Specific tests, not a gut check
"Does the equity curve look good" isn't a test — it's the exact question an optimizer is built to satisfy. The tests below are specifically designed to be hard for a fitted strategy to pass, which is the whole point.
What each test actually catches
| Test | What it catches | How you run it |
|---|---|---|
| Out-of-sample holdout | Parameters tuned to noise that doesn't exist outside the tuning window | Set aside the most recent 20-30% of history (or a different symbol) before tuning, and don't touch it until parameters are fixed |
| Walk-forward re-optimization | An edge that only worked during one historical regime | Refit on a rolling window, test the fixed result on the next unseen window, roll forward through the full history |
| Parameter sensitivity | A result that only exists at one exact setting, with no real logic behind it | Nudge each parameter up and down 10-20% and check the result doesn't collapse — a real edge is stable near its optimum, not a knife's edge |
| Transaction-cost stress test | An edge that only exists when fills are free and instant | Re-run with realistic spread, slippage, and commission, then again with those costs doubled |
The sensitivity test is the one most backtests skip, and it's the cheapest to run. If a strategy's returns fall apart when you move a single threshold by a few percent, the optimizer didn't find an edge — it found a coordinate that happened to line up with the noise in that specific sample. This is the same discipline behind writing a trading plan around decidable rules instead of vague ones: a parameter that only "works" at one exact value isn't a rule, it's a coincidence with a number attached.
Costs are not a footnote
A backtest that fills every order at the exact printed price is testing a strategy that doesn't exist. Spread, slippage, and commission take the biggest bite out of exactly the trades that made the backtest look good in the first place — the small, frequent ones with a thin edge per trade, which is what a lot of tuned day-trading strategies end up looking like once curve-fitting has done its work.
Signs a backtest is fitted to the sample, not the market
-
Performance holds on a genuine out-of-sample holdout — Fixed parameters, tested once, on data the optimizer never saw during tuning.
-
Returns survive a 10-20% nudge on every parameter — A real edge is stable near its optimum. A fitted one is a single sharp peak surrounded by nothing.
-
Realistic spread, slippage, and commission are modeled from the first run — Not added afterward as a discount — built in before the parameters are ever tuned.
-
Parameter count is small relative to trade count — A dozen free parameters fit to sixty trades has more degrees of freedom than data points to constrain them.
-
The equity curve looked good, so testing stopped there — The single most common failure mode, and the one that skips every test above.
Why a suspiciously smooth curve should worry you, not reassure you
There's a specific pattern worth naming directly: a backtest with almost no drawdown and a nearly straight 45-degree equity line. Real markets don't produce that shape from a real edge — every genuine strategy has losing stretches, because the conditions it depends on come and go. A curve that's too smooth is usually a curve that's been shaped by exactly the optimization process described above, whether or not the person running it meant to do that.
This is the same failure that shows up in risk-reward thinking when a headline win rate gets quoted without the expectancy math behind it — a single flattering number, taken at face value, standing in for the harder question underneath it. A backtest's equity curve is that same kind of number. It's a summary, and summaries are exactly what curve-fitting is good at producing on request.
What this changes about how you test
None of this means backtesting is worthless — it means a backtest earns trust the same way a trade idea does, by surviving a check it could have failed. Split your data before you look at any of it. Count your free parameters against your trade count. Run the sensitivity test. Model costs from the start. Treat a smooth curve as a flag to check harder, not a result to celebrate. A strategy that passes all four tests still isn't guaranteed to work live — no backtest can promise that — but it's no longer indistinguishable from a coin flip that got tuned until it looked smart.
Once a strategy clears those tests, the same discipline is worth carrying into how you read individual setups it flags. AI chart analysis works from the chart in front of it, with no memory of your backtest results and no stake in whether the strategy behind the trade "worked" historically — a useful, disinterested second look at whether today's setup actually matches the pattern your backtest was built around, not just a name that resembles it.