Articles

Monte Carlo on a backtest: what it can and cannot tell you

What a Monte Carlo of a backtest answers, why it does not correct overfitting and what the report shows as resampled risk, with a computed example.

A Monte Carlo of a backtest reorders or resamples the history's results thousands of times and looks at how much the falls change. It shows the spread of risk that a single order of trades hides. It is not a second test of the strategy: it works with the same data, with its luck and with its biases. Here is what it answers, what it does not answer and what Rigor's report computes, with a synthetic example you can reproduce from its seed.

What resampling trades or a curve means

A backtest delivers a sequence: closed trades or per-period returns of the equity curve. Resampling builds new histories from those same pieces. Shuffling changes only the order and keeps every result. The bootstrap draws pieces with replacement, so a history can repeat some and leave out others. The stationary bootstrap draws blocks of consecutive periods, of random length, to keep part of the dependence between nearby days.

Each method repeats the draw thousands of times and summarises what comes out: the median maximum drawdown, the one reached in 1 history in 20, or the share of histories that fall further than a threshold. The name Monte Carlo refers to that use of repeated draws. It adds no information the history lacks; it organises what is already there.

What it answers: the spread if the order is exchangeable

A backtest's maximum drawdown depends on the order in which the losses arrived. With the same results, another order would have given a milder or a deeper fall, and a shorter or longer losing streak. A Monte Carlo answers that specific question: what range of falls and streaks these results produce if the order is exchangeable, that is, if any order was equally likely.

That assumption is the limit of the answer. If losses depend on the market regime, on positions open at the same time or on changes in size, the order is not exchangeable and shuffling hides it. The block bootstrap keeps the dependence inside each block, but it cannot rebuild regimes the history does not contain.

A computed example with its seed

Declared · A synthetic series of 500 daily returns, drawn with NumPy from seed 20261008, with a mean of 0.05 % and a standard deviation of 1 % a day. In the order drawn, its maximum drawdown is 12.2 %. With shuffled_drawdown, 1,000 random orders of the same returns from seed 20260926, the worst fall runs from 7.5 % to 16.2 % in 9 orders out of 10, with a median of 10.6 %.

Declared · With drawdown_risk, the report's stationary bootstrap (2,000 histories of 252 days, an expected block of 5 days and seed 12345), the one-year maximum drawdown has a median of 8.5 % and reaches 15.5 % in 1 history in 20. In all, 32.9 % of the histories fall at least 10 % and 0.8 % at least 20 %.

The two figures measure different things: the first covers the series' 500 days; the second, one year. Neither is a measurement of a strategy or a forecast. The article's code fixes the seed and the parameters, so the same calculation always gives these figures.

What it does not answer: overfitting and the search

A Monte Carlo cannot tell a result with an edge from one picked by luck among many. It resamples what is there: if what is there is the best of a hundred attempts, it resamples that luck. Nor does it say how the strategy would behave on data it never saw, or deduct costs the backtest left out.

The search for configurations has another tool: deflated Sharpe compares the observed Sharpe with the one chance would give across the attempts made. Fitting to the past calls for an out-of-sample stretch fixed before looking at the results. The linked articles on deflated Sharpe and on what to do after a backtest explain both; the linked luck calculator does the first calculation with declared figures.

Why the Monte Carlo of an optimised backtest inherits its bias

The optimiser picks the configuration whose history came out best. Because of that selection, the history has more favourable streaks and milder falls than the process that produced it. Resampling starts from those same returns, so its histories inherit the inflated mean and the smoothed falls. A narrow, rising fan of curves can be nothing more than the trace of the choice.

Declared · We drew 100 synthetic series of 250 daily returns with zero mean and a standard deviation of 1 % (seed 20261009) and kept the one with the highest final result. With drawdown_risk and the same parameters, its one-year maximum drawdown has a median of 7.4 %, and 15.6 % of the histories fall at least 10 %. Another 5,000 days from the same generator (seed 20261010) give a median of 14.8 %, with 83.4 % of histories falling at least 10 %. No series had an edge: the difference comes from the selection.

What Rigor's report computes

The report calls this resampled one-year risk. The drawdown_risk function builds thousands of one-year histories from blocks of your file's returns, with a stationary bootstrap whose expected block starts at 5 periods and grows when returns cluster. It shows the median maximum drawdown, the one reached 1 time in 20 and 1 time in 100, the probability of falling at least 10, 20, 30 or 50 % and the consecutive periods below the peak. The seed is fixed and the report prints it.

Next to it, shuffled_drawdown compares the file's worst fall with that of the same returns in random orders, which keep the Sharpe, the volatility and the final result. If the file's fall is milder than in nearly every order, losses followed losses less often than chance would give, as in a smoothed curve; if it is deeper, they came in streaks. With closed trades, loss_streak_review makes the same comparison for the longest losing streak, and the challenge simulator walks histories resampled the same way through each challenge's rules.

One resampling does count towards the class: the 5th percentile of the Sharpe in a stationary bootstrap is part of the test against chance. The resampled risk, the random orders and the streaks are informational and do not change the class. All of them describe the supplied history; none is a prediction.

See it in the sample report

Open the linked sample report to see resampled risk on synthetic data, then compare it with the one from your own file.

FAQ

Does a favourable Monte Carlo rule out overfitting?

No. It resamples the history that already exists; if that history came from picking the best of many configurations, resampling keeps the bias. Deflated Sharpe addresses the search, and an out-of-sample stretch fixed before looking addresses fitting to the past.

How many simulations are needed?

More simulations sharpen the percentiles but do not change what is resampled. The report uses thousands of histories with a fixed seed so the same file gives the same numbers. The main limit is the length and representativeness of the history, not the number of draws.

Related

Put it to the test with your file

Upload the file your platform already exports, without converting it. With an account, your first full report is free, with the PDF.