An expert advisor is overfitted when its parameters were tuned so tightly to past data that the backtest result no longer says anything about the future. The clearest sign is a high Sharpe that fails three tests: discount it by the number of configurations the optimiser tried, double the costs, and run the robot on a stretch of data nobody touched. Here is how to run each test with the tester report and the optimisation XML you already have.
What the optimiser does
The MT4 and MT5 optimiser walks through thousands of parameter combinations and stores the result of each one. A range of 1 to 50 on a moving-average period and 1 to 100 on the stop already gives five thousand combinations. The optimiser knows nothing about the market: it only measures which combination would have done best on that exact history.
That history contains noise: moves that will not repeat. Try enough combinations and one of them fits the noise by chance and lands at the top of the table. The vendor publishes that row and you only see that one.
Why the best of many looks good by luck
You can put a number on that effect. With three years of daily returns and no real edge, the best of 10 configurations shows a Sharpe near 0.9. The best of 100 reaches 1.5, the best of 200 reaches 1.6 and the best of 1,000 reaches 1.9. None of that is skill: it is the expected result of picking the maximum.
A short history makes everything worse. With one year of data, the best of 200 configurations sits near a Sharpe of 2.8; with two years, the best of 500 sits near 2.2. With five years, the best of 50 drops to 1.0.
Isolated peak or plateau
Open the optimisation XML and sort the passes by result. Take the chosen configuration and look at its neighbours: the same settings with one parameter one step up or down, everything else unchanged. If the stop goes from 40 to 45 and the result halves, you are standing on an isolated peak.
A strategy that captures something real sits on a plateau: the neighbours keep most of the result. Count how many neighbours stay positive and what share of the result their median keeps; that figure tells you more than the equity curve.
Costs at 2x
Many robots live on small, frequent trades, where cost weighs more than the signal. Run the backtest again with twice the per-trade cost the vendor declares. If the result turns negative, the EA had no margin, only optimistic costs.
Do the same test at three times the cost and note at which cost the result reaches zero. That break-even cost is the number to compare with your broker's real spread during the hours the robot trades, not with the daily average.
A holdout you never touched
Before optimising, set aside the last stretch of history and do not open it until the end. The optimiser must not see it, and neither should you. Once you have the final configuration, run it once on that stretch and compare the Sharpe with the one from the optimisation stretch.
If you go back to the optimiser after seeing the result and adjust something, the stretch is contaminated and counts as one more pass. An out-of-sample stretch works once. That is why it matters that the vendor states which dates were used to optimise and which were not.
What to ask the vendor
Ask how many passes the optimiser ran in total, not only how many versions were published. Ask for the complete optimisation XML, not a screenshot of the best row. Ask which dates were left out of the optimisation and what per-trade cost the backtest used.
Also ask for the Strategy Tester report with the parameters printed on it, to check that they match the pass being sold to you. Without those four answers you have no way to measure what you are buying.
How to measure it on the real optimisation XML
The MT5 optimisation XML stores every pass with its parameters and its result, so the number of configurations tried becomes a measured figure, not a declaration. With that number and the backtest's years, Rigor's calculator tells you what Sharpe luck alone would give and how much is left after the discount. It is free and asks for no signup.
If you upload the report and the XML for an audit, Rigor counts the configurations tried, computes the deflated Sharpe, reruns the costs at 1x, 2x and 3x with the break-even cost, and checks the out-of-sample stretch you declare. Every number comes tagged Measured, Declared or Not measured, so you know what comes from the file and what comes from the vendor.
FAQ
How many years of data do I need to trust a Sharpe of 1.8?
It depends on how many configurations were tried. With 1,000 tries, about 3.3 years of daily returns are needed for luck alone to fall below 1.8. With less history, that Sharpe fits inside what chance explains.
Does an MT5 forward test replace the holdout you never touched?
Only if you ran it once and did not optimise again after seeing it. If the vendor repeated the cycle several times until the forward looked good, the forward is one more configuration, chosen for its result.
Does genetic optimisation change anything?
Yes: the genetic algorithm does not try every neighbour of the chosen configuration, so the XML may lack the passes that show the plateau. When they are missing, run a complete optimisation over a narrow range around the final configuration.