Articles

I built a trading bot with AI in 30 minutes: how to tell if the backtest is luck

A quickly generated bot still accumulates trials. Count its variants, compare Sharpe with the calculator and inspect the effect of trading costs.

Declared · The thirty minutes in the title describe an illustrative situation, not development timed by Rigor. A viral thread shows a prompt, generated code and an upward curve. What often disappears is the path between those images: discarded instructions, added filters and changes of market after inspecting a result. Producing code quickly does not produce fresh evidence. Before interpreting the curve, reconstruct how many opportunities the process had to discover a favorable coincidence in the same history. The coding tool cannot recover that missing record for you.

Count the complete search

Whenever the agent tests a rule and receives a result, that information can guide its next change. Switching the indicator, exit, session or market belongs to the search even if the final file keeps its original name. Versions omitted from the thread count too. Chat messages are not a configuration count: an instruction can launch many combinations, while a syntax correction may leave every trading rule unchanged.

Keep optimizer logs, saved versions and the selection criterion. If part of the search is missing, declare that gap. A partial count should not appear as a measured total simply because it is the only count available.

Compare against a search without skill

The public calculator takes annual Sharpe, history length and configuration count. It compares the declared result with the expected Sharpe of the best unskilled variant under its assumptions. This reference does not describe a particular account. It is not the probability that the bot will work later. It summarizes how selecting among alternatives can lift the largest observed result even without skill.

The table uses declared inputs and daily returns with normal tails. It holds the input Sharpe constant while varying duration and trials. Similar variants are not independent: keep that limitation visible without using it as permission to erase attempted configurations from the record.

Read the table before celebrating the curve

The table crosses search sizes with history lengths. Every row calls the calculator again; no cell comes from a viral screenshot. A larger search raises the luck reference, while a longer history generally narrows its dispersion. This helps explain why a curve selected after extensive testing needs more context than a rule specified beforehand. It does not establish a universal threshold for accepting a strategy.

Years and trades are different quantities. Many entries clustered in the same market episode may offer little variety. Inspect calendar coverage, inactive stretches and changing conditions as well as the number of rows in the exported file.

Stress costs at twice the baseline

Declared · The 2x cost test is a scenario: double the cost assumptions and compare with the baseline calculation. Record commission, spread, slippage and financing where relevant. Avoid subtracting a cost again if it is already included in net returns. If only aggregate returns are available and trades or expenses are unknown, that effect remains Not measured. A convenient guess does not fill the evidence gap.

Rigor examines cost sensitivity when the uploaded file contains the necessary information. Inspect the change and identify what remains unmeasured. A cost scenario cannot reproduce every execution condition on a platform, including the timing of fills during a disruption.

Expected Sharpe from luck

Declared · Assume 100 independent variants and 3 years of daily returns, with 252 periods per year and a declared annual Sharpe of 1.8. The calculator puts the expected Sharpe of the best unskilled variant at 1.47. This calculation assumes no skew and normal tails; it is not a measurement of a portfolio.

Declared · Illustrative inputs and computed results; these are not file measurements.
TrialsYearsSharpe from luck
Declared · 10Declared · 1Declared · 1.58
Declared · 10Declared · 3Declared · 0.91
Declared · 10Declared · 5Declared · 0.71
Declared · 100Declared · 1Declared · 2.54
Declared · 100Declared · 3Declared · 1.47
Declared · 100Declared · 5Declared · 1.14
Declared · 1,000Declared · 1Declared · 3.27
Declared · 1,000Declared · 3Declared · 1.89
Declared · 1,000Declared · 5Declared · 1.46

Separate development from evaluation

Reserve a period the agent has not seen and specify the comparison beforehand. If you change the strategy after observing that result, the period has influenced development. Keep it in the search record and explain the change. Repeating this cycle until an attractive curve appears does not restore the lost independence. An untouched period matters because it limits feedback into selection.

Also inspect whether the prices were available at decision time, how adjustments were applied and whether the universe retains discontinued instruments. A timestamp error can dominate any statistical correction. Ask the agent to explain choices and assumptions; a confident explanation cannot replace a reproducible check of the underlying data.

Take the evidence to the public reader

The linked public reader provides a starting point with declared figures. The calculator helps explore other search sizes. To examine the history itself, preserve the original file alongside the known attempts, frequency and costs. A balance screenshot does not contain this record, and a narrative cannot reconstruct missing observations.

In a report, Measured marks calculations from files, Declared marks information you supplied, and Not measured marks what could not be evaluated. The free first full report uses the same access as the other articles. Its findings can identify missing evidence or weaknesses in a backtest; they do not make a decision for you or connect the bot to an account.

FAQ

Does AI change how a backtest should be read?

It changes the speed of exploration, but data available at decision time, costs and discarded variants still matter. If the agent received earlier results, record that process. The selected file alone cannot reveal the entire search or show which alternatives influenced its selection.

Related

Put it to the test

The luck calculator is free and needs no sign-up. With an account, your first full report is free too.