Articles

Independent backtest audit: scope and evidence

What a statistical review of a trading strategy examines: variant selection, costs and the difference between measured and declared evidence.

A replication asks whether the same rules, data and assumptions reproduce a track record. A statistical audit asks what evidence that record contains after accounting for uncertainty, selection and missing information. These are complementary tasks. Replication can reproduce a choice fitted to historical data exactly; statistical review can identify weaknesses without rebuilding the logic behind the trades. Rigor analyses the supplied files and the declarations that accompany them. That scope matters when a fund compares research projects or a signal provider prepares a record for someone else's review.

What independent means in this review

Independence starts with separating the analysis from the product being examined. Rigor sells no robots or signals, and the report price does not depend on the resulting class. Its methodology publishes the questions, assessment rules and limitations. That makes it possible to discuss a particular conclusion and the file supporting it. It does not make the reviewer an observer of the entire research process: discarded experiments, earlier changes and unrecorded decisions can remain outside the materials supplied. Documenting those gaps is part of reading the review.

Statistical evidence against chance

The first check examines the support for the observed Sharpe ratio. The probabilistic Sharpe ratio accounts for record length, skewness and kurtosis. The report also examines serial dependence and resamples return blocks with a stationary bootstrap. Blocks preserve some local structure that shuffling isolated observations would discard. The result still depends on the supplied sample and the method. A short, irregular record or a record dominated by particular episodes calls for reading uncertainty alongside the headline figure. Repeating the same calculation does not remove that uncertainty or add observations.

Selection across variants

The next question is how many opportunities existed to find an attractive curve. The deflated Sharpe ratio compares the result with a reference that rises as more trials are considered. Rigor uses the largest count supported by the declaration, uploaded variants or optimisation passes. With a variants matrix, it can also examine overfitting through combinatorial cross-validation. The winning curve alone does not describe the discarded alternatives or the decisions that selected it for publication.

Declared · Assume 100 independent variants and 3 years of daily returns, with 252 periods per year and a declared annual Sharpe of 1.8. The calculator puts the expected Sharpe of the best unskilled variant at 1.47. This calculation assumes no skew and normal tails; it is not a measurement of a portfolio. The calculator illustrates selection under its assumptions; it does not measure the reader's strategy. Independence across trials is a simplification: similar variants can share much of their behaviour. Nor can this example reconstruct earlier searches that nobody documented.

Sensitivity to trading costs

The cost check recalculates trades at different friction levels and examines the break-even cost. It requires trade details and identifiable assumptions. A net return curve alone cannot separate commission, spread and slippage or establish how they would change at another size. If trades are missing, this section can remain unmeasured even when other statistics are available. The distinction matters for replication: reproducing the same cost assumption establishes consistency of the calculation, while sensitivity to different assumptions remains a separate question. Cost labels should accompany any comparison between records.

Behaviour outside the fitting sample

The report compares the period after the declared out-of-sample start with the preceding period. It examines their Sharpe ratios and the gap between them. The date is a client declaration; the file alone cannot establish that it was chosen before the results were seen. Changing rules after looking at that period changes its interpretation even when the dates remain intact. If the date is missing or either part lacks enough observations, the report identifies the limitation instead of constructing a comparison from an unsuitable split.

Data quality

Another dimension looks for problems in the materials: duplicates, spikes, frozen marks, deposits and patterns associated with martingale or grid behaviour, among others. A flag calls for examining its cause and context; it does not automatically reconstruct the original history. The absence of flags does not establish the file's provenance either. Rigor reads what is supplied and does not reconcile records with a broker. Keeping the original export and explaining transformations helps a fund respond to specific observations without confusing statistical cleanliness with evidence of origin.

Comparison with a benchmark

The review compares the record with the supplied benchmark using excess return, the drawdown ratio and the information ratio. Dates must overlap sufficiently to support that comparison. Choosing a relevant reference remains a research decision worth documenting: comparison with a different exposure may answer a different question. Without the benchmark or sufficient overlap, the comparison cannot be inferred from the strategy name. The dimension records the missing evidence instead of replacing it with an expectation. A benchmark comparison describes the supplied period, including its particular market conditions.

Measured, declared and awaiting data

Measured means the report computed a value from the supplied file. Declared identifies something stated by the client or platform that the analysis cannot establish. Not measured indicates insufficient information for that measurement. These labels apply to individual values: a computed Sharpe can sit alongside a declared trial count. Measuring a mathematical operation does not turn its premises into observed facts. When sharing the report, keep the labels and limitations next to the conclusions, including when a finding makes the original presentation harder to support or leaves a question unresolved.

FAQ

Does the review rebuild the strategy?

No. It analyses the received files and associated declarations. Rebuilding the signal or repeating the backtest requires rules, code and data that this review does not reconstruct.

Does a measured value describe future results?

No. It describes a calculation on the supplied material. Dependencies, assumptions and missing information still limit its interpretation; the report does not recommend buying, selling or investing.

Related

Put it to the test

The luck calculator is free and needs no sign-up. With an account, your first full report is free too.