A published track record usually shows the version that survived the research process. Readers see its Sharpe ratio, but rarely see the discarded settings, revised windows or alternative universes tried before it was selected. The deflated Sharpe ratio addresses that selection: it asks how much statistical support the observed Sharpe retains against a benchmark that accounts for the search. The calculation describes evidence in the available history; it does not establish how the model will behave after publication.
What the deflated Sharpe ratio measures
The Sharpe ratio summarizes average return relative to its dispersion. Its estimate is uncertain: the uncertainty depends on history length, skewness and the tails of the return distribution. The probabilistic Sharpe ratio evaluates the observed estimate against a benchmark. The deflated Sharpe ratio, or DSR, uses the expected maximum Sharpe among unskilled attempts as that benchmark, accounting for how many attempts were made and how dispersed their estimates are.
DSR is expressed as a statistical probability under those assumptions, not as an adjusted annual Sharpe ratio. It is also not the probability that the next period will be positive.
The research ledger belongs with the result
Each choice considered when selecting the final version can extend the search: parameters, entry rules, filters, instruments and dates. Keeping only the selected file hides that context. Retain the optimizer log, discarded variants and the date when you fixed the selection criterion.
Rigor uses the largest count among the declaration and counts supported by compatible files. If the total comes from the author's estimate, it remains Declared. Counting columns or optimizer passes provides Measured evidence about that file, although it does not establish that no earlier experiments existed.
The calculator expresses the search in Sharpe units
The public calculator shares the report's luck calculation. It shows the expected Sharpe of the best unskilled configuration, the Sharpe remaining after a Bonferroni correction and the history length at which that benchmark would fall below the entered Sharpe. These quantities concern selection; none of them is the DSR probability.
The Bonferroni correction adjusts the p-value for the number of attempts and converts it back into Sharpe units. The calculated history length holds the scenario's assumptions and Sharpe constant: it is not a waiting period that resolves uncertainty. Changing the history, costs or research search requires reassessing the evidence.
How to read the scenario table
The table combines different trial counts and history lengths while holding the assumed Sharpe constant. Its inputs and outputs are labelled Declared: they are calculator scenarios, not measurements of a portfolio. Each cell reports the expected luck Sharpe, not a DSR or a strategy's return.
First compare a fixed history length as the search expands. Then hold the trial count constant and change the length. This separates the effect of selecting among more variants from the effect of estimating with more history. A scenario describes that relationship under daily, normal-return assumptions; it does not substitute for the distribution in a client's file.
Expected Sharpe from luck
Declared · Assume 100 independent variants and 3 years of daily returns, with 252 periods per year and a declared annual Sharpe of 1.8. The calculator puts the expected Sharpe of the best unskilled variant at 1.47. This calculation assumes no skew and normal tails; it is not a measurement of a portfolio.
| Trials | Years | Sharpe from luck |
|---|---|---|
| Declared · 10 | Declared · 1 | Declared · 1.58 |
| Declared · 10 | Declared · 3 | Declared · 0.91 |
| Declared · 10 | Declared · 5 | Declared · 0.71 |
| Declared · 100 | Declared · 1 | Declared · 2.54 |
| Declared · 100 | Declared · 3 | Declared · 1.47 |
| Declared · 100 | Declared · 5 | Declared · 1.14 |
| Declared · 1,000 | Declared · 1 | Declared · 3.27 |
| Declared · 1,000 | Declared · 3 | Declared · 1.89 |
| Declared · 1,000 | Declared · 5 | Declared · 1.46 |
Where independence becomes a limitation
The expected-maximum benchmark treats attempts as independent. In a research search, nearby configurations often share signals, positions and returns. The raw count and the effective number of independent attempts can differ. Do not reduce the count until the reading looks desirable: describe the dependence and show sensitivity to other counts. The calculator does not estimate that correlation from the fields it receives.
Consecutive returns within the same history can also be dependent. That affects how much information each observation contributes and is distinct from correlation between variants. The report incorporates dependence adjustments into its statistical analysis; even so, data quality and how faithfully the search process is represented continue to limit interpretation.
Measured and Declared answer different questions
Measured identifies a quantity obtained from the supplied file or a calculation using it. Declared identifies information supplied by the author that the file does not establish. Not measured indicates missing evidence for a check. A calculated statistic can depend on a declared trial count, so read the labels on the inputs alongside the result.
The calculator has no file: Sharpe, duration and trial count are assumptions supplied by its user. In the report, the series allows return moments to be measured and its time structure to be examined. That difference in evidence explains why a public illustration and an analysis of actual data can produce different readings.
What to publish alongside the track record
State the period examined, observation frequency, cost treatment and separation between research and out-of-sample evaluation. Add how the trial count was obtained and which part of the search you could not reconstruct. If a variants matrix exists, retain its dates and the relationship between each column and the version tested.
Readers need those conditions to interpret DSR. A favorable multiplicity result still leaves data quality, exposure, costs and stability over time to examine. Link the methodology and deliver the limitations with the same history you show the committee. Statistical review contributes documented questions; it does not replace the investment decision or settle whether the proposed implementation matches the supplied series.
FAQ
Is the calculator's remaining Sharpe the DSR?
No. It is a Sharpe after a Bonferroni correction. DSR is a statistical probability against a selection benchmark; these are different quantities.
What if the number of attempts is unknown?
Declare that uncertainty and present scenarios with other counts. Gather logs and variants before interpreting the figure as a description of the entire research process.