Check your backtest

Paste or drop the CSV of a backtest you already ran. The page computes Sharpe, PSR, DSR, minimum track record length, skewness, kurtosis and maximum drawdown. These are conditional statistical readings. The strategy verdict stays NOT CONCLUSIVE because a CSV cannot verify the app's out-of-sample, cost, benchmark and selection prerequisites. It is computed on the server, without the desktop app.

Three formats are read: returns per period, an equity curve, or trades with a date and a profit or loss. A fourth, a matrix with one column of returns per configuration tried, adds the probability of backtest overfitting. The page tells you which format it detected before it shows a figure.

Public results

Shared results: 1. Desktop app: 0 strategy verdicts, including 0 GO and 0 NO-GO. CSV check: 1 published statistical reports, all NOT CONCLUSIVE for a strategy verdict. With fewer than 30 shared results, these counts are too small a sample to read. Only results their authors chose to share are counted; a check whose box is left unticked leaves no trace.

Check a backtest

The example is generated noise with a small drift, 300 trading days. It is not market data.

For a winner selected from a population of tested configurations, enter that population's size. DSR then describes selection luck under its assumptions. Adding a separate lever is not itself a reason to tighten a strategy verdict. Enter 1 for a preselected single series.

Used to annualize. For trades, the number of trades per year is measured from the dates instead.

If you leave it empty, the page assumes the dispersion of N strategies without skill and says so in the report.

How the figures are computed

The server reads the CSV cell by cell as numbers and dates. It runs no formula, macro or code from the file, and no language model produces any figure. Returns come from the file as decimals (0.01 is 1%; a value ending in % is divided by 100). An equity curve is turned into returns from one point to the next. Trades are read as profit and loss per trade; the number of trades per year is measured from the first and last dates.

SR = mean / standard deviation          (per period, biased moments)
PSR = Phi((SR - 0) * sqrt(T - 1) / sqrt(1 - skew * SR + (kurt - 1) / 4 * SR^2))
SR0 = sqrt(V) * ((1 - g) * Z(1 - 1/N) + g * Z(1 - 1/(N e)))     g = 0.5772...
DSR = PSR evaluated at SR0 in place of 0
MinTRL = 1 + (1 - skew * SR + (kurt - 1) / 4 * SR^2) * (Z(0.95) / (SR - SR0))^2

Phi is the standard normal distribution function and Z its inverse. Kurtosis is raw (3 for normal returns). V is the variance of the Sharpe ratios of your N trials. If you do not give their dispersion, the page assumes V = 1/T, the variance of a Sharpe ratio measured on T returns when the true Sharpe ratio is 0, and the report prints that assumption under V. With a matrix, V is measured on its columns and N is at least the number of columns. The formulas are those of the desktop app and of the deflated Sharpe ratio calculator, checked against the app on the same series.

Sources: Bailey and Lopez de Prado (2012), The Sharpe Ratio Efficient Frontier, Journal of Risk 15(2), equations 11 and 13. Bailey and Lopez de Prado (2014), The Deflated Sharpe Ratio, Journal of Portfolio Management 40(5), equation 2.

Why N matters

When a reported series was selected as the winner from a population of tested configurations, N describes that selection population. SR0 estimates the best Sharpe ratio expected from that population under the model's no-skill assumptions. The DSR reading depends on whether the declared N and dispersion describe that selection. Adding a separate lever is not itself a reason to impose a stricter strategy verdict.

Scope of the result

The CSV report is NOT CONCLUSIVE about the strategy, whatever its DSR. The app's strategy verdict needs verified in-sample and out-of-sample observations, charged costs, an out-of-sample directional benchmark, and selection-luck evidence. A returns, equity or trades CSV cannot verify those prerequisites. With fewer than 30 observations, even the conditional statistical reading is especially fragile. DSR and MinTRL remain displayed with their assumptions; neither establishes an economic verdict on its own.

PBO, computed from a matrix

The probability of backtest overfitting needs the returns of the configurations in the declared selection population, one per column. From a single series it cannot be measured. With a matrix it is computed by CSCV with up to 16 blocks and the Sharpe ratio as the selection score. With fewer than about 100 configurations it is exploratory. The PBO calculator lists the result of each split.

Limits and privacy

Files are limited to 2 MB, 100,000 rows and 200 columns, and each address to 30 checks per 10 minutes. Dates must be written YYYY-MM-DD, with an optional time. A check is computed in memory and is not stored. Pressing Publish stores the report: the figures and the date range are kept, the file and its column names are discarded, and the report gets a public page at a random address that search engines may index. It counts in the totals above. You receive a deletion code when you publish. Details are in the privacy policy.

The full backtest

The check above reads a result you already have. The desktop app builds the backtest itself: costs per market, look-ahead guards, walk-forward validation, CSCV over every configuration it tries, and the first-passage probability of passing a prop firm challenge before breaching its drawdown limits. Access is private for now; request an invitation to the app, or estimate pass rates first with the prop firm challenge simulator.