JTOS Verdict · free
Does your strategy have an edge, or did you just test enough of them?
Paste the monthly, weekly or daily returns of any strategies you are considering. You get one page back: which of them clear the bar that the luckiest no-edge strategy would set on that much data, each Sharpe with its standard error beside it, which streams are really the same stream, what allocating on the evidence would have done against equal weight, whether choosing the best of them is a procedure that works at all, and whether any of them has quietly stopped being the thing you measured.
What the page says
- The bar
- Test ten strategies and the best of them looks good by chance alone. The bar is the annualised Sharpe the luckiest strategy with no edge would show on your data (Bailey & López de Prado). A stream clears it only if the probability its true Sharpe is above the bar is 0.95 or better, after its own skew and fat tails.
- The standard error
- Three years of a Sharpe-0.5 strategy carries a standard error near ±0.6. Most rankings on short records are noise, and the page says so instead of sorting the noise.
- One stream or several
- Two strategies correlated above 0.9 are one bet with two names. The correlation matrix flags them.
- Allocation, against equal weight
- Capital follows the posterior probability that a stream is the best one, sized by risk and capped. Equal weight is printed beside it because that is what it has to beat, and the walk-forward reports whether it did.
- Whether picking works at all
- A set of strategies can hold a genuinely good one and still be impossible to choose from, because the choosing overfits even when no single strategy does. Every way of splitting your record in half is tried, the winner of one half looked up in the other. At 50% the ranking carries no information; above it, picking the leader is worse than picking at random (Bailey, Borwein, López de Prado & Zhu).
- Whether it is still the same thing
- Everything above judges a fixed record. The last section asks whether the record still describes what these strategies are doing now: correlations and independent bets over the recent window against everything before it, and each stream against its own prior — measured on the standard error of the difference, so a short recent window is treated as the weak evidence it is rather than as a verdict.
A registration date turns a backtest into evidence: only periods after it count. Leave it blank and the page treats everything as in-sample and says so in the first caution. Returns should be excess over cash; if they are not, subtract your cash rate first. This is a statistical reading of what you supplied, not investment advice.