← run it on your own streams · this is a sample: four made-up streams over five years, no registration date, so the first caution applies to it too
JTOS Verdict
Each stream against the bar
The bar is the annualised Sharpe the luckiest of 4 strategies with no edge would show on this much data: 2.87. “Deflated” is the probability a stream's true Sharpe is above that bar, after its own skew and fat tails. It clears at 0.95.
| Stream | Periods | Return/yr | Vol/yr | Sharpe ± SE | Max DD | P(edge>0) | Deflated | Bar |
|---|---|---|---|---|---|---|---|---|
| trend-following | 60 | 47.2% | 11.7% | 4.04 ± 0.58 | -4.4% | 100% | 0.981 | clears |
| carry | 60 | 3.8% | 5.3% | 0.71 ± 0.45 | -6.0% | 94% | 0.000 | does not |
| my-backtest | 60 | 98.3% | 18.0% | 5.47 ± 0.67 | -5.7% | 100% | 1.000 | clears |
| discretionary | 60 | -6.4% | 20.7% | -0.31 ± 0.45 | -48.0% | 24% | 0.000 | does not |
Correlation
Shrunk (Ledoit-Wolf, intensity 0.93). Two streams above 0.9 are one stream.
| trend-following | carry | my-backtest | discretionary | |
|---|---|---|---|---|
| trend-following | 1.00 | -0.00 | -0.01 | 0.00 |
| carry | -0.00 | 1.00 | -0.01 | 0.01 |
| my-backtest | -0.01 | -0.01 | 1.00 | -0.01 |
| discretionary | 0.00 | 0.01 | -0.01 | 1.00 |
If you had to allocate today
P(best) is the posterior probability each stream has the highest true Sharpe, under a skeptical prior that tightens with the number of streams. Capital follows it as a risk budget, sized by volatility, capped. Equal weight is shown because it is what this has to beat.
| Stream | P(best) | Weight | Risk share | Equal weight |
|---|---|---|---|---|
| my-backtest | 57.7% | 50.0% | 70% | 25.0% |
| trend-following | 41.8% | 50.0% | 30% | 25.0% |
| carry | 0.5% | 0.0% | 0% | 25.0% |
| discretionary | 0.0% | 0.0% | 0% | 25.0% |
Effective number of independent bets: 1.7 of 4 · diversification ratio 1.39 · ex-ante vol 10.7%
Walk-forward: does allocating on this evidence beat equal weight?
48 rebalances, deciding each time only with data before that date.
| Evidence-weighted | Equal weight | Inverse vol | |
|---|---|---|---|
| Annualised return | 50.0% | 35.4% | 25.5% |
| Sharpe | 6.49 | 5.02 | 4.88 |
| Max drawdown | -0.3% | -2.0% | -1.0% |
Evidence-weighting added 1.61 Sharpe over inverse-vol out of sample.
Is picking the best one a procedure that works?
The deflated Sharpe above asks whether the best stream is real. This asks something the deflated Sharpe cannot: whether SELECTING it is a method that survives. Every way of splitting the record into two halves is tried; the winner of one half is looked up in the other. PBO is how often that winner turns out to be below average.
20 splits of 6 blocks · median out-of-sample rank of the in-sample winner 1.00 (0.5 is a coin flip) · mean Sharpe 1.60 in sample against 1.50 out. PBO recombines blocks in every order, so it is blind to time: an edge that worked for years and has since died still scores well here. The walk-forward above is what answers that question.
Is it still the same thing?
Every number above is a verdict on a fixed record. This asks whether the record still describes what these streams are doing now: the most recent 12 periods against everything before them.
| Stream | Sharpe before | Sharpe recent | Std errors apart | State |
|---|---|---|---|---|
| trend-following | 4.37 ± 0.67 | 2.85 ± 1.16 | -1.1 | within range |
| carry | 0.83 ± 0.51 | 0.12 ± 1.00 | -0.6 | within range |
| my-backtest | 5.25 ± 0.73 | 6.80 ± 1.71 | 0.8 | within range |
| discretionary | -0.40 ± 0.50 | 0.10 ± 1.00 | 0.4 | within range |
Average correlation 0.00 → -0.07, and 3.9 → 5.1 independent bets. The standard errors are on the difference between the two windows, so a short recent window is treated as the weak evidence it is rather than as a verdict.