← run it on your own streams · this is a sample: four made-up streams over five years, no registration date, so the first caution applies to it too

JTOS Verdict

4 streams · 60 common periods at 12/yr · generated 2026-09-13 02:42 UTC · nothing you uploaded was stored
2 of 4 streams clear the multiple-testing bar (trend-following, my-backtest) on 5.0 years. The rest have not demonstrated an edge that survives selection.
No registration date was given, so every period counts as evidence. If these are backtests, they are in-sample: the numbers below are the best case, not the expectation.

Each stream against the bar

The bar is the annualised Sharpe the luckiest of 4 strategies with no edge would show on this much data: 2.87. “Deflated” is the probability a stream's true Sharpe is above that bar, after its own skew and fat tails. It clears at 0.95.

StreamPeriodsReturn/yrVol/yrSharpe ± SEMax DDP(edge>0)DeflatedBar
trend-following6047.2%11.7%4.04 ± 0.58-4.4%100%0.981clears
carry603.8%5.3%0.71 ± 0.45-6.0%94%0.000does not
my-backtest6098.3%18.0%5.47 ± 0.67-5.7%100%1.000clears
discretionary60-6.4%20.7%-0.31 ± 0.45-48.0%24%0.000does not

Correlation

Shrunk (Ledoit-Wolf, intensity 0.93). Two streams above 0.9 are one stream.

trend-followingcarrymy-backtestdiscretionary
trend-following1.00-0.00-0.010.00
carry-0.001.00-0.010.01
my-backtest-0.01-0.011.00-0.01
discretionary0.000.01-0.011.00

If you had to allocate today

P(best) is the posterior probability each stream has the highest true Sharpe, under a skeptical prior that tightens with the number of streams. Capital follows it as a risk budget, sized by volatility, capped. Equal weight is shown because it is what this has to beat.

StreamP(best)WeightRisk shareEqual weight
my-backtest57.7%50.0%70%25.0%
trend-following41.8%50.0%30%25.0%
carry0.5%0.0%0%25.0%
discretionary0.0%0.0%0%25.0%

Effective number of independent bets: 1.7 of 4 · diversification ratio 1.39 · ex-ante vol 10.7%

Walk-forward: does allocating on this evidence beat equal weight?

48 rebalances, deciding each time only with data before that date.

Evidence-weightedEqual weightInverse vol
Annualised return50.0%35.4%25.5%
Sharpe6.495.024.88
Max drawdown-0.3%-2.0%-1.0%

Evidence-weighting added 1.61 Sharpe over inverse-vol out of sample.

Is picking the best one a procedure that works?

The deflated Sharpe above asks whether the best stream is real. This asks something the deflated Sharpe cannot: whether SELECTING it is a method that survives. Every way of splitting the record into two halves is tried; the winner of one half is looked up in the other. PBO is how often that winner turns out to be below average.

Probability of backtest overfitting 0%, over 20 splits of 6 blocks. The in-sample winner usually stays a winner, keeping 94% of its Sharpe out of sample. This is the evidence that a ranking on this data means something.

20 splits of 6 blocks · median out-of-sample rank of the in-sample winner 1.00 (0.5 is a coin flip) · mean Sharpe 1.60 in sample against 1.50 out. PBO recombines blocks in every order, so it is blind to time: an edge that worked for years and has since died still scores well here. The walk-forward above is what answers that question.

Is it still the same thing?

Every number above is a verdict on a fixed record. This asks whether the record still describes what these streams are doing now: the most recent 12 periods against everything before them.

No drift worth acting on: average correlation 0.00 → -0.07, 3.9 → 5.1 independent bets, and every arm is inside two standard errors of its own record over the last 12 periods.
StreamSharpe beforeSharpe recentStd errors apartState
trend-following4.37 ± 0.672.85 ± 1.16-1.1within range
carry0.83 ± 0.510.12 ± 1.00-0.6within range
my-backtest5.25 ± 0.736.80 ± 1.710.8within range
discretionary-0.40 ± 0.500.10 ± 1.000.4within range

Average correlation 0.00 → -0.07, and 3.9 → 5.1 independent bets. The standard errors are on the difference between the two windows, so a short recent window is treated as the weak evidence it is rather than as a verdict.

Sharpe is measured on the returns as uploaded; if they are not excess over cash, subtract your cash rate first. Standard errors follow Lo (2002); the multiple-testing bar and deflated Sharpe follow Bailey & López de Prado (2014); the probability of backtest overfitting is combinatorially symmetric cross-validation, Bailey, Borwein, López de Prado & Zhu (2015); allocation is risk-budgeted equal risk contribution on a Ledoit-Wolf covariance. This is a statistical reading of the data you supplied, not investment advice.