Retrieved article excerpt
Open article · Retrieved 2026-09-18T11:22:22.198238+00:00
# beatnothing
[tests](https://github.com/kaustubhspatil/beatnothing/actions/workflows/tests.yml)
[PyPI](https://pypi.org/project/beatnothing/)
[python](https://pypi.org/project/beatnothing/)
[license](https://github.com/kaustubhspatil/beatnothing/blob/master/LICENSE)
**Can your model beat doing nothing, after costs, without knowing the future?**
Every quant paper has a chart that goes up and to the right. This benchmark asks that
chart one question. It hands the same prices to the dumbest strategy imaginable, hold
everything at equal weight and never think again, charges both of them the same fee on
every trade, refuses to let either one see tomorrow, and measures the gap with an error
bar. That gap is the **Net Edge**. Twenty five contestants have tried so far: four neural
architectures, gradient boosted trees, the classic cross sectional factors, a cross
sectional ranking network, and a 2026 financial foundation model used with no training at
all. **None of them clears the bar.**
[Net Edge with 95% intervals for all twenty five contestants on the point in time universe, on the sealed 2022 to 2025 window and on 2026 to date](https://raw.githubusercontent.com/kaustubhspatil/beatnothing/master/leaderboard/pit/figures/net_edge.png)
*Every contestant, both windows, against the point in time bar. Green would mean the whole interval clears zero. Nothing is green.*
```
pip install beatnothing
```
## The rules of the game
| Rule | What it means in practice |
| --- | --- |
| **One engine** | Predictions become positions by one of three fixed rules a contestant declares: long every name with a positive value (the default), long the top decile, or long the top and short the bottom decile, dollar neutral, with a 50 bps a year borrow charge. Weights are used as given, gross exposure at most one, no leverage. Nobody gets a custom backtester. |
| **One bar, and a second one that could not choose** | The universe bar: always long, equal weight, same universe, same days, same costs. The thing you would earn with zero skill on the names you picked. Beside it, the investable bar: RSP, the equal weight S&P 500 ETF, which holds every index member by construction, dead ones included, and cannot have picked its universe with hindsight. The gap between the two bars is what picking the universe was worth. |
| **Costs on every trade** | 10 bps per unit of turnover, day one included. Gross and 20 bps numbers sit beside the net number so you can see who only wins for free. |
| **One score** | Net Edge = your net Sharpe minus the bar's net Sharpe on the same days, with a 95% paired stationary block bootstrap interval. You clear the bar only when the whole interval is above zero. |
| **Frozen means frozen** | Every submission records the sha256 of its model files and a registration date. A monthly job pulls new prices and rescores everything on the days that arrived after registration. Signals never change; only the calendar does. |
## Scoreboard, first season
Sealed window 2022 to 2025, 48 large cap US stocks, net of 10 bps. Full detail with
drawdowns, exposure and dollars in [`leaderboard/LEADERBOARD.md`](https://github.com/kaustubhspatil/beatnothing/blob/master/leaderboard/LEADERBOARD.md).
[Net Edge with 95% intervals for the nine season one contestants on the 48 name survivor universe](https://raw.githubusercontent.com/kaustubhspatil/beatnothing/master/leaderboard/figures/net_edge.png)
*The nine season one entries on the survivor universe. The chart at the top of this page is the same picture with every contestant, on the universe that did not know the future.*
| Contestant | Net Edge | 95% interval | Edge vs RSP | Net Sharpe | Gross Sharpe | Turnover a year |
| --- | --- | --- | --- | --- | --- | --- |
| Always long, the universe bar | 0.00 | | +0.44 [+0.19, +0.76] | +0.88 | +0.88 | 0.3× |
| Feedforward network | −0.03 | [−0.06, −0.01] | +0.41 [+0.17, +0.72] | +0.85 | +0.88 | 4.3× |
| Cost aware network, 10 bps term | −0.04 | [−0.10, +0.01] | +0.40 [+0.15, +0.72] | +0.83 | +0.87 | 2.9× |
| LightGBM, MSE objective | −0.14 | [−0.45, +0.09] | +0.30 [−0.13, +0.72] | +0.74 | +0.79 | 7.2× |
| LSTM, 60 day windows | −0.53 | [−1.00, −0.12] | −0.08 [−0.61, +0.44] | +0.35 | +0.62 | 43× |
| Cost aware network, no cost term | −0.70 | [−1.27, −0.09] | −0.26 [−0.89, +0.35] | +0.18 | +0.53 | 11.5× |
| 1D CNN, 60 day windows | −0.93 | [−1.55, −0.29] | −0.49 [−1.17, +0.16] | −0.05 | +1.08 | 265× |
| Linear regression, the incumbent | −1.17 | [−1.82, −0.62] | −0.73 [−1.42, −0.10] | −0.30 | +0.51 | 159× |
| Kronos small, zero shot | −1.67 | [−2.09, −1.26] | −1.23 [−1.71, −0.79] | −0.79 | +0.52 | 241× |
Read the two edge columns together. Three contestants clear RSP on the sealed window, and
so does the universe bar itself, by the same margin. They are not beating the index; the
universe is. A contestant that clears the investable bar while failing the universe bar
has demonstrated one thing only: that its universe was chosen with hindsight.
On the 176 trading days of 2026 that none of the frozen models had ever seen, the order
reproduces, every interval widens to include zero, and the two bars converge (RSP 1.30,
universe bar 1.44, gap inside the noise). Eight months cannot separate a network from a
bar. The leaderboard says so instead of ranking noise.
## What the first season taught us
[Gross versus net Sharpe for each season one contestant on the sealed window](https://raw.githubusercontent.com/kaustubhspatil/beatnothing/master/leaderboard/figures/cost_inversion.png)
- **Costs reorder the field.** The 1D CNN has the highest gross Sharpe of anything, 1.08, and a negative net Sharpe, because it turns the book over 265 times a year. Gross rank is not net rank. Papers that report gross numbers are reporting a different sport.
- **The MSE optimum is nearly always long.** LightGBM, early stopped honestly on validation loss, stops after three trees and holds 47 of 48 names. Squared error on daily returns is minimised by predicting the drift, and the drift is the bar wearing a different hat.
- **A foundation model obeys the same arithmetic as a linear regression.** Kronos small, pretrained on 12 billion bars across 45 exchanges and used zero shot, has a gross Sharpe of 0.52 and a net Sharpe of minus 0.79, with a 53% drawdown, because it flips positions 241 times a year. A hundred times the parameters of the feedforward network; the same cost inversion as the incumbent.
- **A cost aware objective repairs turnover, not alpha.** A network trained end to end on net Sharpe with a 10 bps turnover term trades three times a year instead of twelve and lands within 0.05 of the bar across three seeds, holding half the book in cash. Remove the cost term and the identical architecture loses 0.7 of Sharpe to fees. Given the true objective, the optimiser finds the bar.
## How we decide something is real
A leaderboard is a multiple test, and a Sharpe ratio is a badly behaved statistic. Three
things stand between a number in the table above and a claim worth acting on.
**The interval.** Net Edge carries a studentized circular block bootstrap interval
(Ledoit and Wolf, 2008): every resample recomputes not only the difference but a
heteroskedasticity and autocorrelation robust standard error for it, using the delta
method over the four moments that define two Sharpe ratios. The block length is chosen by
the Politis and White rule applied to the statistic's influence function, not to the
return series, because the difference of two return series has almost no autocorrelation
in its level while its squares are strongly persistent.
**The correction.** With fifteen contestants tested separately at five percent, a false
winner appears more than half the time. The same joint bootstrap resamples every
contestant on the same dates and feeds a Romano and Wolf stepdown, which controls the
chance of even one false claim across the whole board while keeping far more power than
Bonferroni. The column that decides a verdict is **p adj**, not p.
**The search.** A submitter who tried forty variants and shows you the best one has told
you nothing. `beatnothing.stats` ships the combinatorially symmetric cross validation
probability of backtest overfitting and the deflated Sharpe ratio, so a contestant with
many variants can be scored on what the search itself would have produced.
[Measured size, power, familywise error and overfitting probability of the benchmark's own statistics](https://raw.githubusercontent.com/kaustubhspatil/beatnothing/master/leaderboard/figures/stats_validation.png)
None of that is asserted. `scripts/validate_stats.py` simulates markets with fat tails and
clustered volatility, where the contestant is highly correlated with its bar because real
contestants are, and measures what the machinery actually does. On 250 simulations of four
years of daily data:
| Experiment | Result |
| --- | --- |
| No real edge exists. How often is one claimed? | studentized test **4.8%** against a promise of 5%; percentile interval 2.4% |
| A real edge of a third of a Sharpe exists. How often is it found? | studentized test **42.4%**; percentile interval 38.8% |
| Fifteen worthless contestants on one board. How often does one of them win? | tested separately **61%**; after the stepdown **1%** |
| Twelve variants of noise. What is the overfitting probability of the best? | **0.66**, where one half is what pure noise deserves; with one genuinely good variant among the twelve, **0.00** |
The studentized test keeps its word. The older percentile interval fires at about half its
nominal rate, and being conservative is not free: it misses real edges the studentized
test finds. Size stays near five percent as the record lengthens from two years to sixteen,
so what distortion remains is a finite sample effect and not a bug.
The third row is the one that matters most for a leaderboard, and it is why the verdict
column reports an adjusted p value rather than an interval. Six boards in ten would have
crowned somebody. On the real boards here the lowest adjusted p value is 0.997 on the
twenty five contestant board and 0.91 anywhere across both tracks, so nothing comes close.
Getting that experiment right was harder than it looks, and the first attempt was silently
wrong: adding zero mean noise to the bar leaves the mean alone but raises the variance, so
every simulated contestant was genuinely worse than its bar and the board could not have
produced a false winner at all. It reported zero percent both ways, which looked like a
result. Full numbers in [`leaderboard/stats_validation.json`](https://github.com/kaustubhspatil/beatnothing/blob/master/leaderboard/stats_validation.json).
## Verify your verifier
Most leakage is not in the model. It is in the evaluation. So the harness ships
contestants that cheat on purpose, and you run them through *your* pipeline first:
```
from beatnothing import Backtest, peek, off_by_one, hindsight_universe
Backtest(actual, predictions=peek(actual)).stats()["net_sharpe"] # tomorrow's return as today's signal: absurd, or your pipeline is broken
Backtest(actual, predictions=hindsight_universe(actual, 10)).stats() # the ten best names of the whole window, chosen at the start
```
If a canary does not score absurdly well, your pipeline is not measuring what you think.
The `truncation_test` in the same module proves a feature builder is trailing only by
deleting the future, recomputing, and demanding byte identical rows. Run on the point in
time panel it compares 46,605 rows across 60 names either side of a June 2024 cutoff and
finds a maximum difference of exactly zero on every one of the seventeen features
([`data/pit/truncation_test.json`](https://github.com/kaustubhspatil/beatnothing/blob/master/data/pit/truncation_test.json)).
## The universe knew the future
[Members on 31 December 2021, how many left the index, and how many no longer have prices](https://raw.githubusercontent.com/kaustubhspati