Sharpe 15 in training, zero out of sample: how our trading models fooled us
Seven of eight candidate strategies had no edge on held-out data. The most expensive failure was not a bug: we trained on one venue's spot data and executed against another venue's perpetuals.
Across roughly eighteen months of sessions we built eight candidate trading strategies. In training they looked extraordinary: Sharpe ratios of 5 to 15 and higher, on curves smooth enough to put in a deck. When we finally subjected all eight to held-out evaluation, seven of them had no genuine edge at all.
Research write-up, not financial advice.
That is the headline, and it is not the interesting part. Overfitting is the oldest story in quantitative finance and nobody needs another post explaining that a curve fit to noise will not repeat. The interesting part is what the postmortem found underneath the overfitting: our most expensive failure was not a modelling error, not a leak, not a bug in the usual sense. We had trained models on one venue’s spot data and executed them against another venue’s perpetual futures. Nothing was broken. The pipeline ran clean, the tests passed, the numbers came out. They were simply numbers about a market we were not trading.
Our internal catalog of this period lists six failure categories. Publishing all six would make this a list, and a list of failures reads as either a shrug or a confession. It is also the wrong shape, because three of those six turned out to be the same failure at three different altitudes. That is what this post is about.
The spine: measuring one market, trading another
Here is the shape, stated once so you can watch it recur.
A model is trained against some representation of the world. A backtest scores it against some representation of the world. A live system executes against the actual world. Every performance number you produce is a claim about the relationship between those three. If any two of them differ in a way nobody wrote down, the number is not too optimistic or too noisy. It is a measurement of a different object, and no amount of statistical care downstream can recover it. Regularisation does not help. More data does not help. A better model makes it worse, because a better model fits the wrong world more tightly.
We hit this three times, at three scales, and only recognised it as one thing in the postmortem.
One: the instrument
The most expensive instance, and the one with the least excuse.
Our order-flow features - the microstructure signals that were supposed to carry most of the short-horizon information - were computed from Binance spot data. Every execution run targeted Hyperliquid perpetuals. This was not a decision anyone made. It was an inheritance: the research pipeline was built first, against the data source that was easiest to backfill, and the execution venue was chosen later for reasons that had nothing to do with where the features came from.
Spot and perpetual markets for the same asset are correlated. They are not the same market. They have different participants, different leverage, different liquidation dynamics, and a funding mechanism on one side that does not exist on the other. When we finally measured it directly, the spot and perp versions of one core order-flow series agreed 57% of the time.
Fifty-seven percent. On a two-sided directional signal, where a coin flip gives you fifty. Close to half of the order-flow information the models were trained on was, relative to the venue we actually traded, noise.
Everything downstream of that inherits it. Feature importances rank the wrong features. Hyperparameters tune to the wrong distribution. And critically, the backtest confirms all of it, because the backtest was reading the same wrong data. The system was internally consistent and externally about a different market. Months of work were invalidated by a mismatch that nobody had to introduce, because nobody had to state the assumption in the first place.
The fix was not clever. It was a parallel perpetuals data pipeline, so features are computed from the same instrument being traded. It cost an epic’s worth of engineering to buy back a property we had assumed we already had.
Two: the label
The same shape, one level down, and it was in the code for months before anyone saw it.
Our models were trained to predict a fixed-horizon outcome: what happens over the next five bars. Our backtester did something else. It held a position until the model’s signal flipped.
Read those two sentences again, because it took us a long time to. The model was optimised to answer one question and graded on the answer to another. A model can be excellent at predicting five-bar returns and useless at deciding whether to hold until reversal - these are different problems with different optimal policies, and there is no reason the first should transfer to the second. The backtest was not measuring the model. It was measuring the model plus an exit rule the model had never been shown.
Compounding it, our cross-validation was standard k-fold. On time-ordered data, k-fold trains each fold on data that comes after its own test set. The model gets to see the future, in the ordinary sense of the phrase, and every fold-level metric it produces is fiction. This is a textbook error and we made it anyway, for the usual reason: k-fold is the default, defaults are invisible, and nobody has to justify a default in code review.
The fix was two changes made together, and they belong together. Forward-only splits, so a model is never fitted on data that postdates its evaluation window. And triple-barrier labels, which define the training outcome as whichever comes first - a profit target, a stop, or a time limit - because that is what actually terminates a position in the backtest. After the change, the thing the model learns to predict and the thing the backtest scores are the same event.
Note what that fix is. It is not a better model or a stronger regulariser. It is the alignment of two definitions that had drifted apart while nobody was looking at both of them at once.
Three: two portfolios, then two models
The same shape once more, one level up, and this one is subtle enough that it fools people who have already learned the first two lessons.
Walk-forward validation is the honest alternative to a single train/test split, and we used it: retrain on a rolling window, evaluate on the next segment, roll forward, repeat. For February 2026 it reported a Sharpe of +10.38.
Backfill over that same February reported -4.47.
A gap of nearly fifteen Sharpe points between two evaluations of the same strategy and month is not a discrepancy to reconcile. It is an alarm that the evaluations are measuring different objects. Here they simulated different portfolios. Walk-forward applied no concurrent-position cap and no one-position-per-ticker rule. Backfill applied both. When the capped portfolio filled, backfill dropped the remaining trades in alphabetical ticker order rather than by signal quality.
Lifting the concurrent-position cap brought backfill to Sharpe 11.14, directionally matching the uncapped walk-forward result of 10.38. The sign flip was not evidence about edge. It was evidence about which portfolio each system simulated.
The remaining trade-count difference also had concrete causes: an entry-bar barrier effect in backfill and the difference between backfill’s multi-ticker simulation and walk-forward’s per-ticker evaluation. We closed that gap as structural and expected. It was not evidence of an unreplicable ensemble.
There was still a separate model mismatch. On the same February, an otherwise matched run using the fold model trained to that month’s cutoff scored Sharpe 30.42. The single production model scored 11.14. Each walk-forward segment is served by a model freshly fitted to the conditions immediately before it. A single frozen production model does not have that property.
That is the same spine again: the first comparison measured two different portfolios, and the second measured two different models. Neither comparison was about a weaker or stronger version of the same trading system. Walk-forward Sharpe remains a ceiling, not a target. We treat backfill on a single frozen model as the closer proxy, and we treat a gap between the two as a diagnostic rather than a puzzle.
What survives
The eight-strategy consolidation was the first time we subjected every candidate to held-out evaluation at once. Before it, in-sample metrics were the primary development signal, which is the condition under which all three of these failures are invisible - each one produces a better in-sample number, so each one looks like progress.
Seven of the eight had no edge at the signal-generation level. That is worth stating plainly, because the tempting response is to reach for machine learning: put a meta-labeller on top, filter the trades, let a bigger model find the good subset. It does not work, and the reason follows from the spine. A filter on top of an edgeless strategy selects a subset of losing trades with different characteristics. There is no signal underneath for it to recover.
What we take from this period is narrower than “be careful about overfitting” and more useful:
-
A performance number is a claim about a specific market, measured a specific way. It is not a scalar you can carry between contexts. Write down which instrument, which label, and which model produced it, or you will eventually compare two numbers that are not about the same thing.
-
The costliest failures are agreements nobody made. Nobody chose to train on spot and trade perps. Nobody chose fixed labels with reversal exits. Both were the residue of two reasonable decisions made months apart by people who never had to reconcile them. Bugs get found; unstated assumptions do not, because there is no code to inspect.
-
Measure the assumption directly when you can. The spot-perp mismatch was arguable in the abstract for a long time and settled in an afternoon by computing agreement between the two series. Fifty-seven percent ended the conversation. If an assumption is load-bearing, it is usually cheap to put a number on it, and the number is worth more than the argument.
-
Two evaluations that disagree are more informative than one that agrees with you. The +10.38 against the -4.47 is the most valuable single result in this whole period, because it exposed two portfolio simulations that had been treated as comparable.
None of this was recovered by better modelling. All of it was recovered by making the training environment, the evaluation environment, and the execution environment describe the same market. That is unglamorous infrastructure work, it is where our next phase went, and it is a prerequisite rather than an achievement: getting it right does not produce an edge. It only makes it possible to tell whether you have one.
We do not, yet, on seven of those eight. That is the honest state of it.
Research write-up, not financial advice. Every figure above is from our own session logs and internal postmortem across sessions 204 to 327. They are historical research results on our own data, not forecasts, and not a claim that any of these strategies works.
We were not overfitting a market. We were fitting one market and grading ourselves on another.
- +Compute features from the same instrument you execute on, or the backtest is measuring a different market
- +The label the model learns must be the outcome the backtest scores, or the two are answering different questions
- +Cross-validation that shuffles time is not validation on time series: use forward-only splits
- +A fold model scored Sharpe 30.42 where the single production model scored 11.14 on the same February, so walk-forward Sharpe is a ceiling, not a target
- ×Order-flow features came from spot data while every execution ran on perpetuals, and the two agreed only 57% of the time
- ×Models were trained on 5-bar fixed-outcome labels while the backtest held positions until the signal flipped
- ×K-fold cross-validation on time-ordered data let each fold train on its own future
- ×Walk-forward reported Sharpe +10.38 for February 2026 while backfill reported -4.47 for that same month
- ×Seven of eight candidate strategies survived to the consolidation phase on in-sample metrics alone
GPU passthrough into LXC, without the pain
Four identical RTX 5060 Ti cards, one Proxmox container, and a PCIe lane map that quietly decides which card is allowed to talk to which. The passthrough was the easy part.
Designing an AI options trader that cannot go rogue
Seven independent safety rails, each defeatable on its own. Together they leave the current codebase without a live-order path.
What ran, what shipped, what died - with the numbers behind each. No threads, no hype.