Mid-prices never cut it
The assumption your AI will happily optimize straight into a wall
I have quite often read something akin to “when you simulate trading for a sub-hour timeframe you should try to use more precision, but above that the mid-price entry assumption is sufficient for model development”.
Let me offer a heart-felt warning that I am sure will not surprise anyone who has followed this publication.
For strategies operating in the High Frequency to Medium Frequency space, real spread simulation with latency is treated, rightly, as the absolute minimum. The claim I want to puncture is: that above that timeframe, for longer hold times — hours, days, weeks — the mid-price assumption becomes sufficient.
No doubt at that timeframe slippage precision accounts for less. But is it immaterial? I could argue here invoking my own writing as to how the shape of a good trade and a bad trade differ, or that the price itself we get on entry is paramount. But the strongest counter-arguments are more mechanical, and much simpler.
Spreads are non-uniform
Say you have thousands of mid prices in your feed, and the model assumes entry on 20 of them1. Even if the average spread for the instrument is just above a tick, what gives you any confidence that you are not entering or exiting exactly on the odd mega-wide spread?
And your slippage numbers from other models don’t come to the rescue — they just muddy the water further. You would need a slippage distribution for exactly this model constellation. And even if you had it, you adjust one parameter in your model and you need a completely fresh one. No bueno.
There is no static cost you can measure once and bolt on afterwards, because the slippage is a property of the specific strategy.
And then the model finds the edge...
Now imagine you are talking to an AI/agent to do most of the modeling and data fetching for you. You are extremely pragmatic, and know the drill — variety of models working together, in-sample out of sample, stochastic seeds and the lot — you trust nothing but confirmed real numbers. The broad arch of the request is: find me alpha that is the most consistently profitable across the periods. Bit by bit, byte by byte your model homes in on a strategy you know is — just too good to be true. Days and weeks pass staring at the impressive results — again and again it is confirmed. But you are a seasoned professional, and on this timescale — holding periods of a few hours to a few days — that seems very unlikely.
The model learned, through a variety of features, how to find the widest-spread entries and exits.
No amount of post-fact slippage or spread modeling will offset it, because the strategy was selected for the very moments where the mid lies most. Another rabbit hole. For the less doubtful, it could have resulted in real money lost, of course.
None of this scenario needs AI. It is just a simple modeling fallacy — the mid-price assumption.
A big dent in limit executions
I already mentioned execution slippage — but if you intend to use limit orders at any point, or if you (like many others in the space) separate model development from execution, the news is worse still.
With a market order the execution is all but guaranteed; the actual cost becomes the big question mark. A limit order, on the other hand, can save you from this unruly slippage; but you may simply miss the fill. The mid-price assumption quietly grants you both — you fill, and you pay nothing — not something that real execution will ever allow. And on the limit side the failure isn’t a worse price, it’s a missing fill. Whether it happens on entry or exit, the strategy is now on a new path, and the outcomes have been irrecoverably altered.
Easy fix: just tell the AI to calculate the spread
In the proverbial AI model-development scenario, the immediate fix is quite easy: just ask your AI companion politely to use the opposite side — Bid/Ask — for the execution price. Please. And always. It is an order! (Stop being so damn polite with your AI tool. Really!)
Does it solve the limit order problem? Not fully. And is that as good as having latency simulation also? No, of course not. But it is still much, much better than a pure mid-price assumption.
I am not a “decimalist”. It is more a tribute to my 4-year-old who is excited now to count up to 20.


