Most of our work over the past two weeks happened on the machine learning side of Tantoryn AI. We were not just measuring model quality — we were asking a more important question: are the models learning something that is genuinely useful to a trading system?
The first step was finishing a new, causally correct dataset for machine learning. It is built from the state of the market and does not depend on the internal rules of the Candidate Generator. That makes model comparisons fairer and lowers the risk of an algorithm picking up hidden hints, whether from the future or from the trading system's own logic.
Next, we trained and compared several different families of models. Standard ML metrics were only part of the picture: we also checked whether good ranking actually turns into an economically useful result.
That is where an important limitation surfaced. From a machine learning point of view the models could show a useful signal, but that alone was not enough to treat the current setup as economically convincing.
Rather than keep tuning thresholds and model combinations, we stopped that line of work and went back to the learning objective itself.
After a dedicated Target/Objective Redesign stage, we chose a new setup for the Intra models: the model's main task is now to estimate how large the outcome of a market state is over a fixed 12-hour horizon.
Put simply, the model no longer answers "Will this trade be profitable or not?" Instead it has to answer "How good or bad is the expected outcome of this market state?"
For the current design we have switched off the separate binary profit/loss setup, because several different models exposed the same weakness in that formulation.
Along the way, we took a closer look at two problems that matter for any trading AI.
The first: high model accuracy does not by itself make a model useful in real trading. We explain why in "Why an “Accurate” AI Trading Model Can Still Be Useless in the Market".
The second: one lucky backtest, or one strong run of a model, does not prove that the model is robust. That is the subject of "Why One Great Backtest Does Not Prove a Trading AI Works".
We applied both principles in this cycle: repeated training runs, separate validation periods, and no picking a model just because one run happened to look especially good.
The main outcome of these two weeks is not a new return figure, and it is not the announcement of a "best model".
It is that we found a limitation in the previous problem framing before carrying it further into the system, and replaced it with a better-grounded learning objective.
The next step is to train models under the new setup and check whether an economically useful signal holds up outside the data used for tuning.