In machine learning for financial markets there is a problem that can make even a very attractive backtest untrustworthy: the model can accidentally receive information that did not yet exist at the moment the decision was made.

We covered this topic separately for a general audience, in an article explaining in plain language what temporal data leakage is, why splitting history into a training period and a test period is sometimes not enough on its own, and why point-in-time data matters — that is, data in the form in which it was genuinely available at that historical moment.

The English version of that detailed explanation, "Why a Trading AI Bot Can Accidentally “Know the Future” — and How to Prevent It", is published on Medium.

For Tantoryn AI this topic has practical consequences.

While preparing data and validating the ML system, we additionally revisited the principle that determines when information becomes visible to a model. For a historical experiment to be correct, what matters is not only the date a value refers to, but also when that value could genuinely have been known to the system.

This is especially important for data that is published with a delay, may be revised later, or is calculated over a period that has not yet closed.

So when we prepare historical data we follow a simple rule:

if the information did not yet exist at the moment of the decision, the model must not see it in the historical test either.

This approach can make results look less impressive on paper, but at the same time it makes validation of the system stricter and closer to real conditions.

For Tantoryn AI this is part of a broader development principle: first make sure the experiment itself is built correctly, and only then judge the quality of the machine learning.

We continue to apply this approach as we prepare data and further validate the project's ML components.