The past three weeks at Tantoryn AI covered several fronts at once: growing the public side of the project, getting ready to expand the team, and continuing a long cycle of machine learning research.
Preparing for the project's next stage
As Tantoryn AI grows, we have started reshaping how the project is organised internally, so that it works not only for its founder but also for new people joining it.
We have put the description of the project's architecture and its areas of work into a systematic form, brought the roadmap up to date and started building a separate context for future team members. The roadmap deliberately keeps the order of development apart from indicative dates: the sequence of stages matters more than trying to pin research work to a rigid calendar in advance.
At the same time, we are preparing to grow the team. Among the first roles under consideration are a Lead Developer and a Community Manager.
For the Community Manager, a dedicated public-facing area is gradually taking shape: the website, publications, social platforms and feedback from our audience. Later on, we plan to support part of this work through TAICM — the Tantoryn AI Community Manager — but publications and replies will still require human approval. TAICM itself is, for now, a direction of development rather than a running system.
Our public side has grown too
Over the same period, we kept developing our Public Build with Controlled Disclosure model.
The idea behind it is simple: talk not only about successful results, but also about how the research is going — the checks, the hypotheses that turned out wrong and the reasons decisions changed — while not publishing the trading core, private data or details that would effectively reveal how the system's algorithm works.
On the website, we expanded the contacts and the links to the project's public resources. Tantoryn AI's research pieces now appear on Dzen and Medium, we also use HackerNoon and other public platforms, and news about the project itself stays on tantoryn.com.
This separation matters to us. An article should explain a problem on its own terms, independently of Tantoryn AI. A project news post should show where exactly we ran into that problem and what we changed in our work as a result.
That is why we try not to copy the same material from one platform to another, and instead connect popular-science pieces with the practical story of how the project is being built.
In ML, we went through several hypotheses that did not work out
Even so, the ML work led us to an important practical conclusion: a complex model is valuable not because it is complex, but only if its advantage survives a fair comparison with simpler solutions.
Most of the technical work in recent weeks centred on the ML Evaluator — the project's AI-based evaluation layer.
We tested different model families and different ways of framing the task one after another, using development and validation stages separated in time, repeated training runs and experiment rules fixed in advance.
Some of the results looked promising in terms of how well the models could rank market states. But further checks exposed a problem: good ML ranking does not yet mean that an economically useful edge has been found.
We did not turn that result into a claim of a successful trading model. Instead, we stopped that line of tuning and rethought how the task itself was framed. The research branch was closed as a historical experiment, with no claims of profitability or readiness for real trading.
That led to a more fundamental change. The ML Evaluator's role was brought back to a narrower task: estimating how likely a trading scenario is to play out successfully given the state of the market, while the economic result of the whole system is tested separately. For this we set up a new classification task that distinguishes a failed, a partially successful and a successful outcome.
The new model showed a signal — but the next problem appeared
After reframing the task, we prepared a new training cycle and ran the first full check of one of the models.
The result was useful precisely because it did not give a simple answer such as “the model is good” or “the model is bad”.
The model did show an ability to tell more promising market states from less promising ones, and it passed the checks set for this stage.
At the same time, we compared it with a much simpler reference point built on a small group of basic characteristics of the market situation.
On the key independent check, the simple baseline turned out stronger than the complex ML model.
That does not make the training we did useless. On the contrary, a comparison like this shows that our research controls are working: the added complexity has not yet proven that it adds value.
That is why, right now, it matters more to us to understand what information the simple baseline relies on, and why the more complex model could not extract more from the data, than to push a good-looking model number higher at any cost.
We looked into this problem separately
In parallel, we looked at how typical this situation is in modern financial ML — and it turned out to be far from a problem specific to Tantoryn AI.
We shared the findings of this research in more detail in a separate article on Medium: “When a Simple Baseline Beats AI: Why More Complex Doesn’t Always Mean Better”.
Recent research on financial time series paints a mixed picture: modern neural and Transformer-based approaches can indeed outperform classical methods, but not always. In some regimes, linear and autoregressive baselines remain competitive or even beat far more complex architectures.
So the key question is not how complex a model is, but whether its advantage holds up with the same data, the same time split and the same evaluation rules.
We wrote a separate piece on this: “When a Simple Baseline Beats AI: Why More Complex Doesn’t Always Mean Better”.
In it, we look at the problem independently of Tantoryn AI: why financial ML needs simple baselines, why a win on a single metric is not enough, and why a negative result can be more useful than yet another round of model tuning.
What comes next
The next step in ML is not to make the system more complex by default.
First, we need to understand why the simple reference point turned out stronger than the first tested model of the new series. After that, the other approaches have to go through the same principle of comparison: added complexity has to prove that it adds value.
Alongside this, work will continue on Tantoryn AI's public side and on preparing the project for a larger team.
So the main result of recent weeks is not a single new model or feature. The project is gradually becoming more structured in three directions at once: research, public transparency and the organisation of a future team.
And in ML we once again got a result that we believe is worth keeping, even when it is inconvenient: if a simple baseline turns out stronger than a complex AI, the right next step is not to hide the baseline, but to find out why.