Public markets aggregate the views of millions of participants, many with powerful models and data of their own. Any pattern that reliably predicts prices tends to attract capital until the edge disappears — a dynamic that applies to AI just as it does to human strategies.
AI also inherits real limitations: models can be overconfident, can misread noise as signal, and — in the case of language models — can have knowledge cutoffs or reason from stale information. None of that is disqualifying, but it means an AI is not automatically better than a disciplined human or a simple index fund.
The only fair test is a forward-looking, out-of-sample track record measured against a benchmark such as the S&P 500, with every decision logged before its outcome is known. Backtests are easy to overfit; live results, marked at real prices, are the honest scoreboard.
This is exactly what makes public AI-trading experiments interesting: they put the question to a live, dated test rather than a marketing claim.