A couple of months ago, I wrote about an experiment where researchers from the San Francisco Fed asked ChatGPT to forecast inflation. The results were miserable. Once the researchers tried to forecast inflation outside of the training window of ChatGPT, the model broke down, indicating that even if you ringfence your model, the training data used for the model in the first place leaks through and gives you forecasts that look much better in a backtest than when applied to live markets.
A team of Chinese researchers now used state-of-the-art AI trading agents to assess if the same problem exists when developing trading strategies. They asked five LLM-based methods and restricted them to run on GTP-4o so that the training cutoff was known…





