A recent study by the technology firm Saturn found that popular chatbots, such as ChatGPT, Claude, Copilot, Grok, and Gemini, provided incorrect answers to an average of 57% of financial queries. The researchers evaluated 18 models using 121 questions, resulting in over 10,000 responses.

When it came to complex tasks requiring multiple calculations, the error rate jumped to 88%, with some individual models hitting 99%. The best performer was Claude Opus 5 in reasoning mode, though it still made mistakes 39% of the time. In one instance, advice from Claude Haiku 4.5 could have resulted in a tax penalty of £17,500 (approximately $23,400).