A Study Finds AI Chatbots Get Financial Questions Wrong 57% of the Time
A Saturn study put more than 10,000 financial questions to 18 AI models and found errors in 57% of answers on average, with even the best model wrong 39% of the time.
Ask a chatbot about a tax deduction and it will almost never say it doesn’t know. That confidence is the problem. A study by Saturn tested 18 AI models with more than 10,000 financial questions and found errors in 57% of the answers on average.
What the study measured
The study drew on more than a hundred financial questions and ran them against free and paid versions of models from ChatGPT, Gemini, Claude and Copilot, with Grok named alongside those systems as part of the mainstream field. Each question was put to a model up to five times. Repeating the same prompt is a way to check consistency: a model that answers differently on separate runs is harder to use than one that is reliably wrong in the same way. Across 18 models, the study says it asked more than 10,000 questions.
Financial questions make a harsh test case. In most domains, a plausible-sounding wrong answer is merely annoying. In tax and personal finance, the answer can usually be checked against a published rule, and the rule often carries an effective date. There is a right answer, and the right answer changes.
Three ways the answers go wrong
Saturn grouped the errors into calculation mistakes, missed changes to tax policy that were about to take effect, and invented rules, or hallucinations. The first category is the easiest to catch by redoing the math. The second is a data-freshness problem. Tax codes are updated on a schedule, and a model trained on older material can state last year’s threshold as if it were current. The third is the most dangerous for a general audience, because a hallucinated deduction or filing requirement arrives in the same fluent tone as a real one. The study warns that in the worst cases, someone acting on an AI’s tax answer could suffer serious financial loss.
Paid models score better, and newer beats older
Two patterns held: paid models answered more accurately than free ones, and newer models outperformed older ones. The top result came from Claude Opus 5 in reasoning mode, which still produced errors in 39% of its answers. Put another way, the best performer in the test was wrong about two out of every five questions.
That ranking is useful without being reassuring. It suggests the gap between tiers is real but not decisive. Moving from a free model to a paid one gets you fewer mistakes; it does not get you a system that can be handed a tax return and trusted to fill it in.
What it means for using AI with money
The practical lesson is about how you split the task, not which brand you pick. Chatbots are reasonable at the language parts of financial work: explaining what a term means, turning a vague worry into a specific question, drafting a list to bring to an accountant, summarizing a document you already have. They are weak at the parts that require being current and being right, especially rates, eligibility conditions and deadlines.
One low-cost habit is to ask the model for a source and then check that source directly. If it cannot name one, the answer is not a fact yet. Another is to treat any number as a draft until it has been verified against a primary document. A 57% average error rate is high enough that the next model release probably will not settle the question, and 39% from the current leader is not a margin anyone should put on a signed tax filing.