Millions of people are turning to AI chatbots for help with budgeting, savings, and retirement planning — but a new study warns the advice they receive can vary wildly depending on which platform they use. Researchers at the University of Georgia tested seven of the most popular AI systems and found significant inconsistencies, including recommendations that differed based on the user’s race and gender. For anyone using an AI financial advisor to guide personal finance decisions, the findings are a timely reality check.

What the Researchers Found
The study, published in the Journal of Financial Planning and authored by UGA professors Swarn Chatterjee and Brenda Cude alongside Gianni Nicolini of the University of Rome Tor Vergata, put ChatGPT, Claude, Copilot, DeepSeek, Gemini, Meta AI, and Perplexity through identical financial planning scenarios. Emergency savings recommendations ranged from roughly $19,500 to $37,500 — a spread of nearly $18,000 for the same hypothetical household. Claude recommended about $10,000 more than the average of the other six platforms. According to UGA Today, ChatGPT, Copilot, and DeepSeek recommended higher emergency funds for women and African American users than for their white male counterparts, while Meta AI suggested women hold safer, stock-light portfolios.
Why the Inconsistency Matters
The one area where all seven chatbots agreed was the 4% retirement withdrawal rate — a well-established rule of thumb in financial planning. But agreement on a single benchmark doesn’t offset the broader pattern: on questions with no single correct answer, the platforms diverge sharply. A consumer following one chatbot’s advice could end up with a fundamentally different financial plan than a neighbor using a different app — even if both asked the exact same question.
What Users Should Do
Lead researcher Chatterjee summed up the takeaway bluntly: “AI gives people a starting point, not an ending point.” The study does not dismiss AI tools outright — the platforms generally stayed within defensible ranges and reflected sound general principles. But the demographic-linked variation raises a harder question about algorithmic fairness that goes beyond simple accuracy. For high-stakes decisions around retirement, emergency funds, or portfolio allocation, the researchers recommend validating any AI-generated guidance with a licensed human financial planner who can account for individual circumstances.
