Key Takeaways
- A study of AI shopping assistants found that 86% produced a factual conflict to potential customers
- Factual conflicts in the study included “contradictory prices, stale products presented as current, or spec claims engines couldn’t agree on”
- The study focused on the top four AI models, which included free and paid versions of ChatGPT, Gemini, Perplexity, and Claude
Maybe leave ChatGPT out of your enterprise purchasing decisions, because a new study found that these AI models are wrong 86% of the time when it comes to shopping.
Models like Gemini, Claude, Perplexity, and yes ChatGPT were all tested and were consistently found to provide incorrect prices and specs and suggest outdated products.
Even worse, the models weren’t just a little bit wrong; they really whiffed it. When asking about high-cost items like laptops, mattresses, and TVs, AI was incorrect by a median average of $300.
How Accurate Is AI for Shopping?
The study from Product.ai investigated the free and paid tiers for ChatGPT, Claude, Gemini, and Perplexity, asking each of them the same 220 shopping questions.
According to the report, the AI models provided conflicting and incorrect information in responses 86% of the time. Even worse, comparison questions that specifically matchup products against each other were incorrect 97% of the time.
This just in! View
the top business tech deals for 2026 👨💻
How Wrong Is AI for Shopping?
The study goes on to note that AI models aren’t just a little wrong when it comes to prices, features, and specs of high-end items like laptops, mattresses, and TVs. They were way off.
“When a price was wrong, it was wrong by a median of $300…There was almost no middle ground between exactly right and badly wrong.” – Product.ai researchers
On top of that, one out of every ten of these incorrect responses were wrong by as much as $500, which isn’t a small error; that’s a serious problem.
Why This Matters
Right now, this very second, there are people using AI to make purchases. One study from Bain & Company noted that around “30% to 45% of US consumers are using generative AI for product research and comparison.”
If that significant portion of shoppers are getting recommendations and pricing/feature information that is inaccurate, we could be talking about millions or even billions of dollars being misspent.
That doesn’t exactly create a positive customer experience. And in an era where trust is more valuable than anything for businesses, it might be better to leave AI alone for the time being.