← All guides
Research · data collected August 11, 2026 · published September 1, 2026

How accurate are AI shopping assistants about prices?

We asked ChatGPT, Claude, and Perplexity the same three questions about real Amazon products — current price, all-time low, and buy-or-wait — with web search on, then graded all 117 answers against recorded price history, with a 2% tolerance. They answered 57 of 117 correctly. None cleared 60%.

The scores

AssistantQuestionsCorrectAccuracy (95% CI)Fabricated prices
ChatGPT (gpt-5.5)4122 of 4158% (42–72%)2 (5%)
Claude (claude-sonnet-5)137 of 1354% (29–77%)0
Perplexity (sonar-pro)6328 of 6353% (40–66%)5 (9%)

“Fabricated” is the score that matters most: a specific dollar figure that matches no price the product ever sold at — stated confidently, often with a citation. Claude fabricated nothing in our sample; ChatGPT and Perplexity both did.

What a fabrication looks like

Asked the all-time-low price of a Dyson V15 Detect Plus, ChatGPT answered “$489.99” and cited a price-tracking site for that exact figure. The recorded all-time low is $499.99. The number was specific, sourced, and wrong — which is precisely why this failure mode is worse than saying “I don’t know.” No assistant in our sample ever abstained.

Method, and what this doesn’t show

Every assistant got identical questions with web search enabled where the API supports it. Answers were parsed into claims by one fixed extractor model (never graded itself), then graded against Keepa-derived Amazon price history. 30 products made the graded set; 57 were excluded as ambiguous (multiple variants or unstable listings) before any grading.

The limitations, plainly: the samples are small and the confidence intervals above are wide — treat the ranking as provisional, not settled. Claude’s sample (13 questions) is smaller than the others because collection stopped when API credit ran out, and a share of collected answers remains ungraded for the same reason. We will extend the sample and re-run the audit; the method is fixed and pre-registered, so future runs are comparable.

Why we can grade this at all: Espresso Steals tracks daily Amazon price history and publishes its own hit rate on every buy/wait call it makes — see our track record. We hold ourselves to the same standard we graded the assistants against.