Analysis
Researchers at NPR, working with the media-literacy watchdog NewsGuard, built 30 test questions out of false narratives that Russia-, China- and Iran-aligned actors pushed between December 2025 and July 2026, then posed those questions to six of the most widely used AI chatbots -- OpenAI's ChatGPT, Google's Gemini, Microsoft's Copilot, Meta AI, xAI's Grok and Anthropic's Claude -- alongside the largest search engines, NPR reported on Aug. 30. The chatbots correctly identified and pushed back against the false narratives roughly three-quarters of the time, outperforming Google, Bing, DuckDuckGo and Yandex on the same set of questions.
The finding runs against a specific, well-founded worry that has circulated in AI-safety and information-integrity circles for the past two years: that state-aligned disinformation operations could exploit either a chatbot's training data or its live search-integration features to get the model to repeat propaganda as settled fact, especially on topics where authoritative Western sourcing is thin. NewsGuard has spent years cataloging exactly this kind of "LLM grooming" -- flooding the open web with false content specifically designed to be scraped into future model training runs -- and the working assumption going into this test was that chatbots would be meaningfully vulnerable to it.
Why the result is more fragile than the headline
Three-quarters correct is a genuinely good result, but it also means one in four answers either repeated a false narrative, hedged in a way that gave it undue credibility, or failed to push back clearly. NPR's methodology tested a fixed set of 30 questions against narratives already documented by NewsGuard's fact-checkers -- a favorable setup, since well-cataloged disinformation is exactly the kind of content most likely to have triggered safety training and fact-checking guardrails at each lab. Newer, less-documented narratives, or ones crafted specifically to avoid pattern-matching against known fact-checks, would be a harder test that this study didn't run.
The comparison to search engines is also doing real interpretive work here. Search engines were never designed to synthesize an answer -- they surface links and let a user judge sourcing themselves, which means a search engine "failing" this test looks different from a chatbot failing it: a bad search result is one link among ten, while a bad chatbot answer is presented as the single, confident, conversational response most users take at face value. Beating search on this specific test says less about chatbot safety in isolation than it does about how much more scrutiny a synthesized answer needs relative to a link list, precisely because users trust it more.
The backdrop this lands against
NewsGuard's own longer-running audits give this new result useful context. A separate NewsGuard tracking effort found that ten leading generative AI tools advanced Kremlin disinformation goals by repeating false claims from the pro-Kremlin Pravda network roughly 33% of the time in earlier testing, and that DeepSeek's chatbot specifically advanced China's position in response to related prompts roughly 60% of the time, per NewsGuard's AI Tracking Center. Read against that backdrop, NPR's fresh 75% pushback rate looks like real improvement across the mainstream US labs tested -- but it also implies meaningful variance by model and by narrative category that a single aggregate number obscures, and DeepSeek's much weaker performance on China-specific claims in NewsGuard's other testing is exactly the kind of result that gets lost inside a six-chatbot average.
This result also arrives in the same week Anthropic disclosed infostealer malware compromising Claude accounts -- a reminder that AI platform trust is being tested on multiple, unrelated fronts at once right now, not just the disinformation-resistance angle NPR and NewsGuard measured here.
For enterprise buyers evaluating which chatbot to deploy in a customer-facing or research context, this study is a genuinely useful, if imperfect, data point -- three-quarters accuracy on cataloged disinformation is a meaningfully better starting point than the alternative most people assumed going in. It is not, on its own, evidence that the harder version of this problem -- fresh, uncataloged, deliberately evasive disinformation -- is solved.