AI

NPR Tested 6 Chatbots on Propaganda; Most Held the Line

NPR and NewsGuard tested six chatbots against 30 questions built from Russian, Chinese and Iranian false narratives. Bots scored 75%; search summaries did worse.

Muhammet Fatih BatmanAugust 31, 20263 min read3 views
NPR Tested 6 Chatbots on Propaganda; Most Held the Line

In mid-July, reporters at NPR and analysts at NewsGuard sat down and asked six chatbots the kinds of questions a curious, slightly misled reader might ask about stories circulating online. The stories were not random: all traced back to Russian, Chinese and Iranian influence operations. The bots, ChatGPT, Gemini, Copilot, Meta AI, Grok and Claude, debunked them correctly about three-quarters of the time.

How the test worked

NewsGuard selected 15 false narratives that entered circulation between December 2025 and July 2026 and built 30 questions from them. NPR journalists and NewsGuard researchers then put those questions to all six products in the same period. A media literacy expert NPR consulted called the three-quarters result excellent by classroom standards.

Honest framing: 30 questions in a single round is journalism, not science. Its value is comparative, the same questions hitting six products at the same time, and NPR accordingly publishes a collective report card rather than a bot-by-bot ranking.

The results, good and bad

The finding cuts against the reflexive claim that chatbots are misinformation machines, at least on this sample. The recorded weak spots are still instructive. When Claude failed to debunk a narrative, its answers cited state-linked sources more often than the others did, and that detail matters: what a bot reaches for at the moment it fails determines whether an error stays small or compounds.

The genuinely bad grades went elsewhere. AI-generated search summaries performed clearly worse than the chatbots, with Bing's summaries the weakest of all and Google's AI Overviews better than Bing's but still behind the bots.

The weak link is not the chatbot

Here is the practical shape of the result: a user who asks a chatbot gets a more accurate answer than one who reads an AI search summary, yet search summaries are where millions of people actually meet information every day. If your company is ever the subject of a false story, it is likelier to surface in a search summary than in a chatbot answer, so it is worth occasionally checking what the summaries say about the queries that matter to you. The same logic applies to any chatbot you put on your own website: requiring it to cite sources is the single cheapest measure that separates the good performers from the bad in tests like this one. And a closing rule of thumb for internal use: three out of four is a solid grade, but the remaining quarter is exactly why answers without sources deserve a second look.

Sources: NPR

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk