ChatGPT, Claude, Gemini, and Grok Are Not Ready to Brief American Voters
May 20, 2026 – 1:27 pm
A new generation of voters will ask ChatGPT, Claude, Gemini, and Grok how to vote, where the polling station is, and who is telling the truth. The published research is consistent: these models cannot reliably answer those questions. The election will arrive anyway.
In the spring of 2024, a Tow Center researcher at Columbia Journalism School ran a controlled experiment that should, in retrospect, have settled an industry argument. The team fed eight AI search products, including ChatGPT Search, Perplexity, Gemini, Copilot, and the Grok-2 and Grok-3 search modes, a set of 200 news articles drawn evenly from twenty publishers, then asked each tool to identify the article and credit its source. Across 1,600 queries, the models returned the wrong answer more than 60% of the time.
ChatGPT Search, the only tool that consented to answer all 200 queries, was completely accurate on 28% of them and completely wrong on 57%. Perplexity, marketed as the research-grade option, was wrong 37% of the time, the lowest failure rate in the cohort.
Those numbers were published over a year ago. They have not improved. A Bloomberg study summary published on May 20 confirmed that ChatGPT, Claude, Gemini, and Grok remain unreliable when asked about news, including election news. Nieman Lab’s read of the same data set found ChatGPT continues to be the worst of the four at crediting the news outlets it draws from. A separate NewsGuard False Claims monitor has the top ten generative AI chatbots returning false claims to news prompts 35% of the time in August 2025, up from 18% the year before.
The 2026 US midterms are 167 days away from the date of this writing. The first cohort of American voters who will, plausibly, use a chatbot as their primary news interface will go to the polls in November. NOTUS’s reporting on the campaigns has been blunt: ChatGPT and Claude will be a force in this election, and nobody, including the labs that built them, has a defensible plan for what happens when those forces produce confident, eloquent, well-cited answers that are also wrong.
What the published research shows, taken together, is not that chatbots occasionally hallucinate. The hallucination framing is a category error inherited from early 2024 discourse. The research shows something more specific and more dangerous for information integrity: chatbots misattribute quotes systematically, fabricate links that resolve to nothing, and cite syndicated or AI-summarized copies of articles in preference to the originals, severing the chain back to the journalists who produced the reporting. They cannot reliably distinguish between a Reuters wire, a content farm rewrite, and a Russian disinformation site dressed up in the same syndication wrappers. NewsGuard’s tracking of Moscow-seeded fake news sites found the top ten generative AI models mimicking Russian disinformation claims roughly a third of the time, citing the seeded sites as authoritative sources.
The structural reason for this is not a mystery, and the labs do not pretend it is. The training data pipelines that produce the current generation of frontier models have ingested the open web at a scale that includes both The New York Times and laundered outposts.