Skip to content
OmniAxis
Get started
Insights

Confidently Wrong: What the AI-Search Citation Studies Actually Found

Independent studies keep measuring the same thing: AI search answers with confidence and cites with error. What that means for competitive research.

When the European Broadcasting Union and the BBC ran the widest test yet of generative AI assistants, they did not ask anything hard. Professional journalists from 22 public service media organizations across 18 countries evaluated more than 3,000 responses in 14 languages from ChatGPT, Copilot, Gemini, and Perplexity, grading each for accuracy, sourcing, and context. 45 percent contained at least one significant issue.

Not subtly wrong. Hallucinated detail presented as fact, and above all sourcing that did not hold: sourcing was the single largest failure category, at 31 percent of responses. Most instructive for anyone who relies on these systems, the errors arrived without a flicker of doubt.

The Tow Center at Columbia had already named that register a year earlier, testing eight AI search engines on 1,600 news-citation queries and finding more than 60 percent answered incorrectly, almost none of them hedged. They were, in the phrase that stuck, confidently wrong. An error rate is something you can plan around. Errors that arrive dressed as certainties defeat the reader's only defense, which is doubt.

What did the three big studies actually find?

Three independent research efforts, run by different institutions with different methods, converged on the same shape of result during 2025.

StudyScopeHeadline result
CJR / Tow Center8 AI search engines, 1,600 news-citation queriesOver 60 percent answered incorrectly
EBU / BBC4 assistants, 3,000+ responses, 22 broadcasters, 14 languages45 percent had at least one significant issue
NewsGuard10 leading chatbots, monthly audits over one year35 percent false-claim repetition, up from 18

Read the three rows together and the convergence is the story. Different institutions, different languages, different grading rubrics, and the failure rate still lands in the same unhappy band. The numbers underneath the EBU and NewsGuard rows are where the mechanism shows itself.

The EBU and BBC cast the widest net. Professional journalists from 22 public service media organizations across 18 countries evaluated more than 3,000 responses in 14 languages from ChatGPT, Copilot, Gemini, and Perplexity. 45 percent of answers contained at least one significant issue. Sourcing was the largest single failure category at 31 percent, and 20 percent of responses had major accuracy problems, including hallucinated detail presented as fact. Gemini fared worst, with significant issues in 76 percent of its responses. The consistency across languages and territories is the point. This is not an English-language quirk or a single-model defect. It is systemic behavior.

NewsGuard's monitor is the most unsettling of the three because it is longitudinal. Auditing the ten leading chatbots monthly for a year, it found the false-claim repetition rate on news topics nearly doubled, from 18 percent a year earlier to 35 percent by August 2025. A rising error rate on a fixed task, tracked the same way month after month, is not noise. It is a trend, and what drove it is the more damning part of the story.

More capability produces more confident errors

The longitudinal data reframes how buyers should think about "better" AI tools. The jump in wrong answers tracked a single behavioral change: over NewsGuard's audit year the models' refusal rate fell from 31 percent of prompts to essentially zero. More willingness to answer, not more accuracy, is what the added capability bought.

That pattern is the whole story. These systems are being tuned toward answering, not toward being right. An assistant that says "I could not verify this" loses the engagement contest to one that always has something to say. Every commercial incentive points at the second behavior, and none of the three studies found evidence that grounding improved to match the confidence.

Here is the one question this piece will pose: if the leading tools cannot reliably attribute a news article that sits verbatim in their retrieval index, what should we expect when the question has no clean published answer at all?

What does this mean for competitive research queries?

News citation is the easy version of the test. The article exists, the publisher is known, the correct answer is one lookup away. Evaluators in these studies could grade every response against ground truth in minutes.

Competitive questions, the kind operators actually feed these tools, fail all three of those conditions. Does this vendor support customer-managed encryption keys. Is their FedRAMP authorization current. What did their last packaging change do to renewal terms. The ground truth is scattered across vendor documentation, compliance registries, changelogs, and filings. Much of it is vendor-authored, which makes it a claim rather than evidence. And a good portion of the retrievable web content about any vendor was written by a competitor for the specific purpose of being retrieved.

A hypothetical makes it concrete. A procurement lead asks a chat assistant whether Sentrix, an imaginary security vendor, supports SCIM provisioning. The assistant synthesizes from whatever it retrieves: a three-year-old documentation page, a comparison table published by Sentrix's imaginary rival CrowdHaven, a forum thread about a beta feature. The answer comes back fluent, current-sounding, and unsourced. Given the behavior measured above, there is no basis for treating it as reliable, and no qualifier attached to warn anyone off.

The failure loop is longer, too. A wrong news citation gets caught the moment someone clicks the link. A wrong capability answer gets caught during implementation, or in a security review, or after the contract is signed. The cost of confident error scales with how far downstream the correction happens, and competitive research sits about as far downstream as it gets.

What a defensible answer looks like

Reversing that sequence is the entire design premise of OmniAxis. The studies document one reflex, answer first and source maybe; the platform is built to run it the other way around.

Every vendor capability claim in OmniAxis is scored against graded evidence, and that grade stays attached to the claim wherever it later gets quoted. OmniAxis keeps the evidence base current on a managed refresh schedule, weekly on the top plan, with on-demand refreshes on higher plans, and every profile carries its last-refresh date, so a grade reflects a recent read of the documented record and not a training snapshot. The marketing-claim score sits next to the evidence score on purpose, because the distance between what a vendor announces and what independent sources confirm is the intelligence a buyer is actually after.

A claim with nothing independent behind it does not get hidden or smoothed into a fluent paragraph. It is labelled as unsubstantiated, visibly, so the reader knows exactly how much weight it can bear.

The studies are worth reading in full before your next research cycle leans on a prompt. What they measured was news. What they exposed was a posture: systems that answer everything and flag nothing. Competitive decisions deserve the opposite. Evidence first, confidence last, and a grade in between that someone is accountable for.

Sources