A new analysis found Google’s AI Overviews answered factual questions correctly about 90% of the time, yet more than half of those correct responses weren’t fully supported by the sources Google cited.

Google’s AI-powered search is getting better at answering questions. It may still be sending millions of users the wrong information every day.

An analysis commissioned by The New York Times and conducted by AI startup Oumi found Google’s AI Overviews achieved 91% accuracy on factual questions after Google’s upgrade from Gemini 2 to Gemini 3. The research found that 56% of the correct answers lacked full support from the sources Google cited.

That sounds like a strong result until you look at Google’s scale. With more than 5 trillion searches processed each year, a 9% error rate could translate into tens of millions of inaccurate AI-generated answers every hour.

“A recent analysis of AI Overviews found that they were accurate approximately nine out of 10 times. But with Google processing more than five trillion searches a year, this means that it provides tens of millions of erroneous answers every hour (or hundreds of thousands of inaccuracies every minute),” the New York Times reported, citing an analysis done by an A.I. startup Oumi.

The findings point to a larger issue than simple accuracy. Oumi found that 56% of the AI Overviews answers that were technically correct were “ungrounded,” meaning Google’s cited sources did not fully support the claims made in the AI-generated response. That makes it harder for users to confirm whether an answer is backed by reliable evidence.

The research raises fresh questions about how people should use AI-generated search results as Google continues turning Search into an AI-first experience.

Oumi evaluated 4,326 Google searches using SimpleQA, an industry benchmark widely used to measure factual accuracy in large language models. The company tested AI Overviews twice: first in October, when the feature relied on Gemini 2 for more complex questions, and again in February, after Google upgraded the system to Gemini 3.

56% of Correct Answers Generated by Google AI Overviews Lack Supporting Sources Despite 90% Accuracy

Accuracy improved from 85% to 91% across the benchmark. Source grounding moved in the opposite direction. In October, 37% of correct responses lacked sufficient support from the cited sources. By February, that figure had climbed to 56%.

“As Google has improved its A.I. technologies, its A.I.-generated answers have become more accurate. In October, AI Overviews were inaccurate 15 percent of the time,” the New York Times noted, citing Oumi’s analysis.

“But with Gemini 3, Google’s A.I.-generated answers were more likely to be ungrounded than when the system was based on Gemini 2, meaning the websites they linked to did not completely support the information they provided. In October, correct answers were ungrounded 37 percent of the time. In February, with Gemini 3, that figure rose to 56 percent,” the report added.

That distinction matters. A response can be factually correct yet still leave readers without evidence that confirms the claim. In those cases, Google may cite pages that contain incomplete information, conflicting information, or no direct support for the answer at all.

One example involved a search asking when Bob Marley’s home became a museum. Google AI Overviews answered that it opened in 1987. Historical records show the museum opened on May 11, 1986. Google cited a Facebook post, a travel blog, and a Wikipedia page that itself contained conflicting dates.

Another query asked when renowned cellist Yo-Yo Ma was inducted into the Classical Music Hall of Fame. Google’s AI Overview linked to the organization’s website, which lists him as an inductee, yet the AI-generated response claimed there was no record of his induction.

Researchers found similar inconsistencies across thousands of cited sources. Facebook ranked as the second most-cited source in the analysis, and Reddit ranked fourth. Incorrect AI Overviews cited Facebook slightly more often than correct ones.

Google challenged the findings, saying the study relied on a flawed benchmark and did not reflect how people actually use Google Search.

“This study has serious holes,” Google spokesperson Ned Adriance said in a statement. “It doesn’t reflect what people are actually searching on Google.”

Google has long acknowledged that AI-generated responses can contain mistakes. A disclaimer displayed beneath AI Overviews tells users, “A.I. can make mistakes, so double-check responses.”

The company said AI Overviews combine Gemini with Google’s search ranking systems and safety protections, which it says improve answer quality compared with the standalone model.

The report arrives as Google pushes AI deeper into its flagship search product. AI Overviews now appear for a growing share of queries, shifting Google from a search engine that primarily links users to information to one that increasingly generates answers itself.

For publishers, researchers, and everyday users, the findings point to a challenge that goes beyond hallucinations. Accuracy is improving, yet trust still depends on whether readers can trace an AI-generated claim back to evidence that actually supports it. At internet scale, that gap has consequences that extend far beyond a single incorrect answer.