One of the things you learn running infrastructure is that percentages can lie.

Back in the dot-com days, I spent a good amount of time in the hosting industry. Reliability was the name of the game. And very early on, I learned that a number that sounds good in marketing can be terrible in the real world.

Take uptime guarantees.

Early hosting providers used to brag about 90% uptime. On paper, it sounded respectable. In reality, it was a joke, because 10% downtime translates into weeks of outages over the course of a year. No serious business could operate that way.

Eventually, the industry got religion about reliability. Four nines became the standard. Then five nines. Even then, you were still talking about a few minutes of downtime each month, and those minutes mattered when real systems were involved.

I learned the same lesson again with early text recognition software. Vendors proudly claimed their software was 90% accurate. Sounds great until you run a 1,000-word document through it and realize you now have roughly 100 mistakes to fix.

Suddenly, that 90% number feels a lot less impressive.

Which brings us to Google’s AI answers.

Google’s AI Overviews now appear at the top of search results, summarizing information and presenting it directly to users. Instead of sending people to a list of sources, Google increasingly gives them the answer.

Google itself does not claim the system is 90% accurate. That number comes from an independent analysis discussed recently in a New York Times report that evaluated the system using a common AI benchmark. Depending on the version of Google’s Gemini model being used, the answers were accurate roughly 85 to 91% of the time.

That sounds impressive. Until you think about scale.

Google processes more than five trillion searches every year. Even if we assume the optimistic end of that range and say the system is correct 90% of the time, that still means millions of answers are wrong.

Not per year. Not per month.

Per hour.

At Google scale, a 10% failure rate translates into thousands of incorrect answers being served up every minute. That is not a rounding error. That is a systemic issue.

The bigger problem is that the answers look authoritative. They sit at the very top of the page, polished and confident, often with a handful of citations that most people never bother to check.

For many users, the AI answer is the answer.

Which means Google has quietly changed its role on the internet.

For most of its history, Google was a search engine. It indexed information created by others and pointed users to the places where that information lived. Publishers, bloggers, journalists and researchers created the content. Google helped people find it.

AI Overviews change that dynamic.

Now Google reads those sources, synthesizes them and generates its own answer. In many cases, the user never clicks through to the original site.

In other words, Google has shifted from being a guide to the web to being one of the largest publishers on the internet, built largely on other people’s work.

There is another dimension to this shift that is worth acknowledging.

For years, Google was a major traffic engine for the open web. Publishers created content and Google helped users discover it. That relationship was never perfect, but it helped sustain a huge ecosystem of journalism, research and commentary.

AI answers interrupt that flow.

In the interest of full disclosure, this shift has real consequences for companies like ours. At Techstrong, a meaningful portion of the traffic we used to receive from Google simply doesn’t leave Google anymore.

So yes, I have a horse in this race.

But the issue here goes well beyond publishers protecting their turf. When the dominant gateway to information on the internet starts generating its own answers instead of pointing users to sources, the accuracy of those answers becomes critically important.

And that accuracy is more complicated than a simple percentage.

The New York Times analysis highlighted another issue with AI answers. Even when they are technically correct, many of them are what researchers call “ungrounded.” That means the sources cited by the system do not fully support the claims the AI is making.

In some cases, the citations themselves are questionable. Facebook posts and Reddit discussions frequently show up among the references.

That might be fine for casual browsing, but it is shaky ground for answers presented as definitive.

Then there is the gaming problem.

Researchers and journalists have already demonstrated how easy it can be to manipulate AI-generated answers. In one experiment, a journalist wrote a tongue-in-cheek blog post claiming to be one of the best tech journalists in the world at competitive hot dog eating.

Shortly afterward, Google surfaced that claim in an AI Overview as if it were an established fact.

That example is funny, but the underlying lesson is serious.

If someone wants to become known as an expert in something, they can simply create content declaring themselves one and wait for AI systems to ingest it. Marketers, propagandists and bad actors will inevitably figure out how to exploit that dynamic.

AI systems are probabilistic by nature. They generate the most statistically likely answer based on patterns in the data they have seen.

Most of the time, that works remarkably well.

But errors are not anomalies. They are part of the design.

At small scale, that might be acceptable. At Google scale, it becomes a flood of mistakes moving through the information ecosystem.

Which brings us back to that old hosting lesson.

A 10% failure rate might look fine in a marketing brochure. In a real system operating at massive scale, it is a disaster waiting to happen.

The real issue is not that AI systems make mistakes. Every technology does.

The issue is how those mistakes are presented.

Right now, Google’s AI answers look polished, confident and authoritative. The small disclaimer warning that AI may make mistakes is easy to overlook. The presentation implies a level of certainty the underlying technology simply cannot guarantee.

Users deserve a clearer understanding of what they are looking at.

AI answers should be framed as summaries and starting points, not the final truth. Sources should be easier to verify. And the companies deploying these systems should be far more transparent about their limitations.

Because at the scale Google operates, accuracy percentages are not just statistics.

They shape how billions of people understand the world.

Shimmy Says

If you are treating Google’s AI answers as gospel, you are making a mistake.

At 90% accuracy, the odds say you are already reading something that is wrong. And because the answer looks authoritative and sits at the top of the page, most people will never question it.

Google built the most powerful information engine the world has ever seen. But in the process, it quietly shifted from being a search engine to being a publisher, synthesizing other people’s work and presenting it as fact.

At the scale Google operates, even small error rates become enormous problems. Millions of searches. Millions of answers. Thousands of mistakes every minute.

That is not a rounding error. That is a systemic flaw.

Google should be far more transparent about what these systems can and cannot do. Until then, users should treat every AI answer with a healthy dose of skepticism and check the sources themselves.

Nine out of ten might be great when dentists recommend toothpaste.

For AI answers that billions of people rely on every day, it just doesn’t cut it.