Google Analytics has offered a checkbox to filter out known bots and spiders for years. Roughly half of all internet traffic is estimated to be non-human. Both things are true, and the gap between them is where the digital measurement industry now lives.

Performance marketers have described a common pattern: direct-response campaigns posting record click volumes while conversions remain flat. The dashboards look healthier than the bank account. Turning on every bot-filtering option barely moves the numbers.

That gap is the story of the decade in web measurement. It is also the reason the old Google Analytics promise — click a box, exclude the machines, trust the pageviews — reads today less like a feature and more like a period piece.

The original bot-filtering setting did something narrow and useful. It removed traffic from user agents on the IAB/ABC International Spiders and Bots List. That covered Googlebot, Bingbot, and the well-behaved crawlers that identify themselves honestly. It did nothing for the traffic that mattered: scrapers pretending to be Chrome on a MacBook, click farms routing through residential proxies, and, more recently, large language model agents fetching pages at machine speed while presenting as ordinary browsers.

The Wikimedia Foundation offered an unusually candid look at the scale of the problem. In a blog post summarized by the New York Post, Wikimedia reported that human pageviews on Wikipedia were down 8 percent year over year, and that Wikipedia’s bot detection systems concluded that much of the unusually high traffic in May and June 2025 had come from bots designed to evade detection. Wikipedia only found the human decline after it stripped out the machines it had previously counted as people.

Wikipedia has better detection than most publishers. It also has less to lose commercially from telling the truth about its own numbers.

The advertising economy is the part of the internet where the incentive to tell that truth is weakest. Business of Apps, tracking mobile ad fraud through 2025, describes an industry in which generative AI has lowered the cost of producing convincing fake traffic while raising the sophistication of the fakes. Synthetic device fingerprints, AI-generated user behavior patterns, and adversarial models trained specifically to defeat fraud filters are now standard tooling on the fraud side.

The measurement industry’s response has been to add more filters, more signals, more machine-learning classifiers. Each layer catches more of the previous generation of bots. Each layer creates a new target for the next generation.

Analytics leads at regional publishers have run experiments comparing raw server logs, Google Analytics 4 with default bot exclusions, and third-party invalid traffic tools. On some properties, the three systems disagree about total human sessions by a factor of nearly two. The middle number typically gets used for board presentations. Everyone does.

The Super Bowl gave the industry a rare public stress test. Mashable reported on analysis suggesting that the majority of traffic Elon Musk’s X received during the event may have been non-human. X disputed the framing. The underlying question — how a platform of that scale cannot produce an unambiguous number for how many humans watched a Super Bowl ad on its service — went largely unanswered.

The category of fake traffic is itself doing work here. It bundles together several distinct problems that require different responses. There is fraudulent traffic manufactured to steal ad spend. There is scraper traffic from companies training AI models on the open web. There is agentic traffic — LLMs fetching pages on behalf of a human user who asked a question. There is the ordinary background hum of security scanners, uptime monitors, and price-comparison bots.

Only the first is unambiguously bad. The others are, depending on your business model, a cost, a threat, or a customer. A publisher whose article is read by ChatGPT and summarized to a user has been read, in some meaningful sense. The user just never showed up in the analytics.

Wikipedia’s traffic decline sits inside that ambiguity. Wikimedia argued that people were still consuming Wikipedia’s content — through AI answer engines that are now used widely by AI-familiar consumers, through Google’s AI Overviews, through social video summaries. The audience has not shrunk. The measurable audience has.

For advertising-supported publishers, the distinction is fatal. DMG Media told the UK’s Competition and Markets Authority that AI Overviews had cut click-through rates to MailOnline by 89 percent. The readers still exist. The pageviews do not.

This is the environment Google’s bot filtering was designed for a decade ago, and this is why the original promise of the feature has quietly aged out of usefulness. The tool was built to exclude a small, well-behaved population of self-identifying crawlers from a measurement system whose main job was counting humans reading articles and clicking ads. The population is no longer small, no longer well-behaved, and no longer clearly separable from the humans on whose behalf many of the machines now act.

Google’s more recent moves have pushed in two directions at once. In GA4, invalid traffic detection for advertising has become more aggressive and less visible to the user. The choice to filter is less of a checkbox and more of a black-box classification. At the same time, Google’s own ad products have absorbed the reality that a growing share of the funnel involves AI intermediaries. The measurement layer has been redesigned around Google’s ability to see the whole journey, which is convenient for Google and less convenient for anyone trying to audit the numbers independently.

CFOs at B2B software companies have described the problem in plain terms during budget reviews: two dashboards showing two different pictures of the same quarter. One counts every session Google Analytics accepts as human. The other counts only sessions that eventually produce a signed contract. The ratio between them gets worse each year. The first dashboard gets ignored.

That is, in practice, what the sophisticated end of the industry has done. The people who buy media at scale increasingly ignore top-of-funnel traffic numbers and grade themselves on outcomes further down: verified leads, activated accounts, revenue. It is a rational response to a measurement environment where the top of the funnel has become uncountable. It is also a luxury available mainly to businesses whose product is expensive enough to make outcome-based measurement affordable.

For everyone else — the small publisher, the local retailer, the nonprofit trying to prove reach to a funder — the checkbox in Google Analytics is still the primary defense against a flood it was never built to hold back. The broader fake-traffic economy makes clear that fraudulent traffic is easier to buy than legitimate traffic and often indistinguishable from it in the buyer’s dashboard.

There is a version of this story in which better filters win. Detection improves, the industry cleans up, and the pageview reasserts itself as a trustworthy unit of account. There is another version, which the current evidence supports more strongly, in which the pageview quietly stops meaning what it used to mean, and the industry adjusts by measuring different things.

Google’s bot filtering, in either version, is doing what it has always done: removing the machines that admit to being machines. The uncomfortable part is that this now describes a shrinking share of the machines, and an ambiguous share of the traffic that anyone should have been counting as human in the first place.

The dashboard is not lying. It is answering an older question than the one being asked.