
Our spy tools monitor millions of native ads from over 60+ countries and thousands of publishers.
Get StartedFor three decades, the robots.txt protocol operated on a single, generous assumption: every bot is welcome unless a publisher explicitly says otherwise. That default, established in 1994 as an informal gentleman's agreement among early web crawlers, quietly governed the relationship between content creators and automated agents for the entire lifespan of the commercial internet. Now, in a matter of months, some of the most recognizable names in news media have flipped that assumption on its head — and the numbers behind their decision reveal why there's no going back.
Reuters and Time both switched to default-deny configurations in May 2026, joining People Inc. and The Atlantic, which had adopted similar allowlist setups within the preceding year. Instead of maintaining an ever-growing blocklist of known bad actors, these publishers now block everything by default and approve only the crawlers that earn entry. The logic is straightforward but devastating to the old paradigm: you cannot block a bot you've never heard of.
The scale of the problem makes the point viscerally. When People Inc. shifted from a blocklist to an allowlist, the number of user agents it blocked leapt from roughly 2,100 to more than 30,000 — a figure its SVP of innovation, Lindsay Van Kirk, shared at an IAB Tech Lab event in late May. That means tens of thousands of crawlers were slipping through unrestricted, consuming server resources and harvesting content without permission, simply because no human had gotten around to naming each one individually in a blocklist. People Inc. now gives access to just 38 approved crawlers while blocking millions of daily unauthorized scrape attempts.
Compliance — or the lack of it — made this transition inevitable. A Tollbit report cited by Search Engine Journal found that 30% of total AI bot scrapes didn't comply with explicit robots.txt permissions. When nearly a third of automated traffic simply ignores the rules, the honor system ceases to function as a system at all. Publishers discovered they were engaged in an asymmetric game: meticulously updating text files that a meaningful share of crawlers never bothered to read.
The industry's institutional response confirms this isn't a temporary skirmish. IAB Tech Lab released new guidance on bot and crawler management strategies, opening the document for public comment through late June 2026. The guidance acknowledges that blanket bot blocking is no longer practical and instead outlines a spectrum of approaches with distinct operational, financial, and strategic trade-offs. It complements IAB Tech Lab's recently released CoMP API V1, a framework designed to standardize communication and permissions between AI systems and publishers — essentially building the plumbing for a permissioned web.
Meanwhile, collective leverage is growing. The SPUR Coalition, which is building shared standards for licensing and content use, expanded to 36 member organizations after adding 30 new publishers in a single month. As Search Engine Journal noted, one site blocking AI bots is easy to ignore — but thirty-six publishers coordinating a unified stance is a different negotiating proposition entirely.
Reuters head of professional products Josh London told Digiday that the publisher now approves a bot only when it offers a "fair value exchange" spanning four categories: licensing payments, referral traffic, site functionality, or monetization support. The live Reuters robots.txt file reflects this calculus precisely — named crawlers from Amazon, Google, Bing, Yahoo, and OpenAI get through, while everything else hits a wall. The freely crawlable editorial web, in other words, is being replaced by something that looks far more like a gated marketplace — and the gates are closing faster than most advertisers realize.
The numbers behind the AI crawler ecosystem are staggering, and they explain why publishers are panicking. According to internal data analyzed by Search Engine Journal, ChatGPT's user agent now accounts for more than 96% of all AI user bot traffic, hitting sites in real time on behalf of consumers asking questions, comparing products, and making purchasing decisions. But OpenAI's dominance is only part of the picture. GPTBot represents roughly 55% of AI training crawl volume, while OAI-SearchBot handles about 47% of AI search crawl activity — meaning OpenAI is the single largest presence across training, search, and user-facing layers simultaneously. And it has plenty of company.
The growth curves of other crawlers are equally jarring. ClaudeBot, Anthropic's crawler, surged 800% between November and December 2025 alone. ByteSpider, the crawler powering ByteDance's AI products, grew 138% over the same tracking window. Google's NotebookLM crawler climbed 144% as researchers and knowledge workers used it to pull live content, and Gemini user agents are now surging on top of that. Perhaps most overlooked of all, Applebot accounts for roughly 30% of AI search crawl activity — and most search and marketing teams aren't even tracking it.
This isn't a gradual evolution. AI agent activity is growing 150% month-over-month, AI agents now represent 15% of all website traffic, and agent activity is on course to overtake human-driven search before the end of 2026. The crawler landscape has become so chaotic that keeping a blocklist current is effectively impossible. When People Inc. switched from a blocklist to an allowlist approach, the number of user agents it blocked jumped from roughly 2,100 to more than 30,000. That gap — between what publishers thought they were managing and what they were actually exposed to — is what's driving the default-deny response.
Into this environment steps Microsoft with a very different message. Speaking at AdExchanger's Programmatic AI event, Microsoft's VP of publisher product Nikhil Kolar urged publishers and retailers to open their sites to AI crawling rather than fighting the tide. His argument was blunt: with four out of five websites currently blocking AI bots, those publishers' content and products are "not legible to agents," which means "your business is closed" and "you get no discovery, no recommendations that you're a part of and no demand." Microsoft also pointed to its Publisher Content Marketplace as a good-faith mechanism for licensing deals, drawing a distinction between using publisher data for "training" and using it for "grounding" AI models.
Both sides have legitimate positions, and that's precisely what makes this tension so productive for native advertisers who understand its implications. Publishers are right that the sheer proliferation and unpredictability of crawlers makes blanket openness untenable — AI training doesn't follow a linear schedule, and log analysis after the fact is no longer sufficient. Microsoft is right that blocking crawlers wholesale removes publishers from an entirely new discovery layer. But here's the key insight that neither side is talking about enough: every wall that goes up around premium publisher content degrades the quality and completeness of the general-purpose AI tools that marketers, analysts, and executives rely on for competitive intelligence. When ChatGPT, Gemini, or Perplexity can no longer crawl Reuters, Time, or The Atlantic, the synthesized answers they produce become less authoritative, less current, and less complete. The fragmentation isn't just a publisher problem or a platform problem. It's creating blind spots in the very tools that most marketing teams now treat as omniscient.
Here's the structural reality that most marketing teams are missing: the publisher lockdown targets editorial content crawlers, not ad-serving infrastructure. These are two fundamentally different architectural layers of the internet, and the wall going up around one has virtually no effect on the other.
When Reuters updates its robots.txt to whitelist only approved crawlers, or when People Inc. blocks between 30,000 and 35,000 different crawlers daily, they're protecting their editorial content — the articles, investigations, and proprietary journalism that represent their core intellectual property. But native advertising doesn't live in that layer. Native ad creatives, sponsored content placements, and programmatic inventory flow through demand-side platforms, ad networks, and purpose-built distribution systems that operate on entirely separate infrastructure. These systems don't depend on open-web crawling to function. They never did.
Consider how a native ad actually reaches a consumer. A brand creates an ad unit. That unit enters a programmatic ecosystem — a demand-side platform bids on inventory, the ad server delivers the creative, tracking pixels fire, and the entire lifecycle is logged in proprietary databases maintained by the ad network. None of this activity is governed by robots.txt. None of it appears in the editorial CMS that publishers are fortifying. The creative itself, the targeting parameters, the performance data, and the placement history all exist in a parallel data ecosystem that competitive intelligence platforms can monitor directly through ad libraries, network-level APIs, and their own crawling infrastructure that interfaces with ad exchanges rather than editorial pages.
This parallel ecosystem is enormous and growing fast. As MarTech reported, U.S. businesses are expected to spend $57 billion on AI-powered advertising this year, accounting for roughly 12% of total ad spending. That investment is pouring into systems designed for continuous creative optimization, autonomous media buying, and conversational discovery — all of which generate rich competitive data that flows through channels completely unaffected by the publisher lockdown. Purpose-built ad spy tools that monitor these networks will continue harvesting creative variations, tracking placement strategies, and benchmarking performance metrics while general-purpose AI tools lose visibility into the editorial content surrounding those placements.
The irony is that Microsoft's own VP of publisher product has urged publishers not to block AI bots, warning that doing so means "your business is closed" and "you get no discovery, no recommendations that you're a part of and no demand." But even as four out of five websites now block AI crawlers according to Microsoft's data, the advertising infrastructure those same publishers rely on for revenue remains wide open by design. Publishers can't wall off their ad-serving systems without simultaneously destroying their own monetization — the ad exchanges, supply-side platforms, and header bidding wrappers all require open programmatic access to function.
This creates an asymmetry that native advertisers can exploit. The editorial layer is going dark to AI, but the advertising layer remains fully illuminated. Competitive intelligence platforms that tap directly into ad networks, scrape Meta's Ad Library, monitor Google's Ads Transparency Center, and track programmatic placements through their own infrastructure will maintain — and in relative terms, actually strengthen — their visibility advantage. The brands running native campaigns through these systems aren't losing data; they're operating in the one part of the open web that the crawler wars can't touch. While competitors relying on general-purpose AI summarization of editorial content find themselves increasingly blind, advertisers working within the programmatic ecosystem retain full-spectrum visibility into what's working, what's spending, and where the opportunities are emerging.
Ask ChatGPT what native ads your competitor is running on a major news publisher, and you'll get something that looks like an answer. It might reference a campaign from six months ago, extrapolate from a blog post that mentioned the brand in passing, or simply fabricate a plausible-sounding media placement that never existed. The response will arrive with the same confident tone regardless of whether it's accurate, outdated, or entirely hallucinated. And the foundation beneath that answer is eroding faster than most marketers realize.
The reason is structural. As Reuters and Time now default to blocking AI bots — allowing only approved crawlers through strict allowlists — general-purpose AI systems like ChatGPT and Perplexity are losing access to the very publisher environments where native advertising lives. People Inc. discovered that switching from a blocklist to an allowlist increased the number of user agents it blocked from roughly 2,100 to more than 30,000. That's not a marginal reduction in crawl access; it's a near-total lockout of unapproved bots. When the underlying training and retrieval data grows stale or disappears entirely, every competitive intelligence query you run through a general-purpose AI tool becomes less reliable — even as the interface makes it feel more authoritative.
Reuters applies what its head of Reuters Professional called a "fair value exchange" test to every crawler it evaluates: does this bot pay for content through licensing, send traffic back, keep the site running, or support monetization? General-purpose AI crawlers extracting editorial content to power chat responses typically fail all four criteria. But specialized ad intelligence platforms operate on fundamentally different terms. They either connect directly to ad network APIs under partnership agreements, maintain negotiated crawler relationships with publishers, or monitor publicly served ad inventory — the very inventory that publishers want to be visible, because visibility is the entire point of running an advertisement.
This distinction matters enormously. The IAB Tech Lab's new guidance on bot and crawler management explicitly acknowledges that blanket bot blocking is no longer practical because AI systems are deeply embedded across the web ecosystem. Instead, it outlines a permissioning framework through its CoMP API, designed to support structured communication between AI systems and publishers about access terms. Specialized ad intelligence platforms are precisely the kind of tools built to operate within these permissioning structures — they have clear commercial relationships, defined use cases, and they don't extract editorial value without compensation.
The gap this creates is widening in real time. Every publisher that shifts to a default-deny posture degrades the data available to general-purpose AI tools while leaving specialized platforms largely unaffected. A dedicated competitive intelligence tool pulling from Taboola's or Outbrain's API, indexing creatives served through programmatic native placements, or maintaining curated databases of active campaigns will continue to see the full picture. ChatGPT will see whatever it last managed to scrape before the door closed.
Marketers who've grown comfortable asking a chatbot about their competitive landscape are building strategy on an increasingly unreliable foundation. The convenience is real, but the data quality is decaying beneath it. Meanwhile, teams investing in purpose-built ad intelligence are quietly accumulating an advantage that compounds with every robots.txt update — because the more publishers lock down, the wider the informational moat around platforms that were built to operate within the rules rather than around them.
The numbers reveal a market sleepwalking into a structural disadvantage. When 81% of organizations are still filing AI agents under the same bucket as legacy bots, applying access rules designed for an entirely different era of the web, they're telling you something important: most brands haven't thought through what happens downstream when the information supply chain they've come to depend on starts delivering incomplete, stale, or outright fabricated intelligence. And when 77% of the brands that do have an agent policy only block training crawlers — a publisher move, not a brand move — the gap between those who understand the new landscape and those who don't becomes a genuine competitive moat.
That moat isn't built from any single technology. It's built from the combination of specialized ad intelligence tools, direct data access, and human analysts who can synthesize what automated systems cannot. Consider what's at stake: leading advertisers are now deploying continuous creative optimization loops in which AI evaluates engagement signals and automatically evolves messaging to improve performance. These loops depend on two inputs — first-party performance data and competitive context. The first input is largely under a brand's control. The second is where the crawler wars create an asymmetry that most marketers haven't yet recognized.
Speed of creative iteration — testing hundreds of variations, responding to competitive moves, capitalizing on cultural moments — is only as good as the quality of the competitive signal feeding it. If your optimization loop is informed by a general-purpose AI model that last indexed a competitor's native ad placements months ago, or that hallucinates campaign details because it never had access to the ad-serving layer in the first place, your "continuous optimization" is optimizing against a phantom. You're shadowboxing while your competitor is studying your actual punches in real time through ad libraries, network-level monitoring tools, and direct API access to placement data that was never subject to robots.txt restrictions.
This is where the hybrid approach — specialized tools plus human curation — becomes decisive. Crawler-independent intelligence pipelines pull from sources that general-purpose models simply cannot reach: publisher ad libraries, programmatic bid-stream data, social platform ad transparency tools, and proprietary monitoring networks that observe ad delivery at the rendering layer rather than the content layer. Human analysts then contextualize that raw data, identifying patterns in competitor positioning, spend shifts, and creative strategy that no automated system can reliably interpret on its own. As AdExchanger has documented, industry experts broadly agree that AI functions best as an assistant, and that humans shouldn't just be in the loop but should be in the lead.
The brands building this infrastructure today aren't doing it because it's trendy. They're doing it because the open web's fragmentation is making general-purpose AI an increasingly unreliable narrator of competitive reality. Every publisher that tightens crawler access, every platform that restricts data sharing, every new AI agent that gets treated like a decade-old bot instead of a consumer proxy — each of these developments degrades the general-purpose model's view of the advertising landscape while leaving specialized, crawler-independent pipelines completely unaffected.
The moat, then, is straightforward but difficult to replicate: invest in data sources that don't depend on the open crawlable web, pair them with human analysts who can extract strategic insight from raw placement data, and feed that intelligence directly into your creative optimization loops. The brands that do this will operate with a sharper, more complete picture of competitor activity than those waiting for ChatGPT to summarize what it can no longer see.
Receive top converting landing pages in your inbox every week from us.
Guide
AI has enabled advertisers to generate creative at a scale that manual competitor research can no longer track effectively. Instead of reviewing individual ads, marketers need to analyze the full creative landscape, filter out short-lived tests, identify long-running survivors, and map the patterns behind winning hooks, visuals, and offers. This article presents a competitive intelligence framework for turning thousands of AI-generated ad variants into actionable strategic insights.
Dan Smith
7 minAug 18, 2026
Guide
Google has made AI-driven creative generation, testing, and optimization the new baseline inside its advertising ecosystem, while many native advertisers still rely on manual workflows and guesswork. This article explains how independent marketers can use ad spy tools, AI creative tools, and structured testing to build a similar observe → extract → generate → test → scale optimization loop without giving up control to a closed platform.
David Kim
7 minAug 18, 2026
Guide
Google’s AI Overviews are reducing the clicks publishers and marketers receive from organic search, weakening the economics of SEO-dependent acquisition. The article argues that this shift creates a stronger case for native advertising, where marketers can buy distribution directly, control creative and landing pages, and reduce dependence on Google’s changing ecosystem. It also presents a five-step playbook for measuring SEO exposure, researching proven native campaigns, testing native content, and building a more diversified distribution strategy.
Elena Morales
7 minAug 16, 2026



