Tabla de contenidos
Yesterday we learned that Google has been quietly paying dozens of publishers for months when their content contributes significantly to its AI answers. This isn’t an isolated move, but the latest piece that defines a changing of the guard in the relationship between AI bots and our content. But let’s take it step by step.
Not all AI bots that access our website are after the same thing. And starting to tell them apart is becoming increasingly important for our business.
Until now, it’s been common to talk about AI bots as if they were a single category. Basically, most websites either let them in or block them, and that’s that.
But something has started to change.
Just a few days ago, Akamai announced it has updated its bot directory to differentiate AI traffic into three broad groups:
- Training bots: those that collect content to train AI models.
- Search bots: those that crawl and feed the platform’s own results index so that content can be cited and linked in its generative answers.
- AI agents: those that access a website in real time because a user has asked them to look up information, compare products, or carry out a task.
While the differences may seem like minor technical nuances, the consequences have a very clear impact. Let’s see.
We might, for example, not want some of our content (for instance, informational pages, FAQs, studies, analyses, etc.) to be used to train models, but still want it to appear as a source when someone asks ChatGPT or Perplexity about our products or services. And we probably don’t want to block an agent that a customer has asked to consult our catalog, look up a price, compare a solution, or perform an action on our website, either.
Akamai therefore enables much more granular control so website owners can decide when and how their content is accessible.
What’s most striking is the timing of Akamai’s decision alongside a series of changes that suggest the era in which LLMs accessed all online content without permission or control is coming to an end.
In recent months, a regulator has turned it into a right, major media groups are steering it toward their bottom line, the web’s own infrastructure has adopted it as a default setting, and even Google itself, as we saw at the start, has begun paying for the content that feeds its AI answers.
The regulator turns the distinction into a right
The Competition and Markets Authority (CMA), the UK competition regulator, issued on June 3 a binding order (the first of its kind in the world) requiring Google to offer publishers separate opt-out controls for three areas of its results: AI Overviews, conversational AI Mode, and fine-tuning of its AI models, such as Gemini. The real novelty of this order is that it requires the search engine to ensure that a publisher’s adoption of any of these opt-outs does not penalize its rankings in traditional search results.
Until then, Google forced an all-or-nothing choice: either your content fed all its AI features or, if you implemented the relevant block in robots.txt, you disappeared from the search engine. The CMA has put an end to this model, and control can now be exercised at the domain level or on individual pages. In other words, you can let your evergreen content section appear in AI Overviews while protecting your premium or subscription content from Gemini fine-tuning. In addition, Google must attribute—with clear links—the content it uses in its AI answers, and the order expressly prohibits Google from hiring crawlers from other companies to access, through the back door, the content of those who have chosen to block its access.
The UK regulator’s order states that Google has nine months to implement the changes and, for now, the features are being rolled out only to a subset of UK sites. But we can see a clear pattern: what Akamai treats as technical categories, the CMA enshrines as separate rights for the content owner. And where the UK has led the way with its Digital Markets, Competition and Consumers Act, it’s reasonable to expect other regulators, such as the European Commission, to take note.

Media groups use this distinction to boost their bottom line
Without waiting for regulators, major publishing groups and international publishers are already applying the same logic. The pattern repeats: on the one hand, we have media outlets signing licensing deals with certain AI companies and, at the same time, blocking access for those that don’t pay. LLM access to content stops being free and universal and becomes a negotiable asset.
The most sophisticated case is The Atlantic. As Digiday has reported, the magazine’s team tracks in a spreadsheet which crawlers access its site and cross-references that data with referral traffic and subscription conversions generated by each AI platform. Each week they review bot behavior and decide which ones get access and which don’t. That analysis led them, for example, to block a single crawler that had made 564,000 requests in seven days without giving anything back in return. In this way, The Atlantic blocks only the bots whose cost exceeds the value they deliver.
And they’re not alone. Reuters and Time have also moved to blocking AI bots by default, allowing access only to those approved via allowlists. Their teams say the change not only didn’t cost them traffic, but also significantly reduced their server spend—and that it has also helped encourage AI companies to come to the negotiating table. Selective blocking isn’t just a defensive measure; it becomes a commercial lever.
Google puts a price on the other side of the scale
And while publishers figure out how to value each bot, Google has started to value each piece of content. As Digiday reported yesterday, the company has been quietly testing the “AI contribution pilot”: a program built into Search Console that pays publishers when their content contributes “significantly” to AI-generated answers in Gemini, AI Overviews, and AI Mode. Publishers participating in this pilot program have an AI revenue widget in their Google Search Console dashboard showing a monthly figure. Payment doesn’t depend on crawl volume, but on the value contributed by the content. In other words, Google only pays when it judges that a piece has made a meaningful contribution to an answer.
While it’s a step in the right direction, participants themselves describe the calculation as a black box, the revenue level is negligible compared to their ad income, and some analysts interpret the program as a test run of the infrastructure Google would need if it were ever forced to pay for the content it uses. The program has also been more attractive to small and mid-sized publishers—who lack the muscle to negotiate their own licensing deal—than to the largest publishers.
But the conceptual shift is important: through this data, publishers will be able to see which content delivers value to AI systems in the same way Search Console shows us today which content performs in search. This establishes that some content does, in fact, deliver value—and there must be a path for that value to flow back to the people who created it. Exactly the same logic The Atlantic applies to bots, but in reverse.
Infrastructure, set by default
In case there were any doubts that this is the new standard and not a fad: starting today, September 15, new domains signing up to Cloudflare will have training crawlers and agents blocked by default on pages that serve ads, while search bots remain allowed. When companies as large as Akamai or Cloudflare—which underpin a large share of the web—adopt the taxonomy as the default server configuration, the debate is settled: granular management of AI bot access stops being a novelty and becomes the ground for negotiation.

From “something for everyone” to selective tolls
Five different players—an infrastructure provider, a regulator, major media groups, the world’s leading CDN, and Google itself—have converged in just a few months on the same idea: AI access to content has a value that must be measured, negotiated, and increasingly, paid for.
It’s time to change how we think: in the face of the growing zero-click phenomenon, the question we should be asking is no longer whether to allow or block AI bots, but what type of access we allow in each case—and in exchange for what value. Protecting our content and, at the same time, pursuing visibility in generative answers are no longer incompatible goals.
One last note: managing all this takes more than editing robots.txt. According to the Tollbit State of the Bots report, around 30% of AI bot crawls fail to comply with the explicit permissions declared in that file. Bot governance requires log monitoring, CDN or WAF rules, and a periodic review of the value each bot delivers in exchange for the server resources it consumes.
And as AI agents become increasingly widespread for search, comparison, or purchasing tasks, this distinction will undoubtedly become even more important: blocking the wrong agent will increasingly mean shutting the door on the potential customer who sent it.
Our recommendation? Start with the basics: audit which bots access your site today, measure what each one gives you back, and decide—bot by bot—who gets in, under what terms, and at what price.







