Summary

  • Cloudflare's Pay Per Crawl design records a billing event when an authenticated crawler offers payment and receives a successful response carrying a charge header. That is a strong access receipt, not proof of citation, training or commercial use.
  • Cloudflare now says crawling is a crude unit of value. Its Pay Per Use experiments move the trigger downstream: Ceramic.ai pays when publisher content appears in a result, while You.com lets an agent buy a specific premium item on demand.
  • Bot Preference Sync, announced on 21 August, aligns a site's category-level Search, Agent and Training settings with robots.txt. It reduces policy drift, but the last mile of purpose and use still depends on crawler transparency and downstream evidence.
  • Monetization Gateway remains an early-access plan. Cloudflare has not disclosed general availability, payment volume, publisher proceeds, product revenue or a fee schedule, so company growth cannot be counted as conversion of this market.

The delivery event is precise

The cleanest way to understand Cloudflare's emerging content market is to start with the event it can observe without guessing.

In the Pay Per Crawl design, an authenticated crawler requests a resource. It may discover a price through an HTTP 402 Payment Required response and retry with an exact-price instruction, or arrive with a maximum price it is willing to pay. If identity, policy and price checks pass, the edge returns the content with a successful response and a crawler-charged header. Cloudflare records the billing event, aggregates the charges and distributes earnings to the publisher.

That sequence has identifiable parties, a price, a request, a delivery decision and a settlement record. It is materially stronger than a line in robots.txt that asks a cooperative crawler to stay away. It is also narrower than the claim often made for it. A successful request says the buyer received the resource under an access rule. It does not say whether the resource was cited in one answer, used in a thousand answers, retained for retrieval, incorporated into training or never used at all.

Cloudflare says exactly that. In July, while describing the move from Pay Per Crawl to Pay Per Use, it called crawling a crude measure of value. One page may be fetched once and cited thousands of times; another may be fetched repeatedly and contribute nothing. Delivery and use are not merely two names for the same transaction. They have different clocks, evidence and pricing units.

Preference, enforcement and identity form a chain

The access ledger itself has several links. A publisher first states a preference. A network rule then enforces it. A crawler must be identified. Its declared purpose must fit the rule. Only then can an access or payment event mean what its label suggests.

Cloudflare's 21 August announcement of Bot Preference Sync addresses the first two links. It reflects zone-level choices for Search, Agent and Training traffic into robots.txt, while preserving existing disallow directives. Cloudflare is unusually explicit about the problem: a published preference and an enforced rule can disagree. A crawler may treat that inconsistency as permission to ignore the file or to probe around the block.

Synchronisation improves the record. It does not abolish the difference between the record and the control. Cloudflare says the feature works at category level, uses the bots it tracks in BotBase and does not directly absorb bespoke per-bot rules. An operator with a bilateral licence or a special exception may still need a custom configuration. The published file can be internally consistent while the economic deal lives elsewhere.

Identity is the next boundary. Pay Per Crawl relies on Web Bot Auth and signed requests to prevent one caller from impersonating another. Cloudflare's newer transparency policy asks mixed Search-and-Training operators for more: respect a no-training preference, offer an opt-out from AI summaries, provide URL-level visibility into training availability and search performance, and show that declining training does not impair ordinary search. A label becomes more credible when the operator accepts verifiable obligations around it. It still remains a statement about purpose, not a complete observation of everything a model later does.

A scheduled default is leverage, not settlement

Cloudflare's July “Your Content, Your Rules” announcement set 15 September as the deadline for new classifications and defaults. The July version said new sites would allow Search but block Training and Agent activity on ad-supported pages, while mixed crawlers that failed to separate purposes would be blocked on those pages. Existing free customers who had not chosen settings were also in scope.

The August product post is more granular. It describes a publisher onboarding choice that sets Training to Disallow; non-publisher domains start without a block; existing customers using the legacy managed file will be asked to review their preferences as the new sync launches. The evolution is evidence that the control is still being refined. The market test is what is actually active after 15 September, how many owners override it and whether crawler operators separate identities and purposes in response.

Defaults matter because Cloudflare sits at a distribution point. In its own one-year bot report, the company says 52% of crawler requests on its measured network were for training in June 2026, while mixed-use crawlers represented more than 36% of activity. It also says more than 20% of the web sits behind its network. Those figures are Cloudflare's measurements, not a global census, but they explain why a default at this edge can change bargaining conditions.

It creates scarcity by withholding delivery or making purpose separation a condition of access. It does not itself create a licence, allocate the value of an answer or move money to a publisher. Leverage and settlement are different receipts.

The second ledger starts after delivery

Cloudflare's experiments with Ceramic.ai and You.com shift the trigger. In the Ceramic description, an opted-in publisher is paid when its content appears in a search result. The proposed reporting includes the query, page, snippet and result position. In the You.com example, an agent can pay for a particular premium item when it needs it.

These are closer to a use ledger because they attach payment to a result-bearing event or a deliberate purchase. They also introduce harder questions. What contribution should be assigned when an answer synthesises several pages? Does a citation prove causal value? How are near-duplicate sources handled? Who audits a missing attribution? Is a query, a result, a token, a successful task or a licence period the right unit?

Cloudflare does not present one universal answer. It names Pay per Query, Pay per Result and other models as experiments. That restraint is economically important. A request is easy to count, but may be poorly correlated with value. An outcome can be closer to value, but is more contestable and more dependent on the buyer's internal record.

Cloudflare's Attribution Business Insights gives publishers a better negotiating ledger: successful bot access, crawl-to-referral ratios, bandwidth, operator and behavioural category. It reduces information asymmetry. Network observation still ends at a boundary. It can show that a crawler obtained a page and that a service later sent a referral. It cannot, by itself, reconstruct every training corpus, latent model influence or answer-generation path.

x402 widens what can be sold, not what has been proved

The planned Monetization Gateway would extend charging beyond crawlers to pages, datasets, APIs and MCP tool calls. Cloudflare describes an x402 exchange in which a caller receives price instructions, pays in stablecoins, retries with proof and reaches the origin only after verification. Rules could vary by route, method or task. The control is meant to sit close to the buyer on Cloudflare's network and be managed through the dashboard, API or Terraform.

This could lower the fixed cost of selling a very small digital unit. A software agent would not need a prior account or monthly subscription just to buy one useful call. For the seller, edge enforcement could keep unpaid demand away from the origin. For Cloudflare, identity, policy, payment verification and delivery would meet at one control plane.

But the verbs in the announcement are future verbs. The product has an early-access waitlist. The source does not disclose supported settlement assets and jurisdictions in operational detail, dispute or refund rules, transaction volume, seller proceeds, Cloudflare's economics or a general-availability date. A target of sub-second settlement is a design objective, not a service record.

The financial boundary is equally clear. Cloudflare reported US$696.1 million of Q2 revenue, up 36%, and US$56.4 million of free cash flow. Its Form 10-Q says subscription and support represented substantially all revenue; the usage-based consideration it identifies is primarily excess bandwidth. It does not separate revenue, customers or remaining performance obligations for Monetization Gateway, Pay Per Crawl or Pay Per Use. Strong company results are meaningful counterevidence to a distress narrative. They are not a receipt for a market that remains partly beta, experimental and planned.

What closes the two books

The market becomes measurable when the access book reconciles to money and the use book reconciles to attributable outcomes.

On the first side, publishers need successful paid requests, net proceeds, settlement time, failed collections, overrides and disputes. On the second, they need query- or task-level evidence showing which resource contributed, under what licence, with what attribution and payment. Buyers need a reason to trust that prices reflect useful material rather than duplicated or stale pages. Both sides need to know whether a paid access grants only delivery or also a defined right to retain, train, quote or redistribute.

HTTP can carry a price and proof of payment. It cannot supply the missing contractual scope by implication. A 200 response after a 402 exchange is not automatically a copyright licence, a training consent or proof that value was created. Cloudflare is building an increasingly credible access ledger. Its most interesting admission is that the use ledger has to be built separately.

Sources