← Writing

Don't block the bots. Give them an MCP.

Scrapers and site owners both waste resources fighting each other. A public MCP endpoint could be the front door they both need.

Contents 6 sections

Half the web is now busy blocking bots. The other half is busy building bots that get around the blocks. Both sides are burning money on the same fight.

Since July 2025, Cloudflare blocks AI crawlers by default. Publishers stack robots.txt rules, WAF rules and challenge pages. Crawlers answer with rotating IPs and headless browsers that render your full page, JavaScript and all, just to extract three paragraphs of text.

That is absurd. The agent doesn’t want your HTML. You don’t want to serve it your HTML. So why not give it a proper door?

The idea: redirect, don’t reject

Instead of answering a bot with a 403, answer with a pointer: “There’s an MCP server here. Use that.”

An MCP (Model Context Protocol) server exposes tools like search, get_article or list_products. The agent asks a structured question and gets a structured answer. No rendering, no parsing, no guessing which <div> holds the content.

Both sides win:

  • The site serves a small JSON response instead of a full page render. It controls what’s exposed, can rate-limit per tool, and can charge for it.
  • The agent gets clean data with fewer tokens, and doesn’t have to fight a CAPTCHA.

Is anyone doing this already?

Yes. More than I expected.

Shopify is the clearest example. Every Shopify store has a public MCP endpoint. Product search, cart, store policies. Zero setup for the merchant. Millions of sites, agent-ready overnight. Catalog and cart tools now live at /api/ucp/mcp and follow the Universal Commerce Protocol. The old cart tools on /api/mcp were deprecated on August 31, 2026.

Cloudflare is building the full stack. AI Index (private beta since September 2025) creates a search index for your domain and exposes it as an MCP server, a search API and llms.txt. AI platforms can even subscribe to updates instead of re-crawling. In February 2026 they added Markdown for Agents: send Accept: text/markdown and you get markdown instead of HTML. In July 2026 the Monetization Gateway followed, which lets you charge per MCP tool call via HTTP 402 and x402.

Microsoft’s NLWeb turns a site into a conversational endpoint. Every NLWeb instance serves /ask and /mcp. Cloudflare supports it out of the box.

So the “give bots a better door” part exists. The missing piece is the signpost.

The hard part: how does an agent find the door?

This is where it gets messy. There is no single standard yet. There are several candidates:

  • /.well-known MCP Server Cards (SEP-2127, formerly SEP-1649). A JSON file describing the server: name, transports, endpoint, capabilities. Proposed in January 2026, still in review.
  • llms.txt. A markdown file at the root that tells LLMs what’s on the site. Widely adopted, but it’s a map, not an API.
  • WebMCP. A browser API (document.modelContext) where a page registers tools for in-browser agents. Shipping experimentally in Chrome, and Cloudflare can inject it at the edge.
  • The MCP Registry. A central catalog of servers. Useful, but it’s opt-in and centralised.
  • A2A Agent Cards. Same idea, different protocol: /.well-known/agent-card.json for agent-to-agent.

What I’d like to see is the simplest possible glue:

HTTP/1.1 403 Forbidden
Link: </.well-known/mcp/server-card.json>; rel="mcp"
Content-Type: text/plain

Crawling is not allowed. Structured access is available via MCP.

A block response that doubles as a redirect. Or a robots.txt line next to Sitemap:. Nobody has standardised this yet. As far as I can find, the research literature hasn’t either. Papers measure how much blocking is going on. Few ask what should replace it.

The part nobody wants to talk about: trust

An open MCP endpoint doesn’t fix the incentive problem. A scraper that ignores robots.txt will also ignore your MCP pointer if it can still get the HTML.

That’s why the other building blocks matter:

  • Web Bot Auth (IETF draft, built on HTTP Message Signatures). Agents sign their requests, so you know who’s asking.
  • AI Preferences (IETF aipref) and Cloudflare’s Content Signals. Declare what the content may be used for: search, AI input, training.
  • Pay per crawl / x402. Put a price on access.

Combine them and you get a sensible policy: unsigned bots get blocked, signed agents get pointed to the MCP, and the MCP decides what’s free and what costs money.

Where publishing for agents stands today

Here’s what a site can publish for agents as of September 30, 2026:

MethodWhere it livesStatus
robots.txt AI rules/robots.txtUniversal, but mostly written for search crawlers
Content SignalsContent-Signal: in robots.txtCloudflare’s own, around 4% of top sites
AI PreferencesHTTP header and robots.txtIETF drafts: vocab-08, attach-05
llms.txt/llms.txtCommunity convention, no formal spec body
Markdown negotiationAccept: text/markdownCloudflare Pro and up, around 4% of top sites
Link headersLink: on every response (RFC 8288)Plain HTTP, works everywhere
API Catalog/.well-known/api-catalog (RFC 9727)Published RFC, barely deployed
MCP Server Card/.well-known/mcp/server-card.jsonSEP-2127 still a draft
Agent Skills/.well-known/agent-skills/index.jsonEarly, a handful of sites
A2A Agent Card/.well-known/agent-card.jsonPart of the A2A spec
WebMCPdocument.modelContext in the pageChrome origin trial since Chrome 149
MCP Registryregistry.modelcontextprotocol.ioLive, API at v0.1
Web Bot AuthSigned request headersIETF WG draft -00 (September 2026)
x402HTTP 402 + payment headersLive on Cloudflare’s Monetization Gateway

A few things stand out.

The signpost is taking shape. It’s not a magic rel="mcp". It’s a Link header on every response pointing to your sitemap, llms.txt, API catalog and server card, plus the files themselves under /.well-known. Cloudflare’s Agent Readiness scanner (April 2026) grades sites on exactly that list. Run isitagentready.com against your own domain and you get a checklist.

Almost nobody does it yet. When Cloudflare launched the scanner, it looked at the 200,000 most visited domains. Content Signals: 4%. Markdown negotiation: 3.9%. MCP Server Cards and API catalogs: fewer than 15 sites in the whole dataset. Early movers get noticed.

The server card path is still moving. Deployed cards, and the scanner, use /.well-known/mcp/server-card.json. The SEP draft has changed its mind on the exact location more than once, and there is a parallel AI Catalog effort (/.well-known/ai-catalog.json) that wants to sit on top. Publish at the de facto path and expect to move it.

MCP itself got easier to host. The 2026-07-28 spec made MCP stateless: no initialize handshake, no session header. Every request stands on its own, which fits a public, cacheable endpoint on the edge. It also added server/discover, but that tells a client what a server supports after it has found it. Finding it is still the server card’s job.

WebMCP is the browser-side door. Chrome runs an origin trial, with an imperative JavaScript API and a declarative one that annotates plain HTML forms. Cloudflare’s edge injection went into developer preview in August 2026 and can proxy calls to your existing /mcp endpoint. Useful for agents that drive a browser. It doesn’t help a headless crawler.

What I’d do today

If you run a content site:

  1. Publish an llms.txt. It costs you ten minutes.
  2. If you’re on Cloudflare, turn on Markdown for Agents.
  3. Add Content Signals to your robots.txt so the rules say what the content may be used for, not just who may crawl it.
  4. If you have search or structured data, expose it through a small MCP server and publish a server card at /.well-known/mcp/server-card.json.
  5. Add Link headers that point to all of the above, then check the result on isitagentready.com.

Blocking is a wall. MCP is a door with a lock you control. I know which one I’d rather maintain.


Updated September 30, 2026: added an overview of where publishing for agents stands today, and updated Shopify’s endpoints after the move to UCP.