← Writing

Web search vs. MCP: what it costs an AI to answer a Drupal question

I asked Claude 106 Drupal questions four ways. The MCP server matched web search on accuracy, at 3.5× the speed and a fifth of the cost.

Contents 8 sections

I asked Claude 106 Drupal questions four ways. The MCP server was as accurate as web search. It was also 3.5× faster and 5× cheaper.

Here’s how I measured it, and where it still falls short.

The problem

Ask an assistant “is this module ready for Drupal 11?” and it has two options.

It can answer from memory. That memory is months old, and release data changes daily.

Or it can search the web. That means several round trips, fetching drupal.org pages full of markup, and digging the answer out of HTML meant for humans.

drupalreleases.com already has this data in structured form. So I gave it an MCP server (more on that here). MCP, the Model Context Protocol, lets an assistant call a tool like get-project and get a clean answer back. No scraping.

I believed that would beat web search. I wanted to know by how much.

How I tested it

One model, Claude Sonnet. Four setups:

SetupWhat the model can use
MemoryTraining data only
WebWebSearch + WebFetch
MCPOnly the drupalreleases.com MCP server
MCP + webBoth, to see which it prefers

The questions cover 20 projects in three popularity bands. Popular ones like webform, pathauto and paragraphs. Mid-range ones like telephone and config_inspector. And long-tail modules like crawl_control and masquerade_as_role.

Every project got the same five questions:

  • What’s the latest stable version, and when was it released?
  • Is it Drupal 11 compatible?
  • Is it covered by security advisories?
  • What’s its maintenance status?
  • Was anything released in the last 30 days?

On top of that, two scenarios. An upgrade audit: “which modules in this composer.json block Drupal 11?” And a discovery question: “I need to schedule content publishing, what should I use?”

That’s 106 questions, 424 answers, run on 26 September 2026.

Three rules kept it fair:

  1. Clean runs. Every answer is a fresh claude -p in an empty directory. No settings, no CLAUDE.md, no other MCP servers.
  2. Abstaining is allowed. A JSON schema forces structured output, and every field can be null. “I don’t know” isn’t counted as wrong.
  3. Grading uses drupal.org, not my data. The correct answers come from drupal.org’s release-history XML. My server gets no home advantage.

The numbers

How often each setup got it right:

SetupCorrectConfidently wrongAbstained
Memory33%2%41%
Web88%5%0%
MCP90%7%0%
MCP + web92%7%0%

And what each answer cost:

SetupMedian timeCostInput tokensTool calls
Memory5.9s$0.0091,3090
Web20.4s$0.07211,1084.1
MCP5.9s$0.0135,7301.3
MCP + web6.7s$0.0208,8801.7

All 106 questions cost $7.63 with web search. With MCP: $1.42.

MCP is exactly as fast as answering from memory, 5.9 seconds median. Except it’s right.

Split by question type:

TypeMemoryWebMCPMCP + web
Latest stable release15%85%90%90%
Drupal 11 compatible70%100%95%100%
Security coverage65%100%90%100%
Maintenance status0%50%75%75%
Released in last 30 days0%100%100%100%
Module discovery100%100%100%100%

And by popularity:

BandMemoryWebMCPMCP + web
Popular40%97%97%94%
Mid34%86%89%94%
Long tail13%77%83%90%

Memory and web search both get worse as modules get more obscure. Those are exactly the modules you need to check before an upgrade.

What the numbers say

Memory knows it doesn’t know

Memory abstained on 41% of questions. It was fine on things that don’t change: every discovery question right, 70% on Drupal 11 compatibility.

Versions, dates and maintenance status were guesswork. 0% on maintenance status. 13% overall on long-tail modules. Useless for release questions.

Web search is right, eventually

Web search made four tool calls per question on average and read about twice the tokens MCP did. It was usually right. But the time adds up:

TypeWebMCP
Latest stable release19.4s5.4s
Maintenance status28.9s6.3s
Released in last 30 days24.1s5.9s
Drupal 11 upgrade audit77.6s16.4s

And its misses show what happens when a model reads HTML. For social_share it returned a Drupal 7 release from 2017. For crawl_control, an alpha. On maintenance status it scored 50%, often because it couldn’t find the status on the page and gave up.

The upgrade audit is where it matters

“Which of these modules blocks a Drupal 11 upgrade?” is the question people actually ask.

SetupScore (F1)TimeCost
Web1.0078s$0.45
MCP0.8916s$0.11

MCP found all four real blockers, plus one false positive caused by gaps in my release history. Web search was perfect, but took 78 seconds and four times the money.

I’d take 16 seconds. That’s an assistant you don’t have to wait for.

The model prefers MCP

With both available, 74% of tool calls went to MCP. 23% went to WebFetch, 3% to WebSearch.

The model checks the structured source first and only fetches a page when something looks missing.

Why MCP + web won

MCP + web got 98 questions right. MCP alone got 95. Four gained, one lost.

Every gain followed the same pattern: MCP first, then a drupal.org page to fill a gap in my data.

QuestionWhat MCP alone got wrongWhat the web fetch fixed
bootstrap_layout_builder, Drupal 11 compatibleRelease history was incompleteThe project page showed the current release line
amp, maintenance statusI didn’t store maintenance statusIt read the status from the page (after 11 fetches)
masquerade_as_role, security coverageI reported coverage, but it only has pre-releases, which aren’t coveredThe page showed no covered release
social_share, security coverageSame pre-release issueSame

The loss was migrate_source_csv. MCP alone got its maintenance status right. With web access, the model fetched a page and overwrote the right answer with a confident wrong one: “Maintenance fixes only”.

So the web isn’t a better source. It’s a fallback for holes in mine. Every MCP miss came down to data, not the model:

  • incomplete release history
  • no stored maintenance status
  • project-level security coverage for projects with only pre-releases
  • no explicit “no stable release” signal

Each of those is a fix on the server. Each fix pushes MCP alone toward 92%, without paying for web search on every question. The weak spots became a to-do list.

The catch: MCP is confidently wrong

MCP was confidently wrong on 7% of answers. Web search: 5%.

When structured data is wrong, the model believes it. It looks authoritative. That makes data quality more important with MCP, not less. It also means tools should say plainly when something is missing, like “this project has no stable release”, instead of leaving the model to guess.

Caveats

  • One attempt per question, one model. 90% versus 92% is two or three questions, which is noise. The big gaps are solid: memory versus everything else, and web versus MCP on speed and cost.
  • Ground truth was captured on the day of the run.
  • The MCP runs used a local copy of the site, loaded with that morning’s production data.
  • I built the server I’m testing. That’s why grading uses drupal.org, and why everything is public.
  • This is the server as of 26 September 2026. Some of the gaps above may be fixed by the time you read this.

Run it yourself

The harness, questions, ground truth and all 424 raw answers are on GitHub: pietervanleuven/drupal-mcp-comparison. Check my grading, swap the model, or point it at your own MCP server.

In the meantime, if you use an AI assistant for Drupal work, add the server. Stop letting it scrape drupal.org.