Contents 8 sections
I asked Claude 106 Drupal questions four ways. The MCP server was as accurate as web search. It was also 3.5× faster and 5× cheaper.
Here’s how I measured it, and where it still falls short.
The problem
Ask an assistant “is this module ready for Drupal 11?” and it has two options.
It can answer from memory. That memory is months old, and release data changes daily.
Or it can search the web. That means several round trips, fetching drupal.org pages full of markup, and digging the answer out of HTML meant for humans.
drupalreleases.com already has this data in structured form. So I gave it an MCP server (more on that here). MCP, the Model Context Protocol, lets an assistant call a tool like get-project and get a clean answer back. No scraping.
I believed that would beat web search. I wanted to know by how much.
How I tested it
One model, Claude Sonnet. Four setups:
| Setup | What the model can use |
|---|---|
| Memory | Training data only |
| Web | WebSearch + WebFetch |
| MCP | Only the drupalreleases.com MCP server |
| MCP + web | Both, to see which it prefers |
The questions cover 20 projects in three popularity bands. Popular ones like webform, pathauto and paragraphs. Mid-range ones like telephone and config_inspector. And long-tail modules like crawl_control and masquerade_as_role.
Every project got the same five questions:
- What’s the latest stable version, and when was it released?
- Is it Drupal 11 compatible?
- Is it covered by security advisories?
- What’s its maintenance status?
- Was anything released in the last 30 days?
On top of that, two scenarios. An upgrade audit: “which modules in this composer.json block Drupal 11?” And a discovery question: “I need to schedule content publishing, what should I use?”
That’s 106 questions, 424 answers, run on 26 September 2026.
Three rules kept it fair:
- Clean runs. Every answer is a fresh
claude -pin an empty directory. No settings, no CLAUDE.md, no other MCP servers. - Abstaining is allowed. A JSON schema forces structured output, and every field can be null. “I don’t know” isn’t counted as wrong.
- Grading uses drupal.org, not my data. The correct answers come from drupal.org’s release-history XML. My server gets no home advantage.
The numbers
How often each setup got it right:
| Setup | Correct | Confidently wrong | Abstained |
|---|---|---|---|
| Memory | 33% | 2% | 41% |
| Web | 88% | 5% | 0% |
| MCP | 90% | 7% | 0% |
| MCP + web | 92% | 7% | 0% |
And what each answer cost:
| Setup | Median time | Cost | Input tokens | Tool calls |
|---|---|---|---|---|
| Memory | 5.9s | $0.009 | 1,309 | 0 |
| Web | 20.4s | $0.072 | 11,108 | 4.1 |
| MCP | 5.9s | $0.013 | 5,730 | 1.3 |
| MCP + web | 6.7s | $0.020 | 8,880 | 1.7 |
All 106 questions cost $7.63 with web search. With MCP: $1.42.
MCP is exactly as fast as answering from memory, 5.9 seconds median. Except it’s right.
Split by question type:
| Type | Memory | Web | MCP | MCP + web |
|---|---|---|---|---|
| Latest stable release | 15% | 85% | 90% | 90% |
| Drupal 11 compatible | 70% | 100% | 95% | 100% |
| Security coverage | 65% | 100% | 90% | 100% |
| Maintenance status | 0% | 50% | 75% | 75% |
| Released in last 30 days | 0% | 100% | 100% | 100% |
| Module discovery | 100% | 100% | 100% | 100% |
And by popularity:
| Band | Memory | Web | MCP | MCP + web |
|---|---|---|---|---|
| Popular | 40% | 97% | 97% | 94% |
| Mid | 34% | 86% | 89% | 94% |
| Long tail | 13% | 77% | 83% | 90% |
Memory and web search both get worse as modules get more obscure. Those are exactly the modules you need to check before an upgrade.
What the numbers say
Memory knows it doesn’t know
Memory abstained on 41% of questions. It was fine on things that don’t change: every discovery question right, 70% on Drupal 11 compatibility.
Versions, dates and maintenance status were guesswork. 0% on maintenance status. 13% overall on long-tail modules. Useless for release questions.
Web search is right, eventually
Web search made four tool calls per question on average and read about twice the tokens MCP did. It was usually right. But the time adds up:
| Type | Web | MCP |
|---|---|---|
| Latest stable release | 19.4s | 5.4s |
| Maintenance status | 28.9s | 6.3s |
| Released in last 30 days | 24.1s | 5.9s |
| Drupal 11 upgrade audit | 77.6s | 16.4s |
And its misses show what happens when a model reads HTML. For social_share it returned a Drupal 7 release from 2017. For crawl_control, an alpha. On maintenance status it scored 50%, often because it couldn’t find the status on the page and gave up.
The upgrade audit is where it matters
“Which of these modules blocks a Drupal 11 upgrade?” is the question people actually ask.
| Setup | Score (F1) | Time | Cost |
|---|---|---|---|
| Web | 1.00 | 78s | $0.45 |
| MCP | 0.89 | 16s | $0.11 |
MCP found all four real blockers, plus one false positive caused by gaps in my release history. Web search was perfect, but took 78 seconds and four times the money.
I’d take 16 seconds. That’s an assistant you don’t have to wait for.
The model prefers MCP
With both available, 74% of tool calls went to MCP. 23% went to WebFetch, 3% to WebSearch.
The model checks the structured source first and only fetches a page when something looks missing.
Why MCP + web won
MCP + web got 98 questions right. MCP alone got 95. Four gained, one lost.
Every gain followed the same pattern: MCP first, then a drupal.org page to fill a gap in my data.
| Question | What MCP alone got wrong | What the web fetch fixed |
|---|---|---|
bootstrap_layout_builder, Drupal 11 compatible | Release history was incomplete | The project page showed the current release line |
amp, maintenance status | I didn’t store maintenance status | It read the status from the page (after 11 fetches) |
masquerade_as_role, security coverage | I reported coverage, but it only has pre-releases, which aren’t covered | The page showed no covered release |
social_share, security coverage | Same pre-release issue | Same |
The loss was migrate_source_csv. MCP alone got its maintenance status right. With web access, the model fetched a page and overwrote the right answer with a confident wrong one: “Maintenance fixes only”.
So the web isn’t a better source. It’s a fallback for holes in mine. Every MCP miss came down to data, not the model:
- incomplete release history
- no stored maintenance status
- project-level security coverage for projects with only pre-releases
- no explicit “no stable release” signal
Each of those is a fix on the server. Each fix pushes MCP alone toward 92%, without paying for web search on every question. The weak spots became a to-do list.
The catch: MCP is confidently wrong
MCP was confidently wrong on 7% of answers. Web search: 5%.
When structured data is wrong, the model believes it. It looks authoritative. That makes data quality more important with MCP, not less. It also means tools should say plainly when something is missing, like “this project has no stable release”, instead of leaving the model to guess.
Caveats
- One attempt per question, one model. 90% versus 92% is two or three questions, which is noise. The big gaps are solid: memory versus everything else, and web versus MCP on speed and cost.
- Ground truth was captured on the day of the run.
- The MCP runs used a local copy of the site, loaded with that morning’s production data.
- I built the server I’m testing. That’s why grading uses drupal.org, and why everything is public.
- This is the server as of 26 September 2026. Some of the gaps above may be fixed by the time you read this.
Run it yourself
The harness, questions, ground truth and all 424 raw answers are on GitHub: pietervanleuven/drupal-mcp-comparison. Check my grading, swap the model, or point it at your own MCP server.
In the meantime, if you use an AI assistant for Drupal work, add the server. Stop letting it scrape drupal.org.