Two tools, ~1,000 tokens, the whole Cloudflare API

Model Context Protocol (MCP) became the standard way to wire AI agents to external services, but it has a core contradiction: agents need a large tool surface to do real work, while every tool added consumes context that the model needs for reasoning. A huge API can blow past a modern model's limits before any actual task begins.

Cloudflare's answer, first explored in Code Mode, is to stop describing operations as separate tools. Instead, the model writes JavaScript against a typed SDK, and that code runs safely in a Dynamic Worker Loader. The code acts as a compressed plan: the agent can explore endpoints, compose multiple calls, and return only what it needs, without shipping metadata about every possible operation through the context window. Anthropic reached the same design independently in its Code Execution with MCP write-up.

Cloudflare now ships this server-side via a new MCP server covering the entire Cloudflare API — DNS, Zero Trust, Workers, R2, the rest — with exactly two exposed tools. search() and execute() keep the agent's context at roughly 1,000 tokens, regardless of whether Cloudflare has 10 API endpoints or 10,000. For the current 2,500+ endpoint surface that's a 99.9% reduction compared to a conventional server that exports one tool per operation, which would need 1.17 million tokens — well beyond the context of even the largest foundation models.

images/BLOG-3184 1

The server and the technique are both available now. Cloudflare is also releasing the Code Mode SDK inside the Cloudflare Agents SDK, so any MCP server can adopt this pattern.

images/BLOG-3184 3

How server-side Code Mode works

Unlike the client-side experiment, this implementation keeps everything on the server. Instead of a sprawling list of tools, the model context contains just the two tool definitions, which is the total surface area shown below.

images/BLOG-3184 2

Discovery happens through search(). The agent gets a JavaScript object representing the full, de-referenced OpenAPI spec. It can filter by product, path, tag, or other metadata, winnowing thousands of endpoints down to a handful — and the raw spec never enters the context. The agent's code does the filtering, so only the result comes back.

Action happens through execute(). The agent writes code against a cloudflare.request() client that handles authentication, pagination, response checking, and chained calls within one sandboxed run.

Both tools run submitted code inside a Dynamic Worker isolate — a lightweight V8 sandbox without a file system, without environment variables, and with external fetches disabled unless an explicit outbound fetch handler allows them.

A DDoS protection walkthrough

To see the pattern in action, imagine an agent asked to "protect my origin from DDoS attacks." The agent first consults documentation via the Cloudflare Docs MCP Server, a Cloudflare Skill, or a web search. That research points it toward Cloudflare WAF and DDoS rule sets.

Step 1: discover endpoints with search(). The agent writes JavaScript to filter the spec object for zone ruleset and WAF endpoints:

async () => {
  const results = [];
  for (const [path, methods] of Object.entries(spec.paths)) {
    if (path.includes('/zones/') &&
        (path.includes('firewall/waf') || path.includes('rulesets'))) {
      for (const [method, op] of Object.entries(methods)) {
        results.push({ method: method.toUpperCase(), path, summary: op.summary });
      }
    }
  }
  return results;
}

The server returns the matching endpoints:

[
  { "method": "GET",    "path": "/zones/{zone_id}/firewall/waf/packages",              "summary": "List WAF packages" },
  { "method": "PATCH",  "path": "/zones/{zone_id}/firewall/waf/packages/{package_id}", "summary": "Update a WAF package" },
  { "method": "GET",    "path": "/zones/{zone_id}/firewall/waf/packages/{package_id}/rules", "summary": "List WAF rules" },
  { "method": "PATCH",  "path": "/zones/{zone_id}/firewall/waf/packages/{package_id}/rules/{rule_id}", "summary": "Update a WAF rule" },
  { "method": "GET",    "path": "/zones/{zone_id}/rulesets",                           "summary": "List zone rulesets" },
  { "method": "POST",   "path": "/zones/{zone_id}/rulesets",                           "summary": "Create a zone ruleset" },
  { "method": "GET",    "path": "/zones/{zone_id}/rulesets/phases/{ruleset_phase}/entrypoint", "summary": "Get a zone entry point ruleset" },
  { "method": "PUT",    "path": "/zones/{zone_id}/rulesets/phases/{ruleset_phase}/entrypoint", "summary": "Update a zone entry point ruleset" },
  { "method": "POST",   "path": "/zones/{zone_id}/rulesets/{ruleset_id}/rules",        "summary": "Create a zone ruleset rule" },
  { "method": "PATCH",  "path": "/zones/{zone_id}/rulesets/{ruleset_id}/rules/{rule_id}", "summary": "Update a zone ruleset rule" }
]

Inspection of the zone rulesets schema reveals the available phases:

async () => {
  const op = spec.paths['/zones/{zone_id}/rulesets']?.get;
  const items = op?.responses?.['200']?.content?.['application/json']?.schema;
  // Walk the schema to find the phase enum
  const props = items?.allOf?.[1]?.properties?.result?.items?.allOf?.[1]?.properties;
  return { phases: props?.phase?.enum };
}

{
  "phases": [
    "ddos_l4", "ddos_l7",
    "http_request_firewall_custom", "http_request_firewall_managed",
    "http_response_firewall_managed", "http_ratelimit",
    "http_request_redirect", "http_request_transform",
    "magic_transit", "magic_transit_managed"
  ]
}

The agent now knows precisely which phases it wants: ddos_l7 for DDoS protection and http_request_firewall_managed for WAF.

Step 2: act with execute(). The sandboxed client lists the zone's current rulesets:

async () => {
  const response = await cloudflare.request({
    method: "GET",
    path: `/zones/${zoneId}/rulesets`
  });
  return response.result.map(rs => ({
    name: rs.name, phase: rs.phase, kind: rs.kind
  }));
}

[
  { "name": "DDoS L7",          "phase": "ddos_l7",                        "kind": "managed" },
  { "name": "Cloudflare Managed","phase": "http_request_firewall_managed", "kind": "managed" },
  { "name": "Custom rules",     "phase": "http_request_firewall_custom",   "kind": "zone" }
]

Finding the managed DDoS and WAF rulesets already in place, the agent chains calls to inspect and adjust their sensitivity settings in one execution:

async () => {
  // Get the current DDoS L7 entrypoint ruleset
  const ddos = await cloudflare.request({
    method: "GET",
    path: `/zones/${zoneId}/rulesets/phases/ddos_l7/entrypoint`
  });

  // Get the WAF managed ruleset
  const waf = await cloudflare.request({
    method: "GET",
    path: `/zones/${zoneId}/rulesets/phases/http_request_firewall_managed/entrypoint`
  });
}

From first search to finished configuration, that end-to-end job took four tool calls.

Why this beats hand-maintained per-product servers

Cloudflare previously shipped separate MCP servers — one for DNS analytics, one for Workers Observability, others for other products. Each exposed a fixed, handbuilt tool set. That scales fine for a dozen operations, but not for 2,500.

The new unified server removes that maintenance burden entirely. The same search() and execute() code paths work for new endpoints and new products the moment they land in the API — no new tool definitions or MCP servers needed. The server also supports the GraphQL Analytics API.

Security and authorization follow current MCP standards. The server is OAuth 2.1 compliant via Workers OAuth Provider, downscoping tokens to only the permissions the user selected at connection time. The agent gets no capabilities beyond what was explicitly granted.

images/BLOG-3184 4

Context-reduction approaches compared

Code Mode is one of several paths to smaller contexts, each with tradeoffs.

  • Client-side Code Mode, Cloudflare's first experiment, lets a model write TypeScript against typed SDKs and run it in a Dynamic Worker Loader on the client. It works, but the client must ship with, and trust, a secure sandbox. The technique has since been adopted in Goose and in Anthropic's Claude SDK as Programmatic Tool Calling.
  • Command-line interfaces offer self-documenting, progressive discovery. Tools like OpenClaw and Moltworker convert MCP servers to CLIs via MCPorter. The catch: an agent needs a shell session, which isn't universal and expands the attack surface well beyond what a sandboxed isolate offers.
  • Dynamic tool search, as seen in Anthropic's Claude Code, surfaces a task-relevant subset of tools. That cuts context but adds a search mechanism that must itself be maintained and evaluated, and each matched tool still consumes tokens.
images/BLOG-3184 5

Server-side Code Mode combines the strengths: a token cost that doesn't grow with API size, no agent-side modifications, built-in progressive discovery, and isolated execution. The agent sees two tools and writes code; the server does the heavy lifting.

Getting started

The server is live. Point any MCP client at the server URL, authorize through Cloudflare, and choose the permissions to grant the agent. The MCP client config is straightforward:

{
  "mcpServers": {
    "cloudflare-api": {
      "url": "https://mcp.cloudflare.com/mcp"
    }
  }
}

For CI/CD or automation without an interactive OAuth flow, generate a Cloudflare API token with the required permissions. Both user and account tokens work as bearer tokens in the Authorization header. Full setup details live in the Cloudflare MCP repository.

Beyond a single API

Code Mode solves context bloat for one service, but agents rarely talk to just one. A developer's agent reaching across Cloudflare, GitHub, a database, and internal docs will again hit context pressure with each additional MCP server.

Cloudflare MCP Server Portals already address that by compositing multiple MCP servers behind one gateway, with unified authentication and access control. The roadmap is a first-class Code Mode integration across all MCP servers behind that gateway, offering the same fixed-token footprint and progressive discovery no matter how many services are aggregated.