On-demand profiling arrives for Workers and Durable Objects

Aggregate metrics and logs can tell you that a Worker is consuming too much CPU or memory, but not which function is responsible. Profiling closes that gap. Workers and Durable Objects now support on-demand CPU and memory profiles, requested from the Workers Observability page, viewed as an interactive flamegraph, or downloaded for offline analysis. Profiling in production is the point: it captures the code paths your application actually executes under real traffic.

Two entry points are available. With the cf package installed, the CLI is enough:

cf workers versions profile latest \
     --worker-id "$WORKER_ID_OR_NAME" \
     --duration-ms 5000 \
     --profile-type cpu > worker-cpu.pprof

In the dashboard, open your account's Worker list via Build → Compute → Workers & Pages, select a Worker, and switch to its Observability tab. The drop-down there includes a "Flamegraph" option:

From that page you request either a CPU or a memory profile and set a duration for the sampling run. You can also choose which version of the Worker to profile. Because the profiler attaches to an existing isolate rather than starting a new one, pick a deployed version that receives enough traffic — low-traffic Workers are difficult to profile.

When the capture finishes, the results render as a flamegraph. Each rectangle is a function call and its width reflects the CPU time or memory attributed to it. Clicking a function narrows the view to it; hovering reveals details.

Capture more than one profile and compare the widest frames — those are your biggest resource consumers, and outliers are easier to spot across multiple runs. The table view is a useful complement for seeing which functions appear most often in a profile.

TypeScript Workers should enable source maps. Without them, the profile is likely to show obfuscated function names that are hard to map back to your code.

CPU profiling in practice: the R2 binding

The Worker implementing the R2 binding receives heavy traffic, so removing even small amounts of wasted CPU work pays off. A 50-second CPU trace produced this profile:

The widest frames were unsurprising and offered no clear optimization: decryptBlock handles decryption of object data on the way out, and fillResponse moves bytes.

Less visible frames required the table view. Sorting by Samples pointed at genericR2JsonReplacer: over 5% of CPU time, and calling itself recursively.

The replacer was invoked by JSON.stringify on every node of the JSON tree while separately walking that same tree, so a value nested five levels deep was processed five times. Eliminating the duplicate work made genericR2JsonReplacer 2.7x faster.

The second candidate was a duplicate call to metrics:

A single invocation was heavy enough to account for 1% of CPU time in the profile. Caching the first result in a variable rather than calling metrics again removed that cost.

function ship() {
  if (!registry.metrics().length) {
    return;
  }

  // ... more code here ...

  return Response(registry.metrics());
}

Memory: fixing evictions from a partially disabled code path

An internal Worker was being evicted with "Exceeded Memory" errors because its P999 memory hovered around 133 MB against the 128 MB Worker limit. Errors and metrics identified memory as the problem but not its location.

A heap profile taken from a production instance and opened in pprof attributed roughly 66.7% of allocations to the team's Prometheus code. That code was believed to be disabled, yet it still instrumented many of the Worker's code paths, consumed substantial memory, and never exported the data it collected. The profile showed the shutdown was only partial — the memory cost of full instrumentation remained.

Removing the Prometheus code path entirely improved memory usage quickly:

Percentile

Before (MB)

After (MB)

P50

70

54

P90

94

79

P99

113

97

P999

133

118

The resulting P999 left about 10 MB of headroom under the 128 MB limit:

Running a profile on the edge

Local profiling through Chrome DevTools has long been available for Workers, but local traffic does not resemble production traffic in volume or shape — and there was no way to start a session against a Worker running in production. Doing so is complicated by how Workers are distributed: requests are routed to the nearest data center, Workers are replicated across data centers and metals, and a Worker under load can be replicated on the same metal multiple times. Durable Objects add placement indirection, since they can be placed dynamically and can move.

Before sampling can begin, the platform has to establish which version the developer wants to profile, which data center recently ran it, whether that isolate is still loaded there, whether that isolate belongs exclusively to the requesting account, and — for Durable Objects — where the exact live primary actor is. Some of this comes from the client issuing the request, usually the dashboard, though requests can also be sent directly:

curl 'https://api.cloudflare.com/client/v4/accounts/<account_id>/workers/workers/<worker_name>/versions/latest/profile' \
  --data-raw '{"duration_ms":5000,"profile_type":"cpu"}'

The request URL carries both the script name and the version to be profiled. No new isolate is started for a profile; the goal is to observe real production execution.

Sampling without pausing the application

Profiling is only meaningful if the application keeps executing while samples are collected, so the Workers Runtime holds the isolate lock for profiler lifecycle operations alone. A CPU profile proceeds as follows:

  1. Acquire the isolate and its lock.
  2. Create a V8 CPU profiler and begin sampling at a one millisecond interval.
  3. Release the lock so normal requests can run.
  4. Wait for the requested duration.
  5. Reacquire the lock and stop profiling.
  6. Release the lock and serialize the result outside the critical section.

Holding the lock for the full duration would block JavaScript execution and invalidate the profile.

Durable Objects diverge from stateless Workers here. For a Worker, any running isolate of the requested version may be sampled. Because a Durable Object is stateful and named, you select which object to profile by name; the runtime routes the request to the metal owning that specific actor and returns a profile of the isolate running it.

Limitations and what comes next

Two constraints are worth knowing before relying on on-demand profiling:

  • Sessions must be started explicitly, so periods when a Worker misbehaves can pass unobserved.
  • The memory profiler only reports allocations made during the profiling window, which means start-up allocations are invisible to it.

Continuous profiling is in development to address both. It would capture samples automatically and expose them in the dashboard for exploration, making rare events easier to catch.