Our responsibility, and what went wrong

At Vercel, we take ownership for our vendors rather than blaming them. The unexpected outage of AWS us-east-1 (which we call the iad1 region) is public knowledge now, but that doesn't change who is accountable: Vercel is unequivocally responsible for this incident. We built our infrastructure on AWS primitives, participate in the AWS marketplace, and maintain close technical partnerships with AWS — the component choices in our system design are ours alone.

Our promise is to simplify the cloud. Through framework-defined infrastructure, we operate Compute, CDN, and Firewall services across 19 AWS regions, with traffic terminated and secured in 95 cities across 130+ global points of presence. On October 20, we fell short of that promise. While we continued to serve a significant amount of traffic and shielded customers from a single global point of failure, our ambition is that customers never drop a single request — even during an outage.

Two independent incidents

The AWS failure produced two separate incidents on the Vercel side. The first came from the initial loss of the us-east-1 data center and caused temporary disruption to serving production traffic. The second came from an outage of our feature flag provider, itself a cascading effect of the AWS disruption. That second incident did not affect production traffic serving, but it severely disrupted the control plane — the dashboard, APIs, builds, and log processing.

Loss of us-east-1

Our production traffic stack is designed to survive the loss of any one of our 19 operating regions. We've proven this capability before, but Monday's event created an unexpected cascading failure that extended the impact.

In red: Impact of primary outage on global traffic. In red: Impact of primary outage on global traffic. In red: Impact of primary outage on global traffic. In red: Impact of primary outage on global traffic.

Our team received the first alerts at 06:55 UTC. What initially looked like minor disruption prompted us to begin re-routing traffic away from us-east-1 at 07:15 UTC, restoring service for end users connecting directly to that region. At 07:25 UTC we started re-routing function invocations for customers with configured backup or secondary regions, restoring their full service.

At 07:45 UTC, a cascading failure struck a portion of our global caching infrastructure. Static file serving failures peaked at about 22% of traffic — the largest impact of the entire event. Service was restored by 08:18 UTC, and we've already mitigated the root cause of the cascade.

The yellow line shows elevated errors in function invocations during the outage window. The yellow line shows elevated errors in function invocations during the outage window. The yellow line shows elevated errors in function invocations during the outage window. The yellow line shows elevated errors in function invocations during the outage window.

Function invocations were restored at 09:21 UTC for customers using us-east-1 as their only function region.

Control plane disruption

The Vercel control plane runs primarily in us-west-1, so it wasn't directly hit by the us-east-1 outage. But at 19:20 UTC, a feature flag provider used by the control plane suffered a major outage, likely as a downstream effect of the earlier AWS disruption. The control plane supports the dashboard, CLI, API, and services like log forwarding. It's entirely separate from the serving stack, so this had no impact on production traffic — but customers couldn't deploy, roll back, or perform other administrative tasks during the window.

The feature flag provider's unavailability also exhausted resources in our primary Kubernetes cluster, delaying or failing processing. At 19:48 UTC we began experimenting with and incrementally rolling out a mitigation for both the provider outage and the code pattern responsible for resource exhaustion. Full control plane restoration came around 03:00 UTC.

What a request experienced

To illustrate the impact, here's what happened to a request to rauchg.com during the incident:

DNS resolution

DNS lookup resolves the A record. Vercel DNS uses anycasted, intelligent steering to route to healthy endpoints, isolating regional failures. This worked as expected.

CDN IP assignment

The DNS query returns a Vercel CDN IP. Anycast allows a single IP to steer traffic to the nearest healthy edge region, limiting latency and isolating faults. We still recommend CNAME records, especially when not using Vercel DNS, because a CNAME gives our infrastructure teams another traffic-steering point for dynamic updates. This worked as expected.

Edge protections

L3–L5 protections provide managed DDoS defense with automatic detection and mitigation, covering SYN, UDP, ICMP, and TCP floods plus malformed packets. The TLS handshake relies on a global component called the HTTPS Terminator, which requires sophisticated certificate replication across tens of millions of unique domains each week. Certificates are actively pushed to all CDN edge regions on creation and cached locally for fault tolerance. The L7 Firewall, powered by a component called Protectd, operates regionally without global dependencies — a design choice that gives industry-leading mitigation speeds. All of these worked as expected.

Request serving

Once the CDN holds deployment metadata, it can map the domain to its deployment destination. Our globally replicated metadata infrastructure has no single point of failure. This worked globally, but customers hitting iad1 directly saw disruption until we re-routed that region.

Routing then relies on a component called Regional Cache, which stores static assets and cacheable dynamic requests. This system persists data globally on first use and is designed with strict durability and latency guarantees. The origin of routing metadata is us-east-1 S3 — and while metadata propagates into every region, a cascading failure of the regional cache service caused a wider scope of disruption than expected.

Routing can take several paths:

  • Routing Middleware execution: Customers can use Fluid compute to run Node.js logic in the request path. These functions deploy to every CDN region, so total impact was limited — there was initial disruption only for users hitting us-east-1 directly.
  • Function invocation: No impact for customers not using us-east-1 as their only region. Those with iad1 as a primary region were failed over to a backup if configured.
  • Serving Cache-Control content: Previously served and cached responses from the Response Cache continued working as expected.
  • Static and ISR content: Cached content at the edge served fine. ISR content with us-east-1 as origin had temporary disruption if not yet cached at the edge.

How the outage unfolded

The incident began just before 07:00 UTC on Monday, October 20, when automated monitoring detected errors across Vercel services and paged on-call engineers. Within minutes, the initial responder escalated to additional engineers and opened an incident call. By 07:07 UTC, the team had determined that AWS-managed services were likely at fault and alerted Vercel's most senior responders through the "Panic Rotation."

Failures in the serving stack prompted a decision to reroute traffic away from the iad1 data center at 07:15 UTC. The incremental migration began at 07:20 UTC and completed at 07:44 UTC. Meanwhile, provisioning of new build machines in iad1 was also failing, so builds were rerouted to a different region by 07:40 UTC. However, builds that depended on external resources tied to AWS us-east-1, such as databases outside Vercel, continued to fail.

At 08:16 UTC, the team observed increased failures loading the Vercel Dashboard and elevated error rates in the Vercel Teams API. The root cause was traced to elevated request latency to external dependencies, which led to CPU starvation. The service was manually scaled up at 08:51 UTC, and the Teams API recovered by 09:15 UTC.

New deployments using Routing Middleware or Vercel Functions with iad1 as the selected region began failing at 09:45 UTC due to AWS API issues in us-east-1. This affected all projects configured with Secure Compute and iad1 as the specified region. Health checks in the iad1 builds cluster passed at 10:22 UTC, and Secure Compute builds were restored by routing them back to the cluster. Service functionality for APIs that create Vercel Functions and Middleware in iad1 was restored at 11:52 UTC.

Recurring Symptoms in the afternoon and evening

At 17:19 UTC, there was a small increase in failures invoking Vercel Functions in iad1, along with paging for failures scheduling new instances of Functions routing services. The team determined that continued EC2 capacity issues in AWS us-east-1 were causing the scheduling failures. To free capacity for the routing services, they reduced provisioning of non-critical infrastructure in iad1, which led to decreased errors in invoking Vercel Functions. The configuration was also updated at 17:29 UTC to prefer other regions for serving Functions when available. Invocation errors continued at a much lower rate until recovery at 17:49 UTC.

Another degradation of AWS APIs at 17:47 UTC caused build failures due to errors creating Vercel Functions and Middleware in iad1. At 18:50 UTC, the service backing Vercel Runtime Logs began experiencing timeouts. Around 19:20 UTC, the Vercel Dashboard began failing to load, with elevated error responses from Vercel API services and aggressive restarts of containers running Control Plane services.

At 19:31 UTC, automated alerts for elevated API errors paged the on-call engineer, opening a separate incident. By 19:41 UTC, timeouts to Vercel's feature flag provider had been identified as the underlying issue behind 500 responses and container restarts. A change released at 19:48 UTC reduced timeouts to the flag provider and cut container restarts, though API services remained degraded. An experimental fix was rolled out in incremental batches starting at 21:20 UTC, with full rollout to all degraded control plane services completed by 22:31 UTC.

Overnight recurrence and final recovery

At 00:20 UTC on Tuesday, October 21, the Dashboard and API degradation resumed because the experimental fix had not landed in 100% of services. Runtime Logs were also missing in several regions, with backlogs of 5–6 hours in a subset. A change to increase throughput was merged, and by 02:37 UTC, all flags clients across Control Plane services were reconfigured to use the experimental fix.

A second fix for the logs pipeline, released at 03:01 UTC, successfully unblocked log processing, and the backlog age began to decrease across affected regions. The Vercel Dashboard and API services fully recovered at 03:30 UTC, with Runtime Logs following at 06:21 UTC.

Next steps

Vercel is preparing a detailed technical postmortem that will address the failure modes exposed during the disruption. The company will also continue and intensify its rehearsal of major regional disruptions.