Skip to content
Convex · Data and observabilitySep 20, 2025, 12:18 UTC 1 year ago

Convex traffic having downtime

MajorCapacityUpdated 3h ago
Sep 20, 12:18 UTCSep 20, 13:15 UTC
Duration
57m
Impact
Major
Root cause
Capacity
Convex, 90 days
6 incidents
Affected
Not listed by the vendor.
Status page

Final update

From around 5:18am to 5:54am Pacific \(12:18pm to 12:54pm UTC\), Convex had a 36 min period of intermittent downtime that affected all Convex services. The specific issue was a cascading failure in our traffic layer. We had a traffic node \(Caddy\) run out of memory due to an unforeseen load spike and instead of just being restarted/replaced this node was marked as permanently down by our container management layer \(Nomad\) which led to the issue propagating to all traffic servers. Since the incident we've more than doubled the size of our traffic layer, fixed the failover behavior which led to nodes staying failed after OOMing, and will be investigating alternative traffic services. As always data was safe during this incident but we really apologize for the availability impact during that time period.

Timeline

  1. Postmortem · Sep 20, 20:42 UTC
    From around 5:18am to 5:54am Pacific \(12:18pm to 12:54pm UTC\), Convex had a 36 min period of intermittent downtime that affected all Convex services. The specific issue was a cascading failure in our traffic layer. We had a traffic node \(Caddy\) run out of memory due to an unforeseen load spike and instead of just being restarted/replaced this node was marked as permanently down by our container management layer \(Nomad\) which led to the issue propagating to all traffic servers. Since the incident we've more than doubled the size of our traffic layer, fixed the failover behavior which led to nodes staying failed after OOMing, and will be investigating alternative traffic services. As always data was safe during this incident but we really apologize for the availability impact during that time period.
  2. Resolved · Sep 20, 13:15 UTC
    This incident has been resolved.
  3. Monitoring · Sep 20, 12:57 UTC
    An unexpected traffic pattern overloaded some of our services, causing intermittent unavailability across Convex instances. We've added extra capacity and are monitoring to ensure that the system is stable.
  4. Investigating · Sep 20, 12:18 UTC
    We are currently investigating this issue.

More from Convex

Full history
StartedIncidentDuration
Sep 2821:09 UTC5 days agoAI gateway unavailable1h 4m
Sep 1823:01 UTC2 weeks agoDashboard logins degraded1h 7m
Aug 1623:33 UTC6 weeks agoAccount verification/team invite emails not working1h 53m
Aug 1203:41 UTC7 weeks agoSome HTTP Actions returning 40420m
Aug 501:42 UTC8 weeks agoAction failures for business deployments that use createFunctionHandle15m
Jul 1505:35 UTC2 months agoNew deployments unable to push Node Actions to Convex57m

Also caused by capacity and load

All
StartedIncidentDuration
Oct 114:47 UTC2 days agoActions Job DelaysGitHub3h 9m
Sep 2919:43 UTC4 days agoInvestigating service degradation - xAI modelsCursor35m
Sep 2318:42 UTC10 days agoNetwork Performance Degradation , Asia-PacificCloudflare6d 4h
Sep 2213:19 UTC11 days agoWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 1607:20 UTC2 weeks agoDegradation with Gemini 3.8 FlashGitHub10h 28m
Sep 1509:47 UTC2 weeks agoDisruption with some GitHub servicesGitHub1h 30m

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.