Skip to content
Convex · Data and observabilityMar 4, 2026, 17:28 UTC 7 months ago

high error rates on customer backends

CriticalCapacityUpdated 3h ago
Mar 4, 17:28 UTCMar 4, 18:28 UTC
Duration
1h
Impact
Critical
Root cause
Capacity
Convex, 90 days
6 incidents
Affected
Free & StarterLive Traffic - Node.js RuntimeLive Traffic - Database ServicesLive Traffic - Function Runtime
Status page

Final update

Convex had downtime for some customers caused by an anomalous load spike that pushed an internal service into a load-shedding mode, which then triggered an unexpected panic in a caching library we were using. Specifically the load triggered two queue management algorithms: CoDel which proactively drops requests to keep queues small, and adaptive-LIFO which dequeues in reverse order to avoid wasting time on old requests. These are both rather subtle algorithms that large services use to avoid congestion collapse under high load or attack. The panic in the caching library was just a bug that depended on both these algorithms simultaneously. We've made some changes as a result of this incident but the key lesson is that services should try to avoid switching logical behavior during high load. When systems are stressed switching to infrequently-used codepaths can often make matters worse. We're now going to be proactively triggering CoDel and adaptive-LIFO at steady state to ensure that we're exercising this worst-case flow at all times. We apologize to our customers for impacting your products and services. We’re focusing intensely over the next few weeks on hardening our systems to p

Timeline

  1. Postmortem · Mar 5, 18:14 UTC
    Convex had downtime for some customers caused by an anomalous load spike that pushed an internal service into a load-shedding mode, which then triggered an unexpected panic in a caching library we were using. Specifically the load triggered two queue management algorithms: CoDel which proactively drops requests to keep queues small, and adaptive-LIFO which dequeues in reverse order to avoid wasting time on old requests. These are both rather subtle algorithms that large services use to avoid congestion collapse under high load or attack. The panic in the caching library was just a bug that depended on both these algorithms simultaneously. We've made some changes as a result of this incident but the key lesson is that services should try to avoid switching logical behavior during high load. When systems are stressed switching to infrequently-used codepaths can often make matters worse. We're now going to be proactively triggering CoDel and adaptive-LIFO at steady state to ensure that we're exercising this worst-case flow at all times. We apologize to our customers for impacting your products and services. We’re focusing intensely over the next few weeks on hardening our systems to prevent issues like this from happening again.
  2. Resolved · Mar 4, 18:28 UTC
    Incident resolved. We're very sorry for the impact on your projects. Our team will be publishing a detailed postmortem soon.
  3. Monitoring · Mar 4, 18:14 UTC
    We've identified the issue and remediated the problem. We're monitoring before we declare the all clear.
  4. Investigating · Mar 4, 17:28 UTC
    We are currently investigating an issue leading to elevated error rates on customer backends

More from Convex

Full history
StartedIncidentDuration
Sep 2821:09 UTC5 days agoAI gateway unavailable1h 4m
Sep 1823:01 UTC2 weeks agoDashboard logins degraded1h 7m
Aug 1623:33 UTC6 weeks agoAccount verification/team invite emails not working1h 53m
Aug 1203:41 UTC7 weeks agoSome HTTP Actions returning 40420m
Aug 501:42 UTC8 weeks agoAction failures for business deployments that use createFunctionHandle15m
Jul 1505:35 UTC2 months agoNew deployments unable to push Node Actions to Convex57m

Also caused by capacity and load

All
StartedIncidentDuration
Oct 114:47 UTC2 days agoActions Job DelaysGitHub3h 9m
Sep 2919:43 UTC4 days agoInvestigating service degradation - xAI modelsCursor35m
Sep 2318:42 UTC10 days agoNetwork Performance Degradation , Asia-PacificCloudflare6d 4h
Sep 2213:19 UTC11 days agoWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 1607:20 UTC2 weeks agoDegradation with Gemini 3.8 FlashGitHub10h 28m
Sep 1509:47 UTC2 weeks agoDisruption with some GitHub servicesGitHub1h 30m

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.