Degraded kernel acquisition

Incident Report for Hex

Postmortem

After fully analyzing the timeline of events, we’ve confirmed that the root cause of the incident was an issue in the AWS EBS storage system backing our primary database replica. The issue began on May 12 11:00 PM PDT according to AWS, and the effect was gradual degradation of performance on the replica, eventually leading to a cascade where even our primary database got backpressured on critical write operations, which triggered the instability causing the incident. The AWS issue was resolved May 13 9:32 AM PDT, after we pushed a database configuration change that caused the EBS system software to reset. We then stabilized and cleaned up systems on our end before resolving the incident.

While the root cause of this incident was on the AWS side, we are not satisfied that our detection and response were fast enough. We are investing in the following mechanisms to improve:

  • We have automated testing to monitor provisioning of kernels, but it is combined with a larger testing suite that makes the signal less immediately actionable. We will be extracting this out as its own monitor so we can more quickly identify and react.
  • Our database monitoring around latencies will be made more comprehensive, including the replica database, so we do not miss this early signal.
  • We are improving our incident runbook to cover some of these checks earlier in the process so we zero in on the root cause more quickly.
Posted May 21, 2025 - 13:21 PDT

Resolved

This incident has been resolved.
Posted May 13, 2025 - 12:10 PDT

Monitoring

A fix has been implemented and we are monitoring the results.
Posted May 13, 2025 - 11:09 PDT

Identified

Some users are still experiencing issues as we implement a fix.
Posted May 13, 2025 - 10:35 PDT

Monitoring

A fix has been implemented and we are monitoring the results.
Posted May 13, 2025 - 09:58 PDT

Update

We are continuing to investigate this issue.
Posted May 13, 2025 - 09:27 PDT

Investigating

We are currently investigating this issue.
Posted May 13, 2025 - 07:38 PDT
This incident affected: Kernels.