Incident Report: EU Production Unavailability — 7 September 2026
Status: Resolved
Duration: 2 hours, 15 minutes
Summary
On Monday, 7 September 2026, Molecule's EU production environment became unavailable at 07:41 CEST (05:41 UTC) and was fully restored at 10:01 CEST (08:01 UTC). The cause was growth in valuation and event data that outpaced provisioned storage on that environment.
Service was restored by expanding database capacity and draining queued work. Molecule's US and other production environments were unaffected throughout. No customer data was lost or corrupted.
Timeline (CEST, UTC+2)
| Time |
Event |
| 07:41 |
The EU environment stops serving requests. |
| 09:01 |
Incident declared and engineering engaged. |
| 09:56 |
Database capacity expanded; service restored. |
| 10:01 |
Incident resolved. Queued processing drains normally. |
What happened
Molecule computed the full valuation and settlement history for a set of long-tenor, high-resolution contracts at load time. The YTD aggregations produced by that computation, which are generated at sub-daily resolution, expanded rapidly, and the resulting data growth consumed the storage remaining on the EU environment.
The growth was specific to YTD aggregation at sub-daily resolution rather than to the underlying valuation or settlement work. That aggregation is now disabled for large-volume trades while it is reworked.
Contributing factors
Two factors extended the time to restoration:
- Detection did not automatically escalate to a paged incident. Monitoring identified the condition correctly, but the step from detection to an acknowledged page was manual rather than automatic.
- Authority to make emergency infrastructure changes was not clearly documented. The fix was correctly diagnosed early, but the change needed to apply it was not obviously within the responder's authority, which added time to mitigation.
What we are doing about it
Completed
- YTD aggregation is disabled for large-volume trades while the underlying computation is reworked.
- Database capacity expanded, with automatic storage scaling enabled on all production environments.
- A second tier of escalation is in place when a page goes unacknowledged.
- Authority to make emergency infrastructure changes during an incident is documented and findable.
- Large data loads are reviewed by our implementations and engineering teams before they run.
In progress (next 30 days)
- Automatic incident creation and paging for all detectable severity-level alerts.
- Formalizing that review into a documented pre-flight checkpoint, including staging and sign-off.
- Expanded on-call coverage for both support and engineering, covering monitoring as well as escalation.
- Enforced capacity limits sizing database provisioning against application throughput, applied automatically rather than case by case.
- Consolidation of alerting, incident management, and status communication onto a single platform, with a reduction in routine alert volume.
Product changes under development
- An account-level cutover date, so historical trade data loaded for record-keeping purposes does not trigger valuation before a defined date.
- Optimization of YTD aggregations for sub-daily resolution trades, which currently generate unnecessarily high data volume.
We have scheduled a 30-day review to confirm completion of the items above.
We are very sorry for the delays on 7 September, and hope this description and analysis explains how we will prevent a recurrence.
Questions
Customers with questions about this incident should contact their account team or Molecule support.