On June 29, 2026, Preset workspaces in the EU region experienced service degradation. Dashboards and the UI failed to load for approximately 45 minutes. The issue was caused by a software bug that allowed chart metadata to grow to an unusually large size, which overwhelmed the shared metadata database when loading dashboards with many charts.
Affected region: EU (eu-north-1)
Affected services: Dashboard loading, UI access
Duration: Approximately 45 minutes of unavailable service (~09:30 – 10:16 UTC). Service was restored after a database restart. A subsequent database upgrade was performed as a preventive measure with a brief planned maintenance window (~13 minutes)
Customer impact: All tenants hosted in the affected EU cluster experienced dashboard and UI loading failures during the incident window
Time | Event |
|---|---|
~09:30 | First customer reports that EU-region workspaces are not loading. Issue independently reproduced on Preset’s own EU testing workspace. |
09:37 | Engineering begins investigating; elevated queue depth observed in the EU region. |
09:53 | Root cause identified: a single workspace with dashboards containing ~90 charts with oversized metadata is exhausting the metadata database’s CPU and memory. |
10:09 | Database restarted to restore availability. |
10:16 | Database back online; CPU remains elevated due to the same workload resuming. |
10:23 | Offending database connections isolated to relieve pressure. |
10:41 | Decision made to upgrade the database to a larger instance to handle the load. |
10:43 | Affected customers notified of brief planned maintenance. |
10:56 | Upgraded database instance live. |
11:00 | Service fully stable; CPU normalized, queue cleared. |
The ag-grid-table chart plugin contained a bug that caused large internal data structures to be persisted into each chart’s metadata on save, instead of being kept as ephemeral runtime data. For workspaces with datasources containing many saved metrics and calculated columns, this resulted in chart metadata growing to ~7.7 MB per chart.
When a dashboard with ~90 such charts was loaded, the metadata database was required to read ~927 MB in a single query. This exhausted the database’s CPU and memory, causing cascading failures for all tenants sharing that database cluster.
This class of metadata bloat had been identified and addressed previously. The recurrence was a regression introduced in a recent release.
Immediate actions taken:
Restarted the metadata database to restore availability
Isolated the offending workload to relieve database pressure
Upgraded the database instance to a larger size to handle the load, stabilizing CPU utilization at ~30%
Code fixes deployed:
Patched the ag-grid-table plugin to exclude the oversized data structure from chart metadata on save, preventing future bloat
A data cleanup migration has been prepared to strip the bloated metadata from all existing affected charts
To prevent recurrence and improve our incident response:
Code fix (completed): The bug that caused metadata bloat has been patched. New chart saves no longer persist the oversized data.
Data migration (in progress): A database migration will clean up existing bloated chart metadata across all affected charts, eliminating the risk from previously saved data.
Infrastructure hardening: The database has been upgraded to a larger instance class to provide additional headroom.
Monitoring improvements: We are reviewing and improving our alerting to ensure database resource alerts are actionable and not lost in noise, including adding low-memory alerts with shorter evaluation windows.
Long-term architecture: We are evaluating changes to defer loading of large metadata columns so that dashboard endpoints only load the data they need, reducing the blast radius of any future metadata bloat.
We sincerely apologize for the disruption. Reliability is a top priority for us, and we are committed to ensuring this class of issue does not recur. If you have any questions, please reach out to your account team or contact support.