Live service health, ninety days of history and every incident we have written up.
Checked a minute ago
99.97% across all services over the last 90 days.
Every incident gets a public write-up, resolved or not.
Error rates are back to baseline. The slow query now runs against an index.
A fix is deployed and error rates are dropping.
A query without an index was saturating the primary database.
Queue times are back to normal. We added headroom to the runner pool.
Runs are starting slower than usual. We are investigating.