• Azure Databricks reported two separate Central US incidents on 5 September, both affecting compute startup and jobs
  • Both incidents were resolved, and the available status reports do not establish that they had the same cause

The fact

Azure Databricks reported two separate service incidents in its Central US region on 5 September. The first, ES-2194668, affected Classic Compute, Declarative Pipelines and Jobs. Reported problems included failed cluster launches, clusters stuck in Pending or Starting states, notebook attachment failures and jobs or pipelines that failed or remained delayed. Status tracking recorded the incident for more than two hours.

A second incident, ES-2195009, later affected Compute and Unity Catalog in Central US. Customers again encountered cluster launch failures, notebook attachment errors and jobs that failed or stalled while compute was starting. The incident was resolved within an hour. The reports do not say that the two incidents had the same root cause. Both were resolved separately and should be treated as two service events rather than one continuous outage.

The assessment

Both incidents stopped work at the point where Databricks had to provide compute. A customer's data, code and notebooks could still be there, but a scheduled job could not run if its cluster remained in Pending or Starting. That makes cluster startup a separate dependency from simply having data stored in the platform.

For short interruptions, retrying a job may be enough. Workloads tied to deadlines need a different plan if they cannot wait for regional compute to recover. That could mean a tested way to run elsewhere or designing jobs so they can resume after a delayed start. The two incidents have not been shown to share a cause, so one technical fix cannot yet be assumed to address both. For BTW readers, teams using managed compute need to know which jobs can wait and which cannot. That decision determines whether simple retries are enough or whether another region needs to be prepared in advance.

What to watch

Watch for separate root-cause information for ES-2194668 and ES-2195009, further Central US compute incidents and any Databricks guidance on retries or regional recovery. For customers, the useful evidence will be whether critical jobs can restart cleanly after an interruption or move to another region when compute startup remains unavailable.