Run health, retries, failure patterns and runtime degradation across observed Databricks jobs.
LakeOps detected concentrated reliability risk rather than a workspace-wide outage, giving the team a clear first investigation target.
Success, retries and runtime health by logical job
Failure, retry and runtime signals requiring verification
Inspect recent failed runs and upstream/schema/environment changes.
Identify transient dependency failures or unstable task logic causing retries.
Compare data volume, code, compute configuration, and dependency changes.