Overview
The role of Incident Manager at Databricks involves leading critical production incidents, ensuring timely communication with customers and stakeholders while coordinating responses across engineering teams. The position combines operational leadership with technical knowledge in a cloud-native environment.
Key Responsibilities
- Lead critical incidents, coordinating multi-disciplinary efforts to mitigate impact and restore operations.
- Drive technical root cause analysis and reliability improvements through collaboration with engineering teams.
- Summarize key learnings, communicate action items, and ensure follow-through on improvements.
- Own communications during incidents, providing updates to stakeholders and customer-facing notifications.
- Mentor and train peers in incident communication and technical response disciplines.
Requirements
- 5+ years of experience in incident management, site reliability engineering, or production operations.
- Proven ability to lead high-severity incidents and manage multi-team responses.
- Strong understanding of cloud infrastructure (AWS, Azure, or GCP).
- Expertise in log analysis and debugging with familiarity in tools such as Datadog, Elasticsearch, or Splunk.
- Hands-on experience with observability systems (e.g., Prometheus, Grafana).
- Proficiency in a major programming or scripting language (e.g., Python, Go, Bash).
- Experience developing and maintaining incident playbooks and communication templates.
- Excellent contextual interpretation and writing skills for technical and business audiences.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field.
Benefits
Databricks offers comprehensive benefits and perks designed to meet the needs of all employees.
Location
Remote - United States
How to Apply
For more information about the application process, please visit the company’s careers page.
Deadline
No specific deadline mentioned.