Overview
Twilio is seeking a Staff Site Reliability Engineer to join the Platform Engineering organization. This remote role focuses on owning the production health and resiliency, ensuring services remain available, performant, and recoverable for customers.
Key Responsibilities
- Own the reliability posture of production services, including availability, latency, and performance.
- Define, instrument, and operate against SLIs and SLOs, using error budgets to guide engineering priorities.
- Identify trends threatening stability and provide mitigation strategies.
- Drive down repair items and prevent incident recurrence.
- Improve detection, response, and recovery processes.
- Design systems for failure and validate recovery paths.
- Participate in on-call duties and lead response during production degradation.
- Write post-mortems identifying root causes and drive follow-up work.
- Write and deploy code to improve service reliability.
- Lead debugging and troubleshooting efforts across system architectures.
- Facilitate cross-team collaboration to ensure timely project delivery.
Requirements
- 8+ years of engineering experience with a focus on reliability, infrastructure, or platform engineering.
- Accountability for production systems and incident management experience.
- Strong software engineering fundamentals with experience building production code.
- Experience defining and operating SLIs and SLOs, using error budgets.
- Strong background in incident command, post-mortem analysis, and observability.
- Ability to drive cross-team changes and build alignment.
- Experience mentoring and improving engineers through code reviews.
- Experience with large-scale distributed systems in cloud environments.
Benefits
Twilio offers competitive pay, generous time off, wellness leave, healthcare, retirement savings programs, and more, varying by location.
Location
Remote - Ireland
How to Apply
If interested, please submit your application through the provided channels.