Overview
GitLab is an orchestration platform for DevSecOps, focusing on improving developer productivity and operational efficiency. The Site Reliability Engineer role involves maintaining user-facing services and production systems at scale, combining software engineering with operational excellence.
Key Responsibilities
- Maintain user-facing services and production systems for reliability, scalability, and efficiency.
- Build automation and tooling to reduce manual work through infrastructure-as-code-driven workflows.
- Operate and troubleshoot production systems on Kubernetes.
- Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps.
- Participate in on-call, manage alerts, and improve runbooks.
- Contribute to the observability stack to detect symptoms early.
- Take part in incident response and post-incident reviews.
- Document runbooks, architectural decisions, and reviews.
Requirements
- Experience in keeping production systems reliable, with a combination of operations mindset and software engineering practice.
- Ability to build new infrastructure tooling and automation.
- Experience with infrastructure as code, Kubernetes, and cloud providers like GCP or AWS.
- Familiarity with observability practices and metrics.
- Strong written communication skills, with the ability to function in an async, distributed environment.
- Alignment with GitLab's values.
Benefits
- Health, finance, and well-being support.
- Flexible Paid Time Off.
- Equity Compensation & Employee Stock Purchase Plan.
- Growth and Development Fund.
- Parental Leave.
Location
Remote, eligible in Canada, United Kingdom, and the United States.
How to Apply
Interested candidates can apply through the GitLab careers page.