Overview
The role involves designing, building, and maintaining the multi-cloud platform for hosting Elastic's internal and external services, contributing to the automation and reliability of the global Elastic infrastructure.
Key Responsibilities
- Lead technical initiatives to automate system engineering efforts ensuring infrastructure reliability.
- Develop and maintain software and tools to scale the global Platform infrastructure.
- Foster a collaborative environment focused on operational excellence.
- Manage incidents and prioritize problem management to enhance customer experience.
Requirements
- Experience in a customer-oriented approach to operational problems from an SRE perspective.
- Background in software engineering with an understanding of public cloud and managed Kubernetes services.
- Strong communication skills and experience in distributed teams or remote work.
Bonus Points
- Experience operating a SaaS product in a public cloud with Infrastructure-as-Code tooling.
- Experience with Kubernetes-at-scale infrastructure.
- Proficiency in programming languages such as Golang.
- Familiarity with containerized services like Docker.
- Experience in incident management processes and Linux system administration.
- Experience with the Elastic Stack.
- Ability to mentor and coach team members.
Benefits
- Competitive compensation based on market benchmarks.
- Health coverage for you and your family.
- Flexible work locations and schedules.
- Generous vacation days.
- Financial match for donations and volunteer hours.
- Parental leave of at least 16 weeks.
Location
Remote - this role is based in Portugal.
How to Apply
Applicants can apply through the company's career portal or by sending an email to candidate_accessibility@elastic.co.