Overview
The role involves working as a Site Reliability Engineer in the global Platform SRE team at Kong, responsible for building, operating, and scaling the company's multi-region SaaS platform.
Key Responsibilities
- Operate and scale Kong’s global SaaS platform (Konnect) across regions and clouds.
- Build, automate, and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD.
- Design, maintain, and optimize multi-region data and caching layers including PostgreSQL, Redis, ClickHouse, and Druid.
- Improve Kong Gateway and Kong Mesh environments for hybrid and distributed architectures.
- Develop CI/CD pipelines and GitOps workflows for consistent infrastructure changes.
- Enhance observability and incident response with tools like Datadog, Prometheus, Grafana, and Thanos.
- Collaborate with development and security teams ensuring compliance with reliability and security standards.
- Participate in a global 24/7 on-call rotation and improve operational playbooks.
- Lead scaling initiatives to improve elasticity, reliability, and cost-efficiency.
Requirements
- BS in Computer Science or equivalent practical experience.
- Experience managing SaaS or PaaS systems at enterprise scale.
- Deep expertise in Kubernetes.
- Proficiency with Infrastructure as Code tools like Terraform.
- Experience with CI/CD pipelines and GitOps workflows.
- Proficiency in programming languages (Go, Python, Bash).
- Solid understanding of Linux/Unix systems and networking.
- Experience with API gateway and service mesh technologies.
- Familiarity with observability platforms and streaming systems.
- Experience in a 24/7 production support environment.
Benefits
No specific benefits information provided.
Location
Remote - United States
How to Apply
Interested candidates are encouraged to apply through the application process outlined by the company.
Deadline
No specific deadline mentioned.