Overview
The Site Reliability Engineer (SRE) – NOC role at NiCE combines traditional Network Operations Center responsibilities with engineering-driven reliability practices. This position emphasizes 24/7 service reliability, incident response, operational automation, and observability, aiming to reduce operational toil through software and automation.
Key Responsibilities
- Act as a primary or escalation responder in a 24x7 on-call rotation.
- Lead or support Major Incident (MI) response, including triage, mitigation, and resolution.
- Coordinate with Engineering, Infrastructure, Security, and Product teams.
- Execute and enhance runbooks, playbooks, and escalation paths.
- Conduct blameless post-incident reviews (PIRs) and track corrective actions.
- Own service health monitoring across infrastructure, applications, and dependencies.
- Design and maintain alerting strategies that align with SLIs/SLOs.
- Minimize alert fatigue through improved signal-to-noise ratios.
- Build dashboards with tools such as Grafana, Prometheus, Datadog, Splunk, or CloudWatch.
- Automate repetitive operational tasks to decrease manual toil.
- Enhance mean time to detect (MTTD) and mean time to resolve (MTTR).
- Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows.
- Implement self-healing and auto-remediation where feasible.
- Collaborate with engineering teams to enhance system design for reliability.
- Support and troubleshoot Linux-based systems and cloud platforms (AWS, Azure, GCP).
- Assist with capacity planning and availability reviews.
- Ensure operational readiness for production releases.
Requirements
- Strong Linux systems administration skills.
- Experience with incident management and production support.
- Familiarity with cloud infrastructure, preferably AWS.
- Experience with containers and orchestration (Docker, Kubernetes).
- Knowledge of monitoring/alerting platforms.
- Proficiency in scripting or programming (Python, Bash, Go, etc.).
- Understanding of networking fundamentals (DNS, TCP/IP, load balancing).
- Experience in 24x7 NOC or production operations environments.
- Ability to calmly handle high-pressure incidents.
- Strong written and verbal communication skills for incident coordination.
- Comfortable working from runbooks and improving them when necessary.
- Preferred: Experience with defining or operating to SLOs/SLIs.
- Prior experience migrating from traditional NOC to SRE model.
- Infrastructure as Code experience (Terraform, Ansible, etc.).
- Exposure to security, compliance, or regulated environments.
Benefits
Details of the benefits package are not provided in the listing.
Location
Remote role based in the United Kingdom.
How to Apply
Application instructions are not provided in the listing.
Deadline
No application deadline is mentioned.