Nice
Nice

NOC Engineer / SRE

engineeringfull-timeUnited Kingdom - Remote
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

So, what’s the role all about?

The SRE – NOC role sits at the intersection of traditional Network Operations Center (NOC) responsibilities and engineering‑driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation.

Unlike a traditional NOC analyst, an SRE‑NOC is expected to engineer problems away, not just respond to alerts.

How will you make an impact?

Incident Response & Operations

  • Act as a primary or escalation responder in a 24x7 on‑call rotation
  • Lead or support Major Incident (MI) response, including triage, mitigation, and resolution
  • Coordinate across Engineering, Infrastructure, Security, and Product teams
  • Execute and improve runbooks, playbooks, and escalation paths
  • Drive blameless post‑incident reviews (PIRs) and track corrective actions

Monitoring, Alerting & Observability

  • Own service health monitoring across infrastructure, applications, and dependencies
  • Design and maintain alerting strategies that align with SLIs/SLOs
  • Reduce alert fatigue through signal‑to‑noise improvements
  • Build dashboards using tools such as:
    • Grafana
    • Prometheus
    • Datadog / Splunk / CloudWatch

Reliability Engineering & Automation

  • Automate repetitive operational tasks to reduce manual toil
  • Improve mean time to detect (MTTD) and mean time to resolve (MTTR)
  • Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows
  • Implement self‑healing and auto‑remediation where possible
  • Partner with engineering teams to improve system design for reliability

Platform & Infrastructure Support

  • Support and troubleshoot:
    • Linux‑based systems
    • Cloud platforms (AWS, Azure, GCP)
    • Kubernetes / containerized environments
  • Assist with capacity planning and availability reviews
  • Ensure operational readiness for production releases

Have you got what it takes?

Technical

  • Strong Linux systems administration
  • Experience with incident management and production support
  • Familiarity with:
    • Cloud infrastructure (AWS preferred)
    • Containers & orchestration (Docker, Kubernetes)
    • Monitoring/alerting platforms
  • Scripting or programming experience in Python, Bash, Go, or similar
  • Understanding of networking fundamentals (DNS, TCP/IP, load balancing)

Operational

  • Experience working in 24x7 NOC or production operations environments
  • Ability to handle high‑pressure incidents calmly and effectively
  • Strong written and verbal communication for incident coordination
  • Comfort working from runbooks—but improving them when they fall short

Preferred / Differentiators

  • Experience defining or operating to SLOs / SLIs
  • Prior migration from traditional NOC → SRE model
  • Infrastructure as Code experience (Terraform, Ansible, etc.)
  • Exposure to security, compliance, or regulated environments

Requisition ID: 11707.

Reporting into: Manager, Network Operations.

Role Type: Individual Contributor.

#LI-Remote

✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now