Jobgether
Jobgether

Senior Site Reliability Engineer (SRE, Compute Node Team)

engineeringfull-timeIreland
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.
    • Analyze and debug complex Linux systems across both user space and kernel space.
    • Investigate system capabilities, limitations, dependencies, and trade-offs across different layers of the operating system and infrastructure stack.
    • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling.
    • Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.
    • Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.
    • Support and improve containerized workloads using Linux-native mechanisms such as namespaces and cgroups.
    • Design and evolve observability for the compute node layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
    • Build reliability signals that provide clear and actionable insight into system behavior.
    • Lead or contribute to incident response, ensuring production issues are diagnosed and resolved efficiently.
    • Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.
    • Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.
    • Identify opportunities to automate operational processes and improve the resilience of compute infrastructure.
    • Collaborate closely with platform, kernel/hypervisor, GPU, and infrastructure teams on system design and operational improvements.
    • Contribute to improving the operability, scalability, and maintainability of node-level services.
    • Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.
    • Help establish reliability and observability as core capabilities of the compute platform.
    • Requirements:

      • Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
      • Deep expertise in Linux, including strong understanding of both user space and kernel space.
      • Knowledge of important Linux kernel subsystems, including scheduling, memory management, filesystems, cgroups, and namespaces.
      • Strong understanding of system boundaries, constraints, dependencies, and trade-offs across different infrastructure layers.
      • Hands-on experience with QEMU/KVM and a solid understanding of virtualization technologies.
      • Understanding of virtual machine lifecycles, performance characteristics, resource management, and failure modes.
      • Practical experience with containers, Linux namespaces, and cgroups.
      • Strong understanding of resource isolation, allocation, and control in containerized environments.
      • Excellent debugging skills and the ability to reason systematically about complex system failures.
      • Structured, hypothesis-driven approach to incident investigation and troubleshooting.
      • Strong understanding of the SRE discipline, including the relationship between software engineering, operations, reliability, and system design.
      • Experience building and operating observability stacks, rather than simply consuming existing monitoring dashboards.
      • Ability to translate complex system behavior into actionable reliability signals, alerts, SLIs, and SLOs.
      • Strong analytical and problem-solving skills, with the ability to investigate issues across operating-system and infrastructure layers.
      • Experience operating production systems and responding effectively to reliability and performance incidents.
      • Strong communication and collaboration skills when working with multidisciplinary infrastructure and engineering teams.
      • Ability to take ownership of complex technical problems and drive them through investigation, resolution, and long-term improvement.
      • Experience with Kubernetes internals or node-level components is an advantage.
      • Hands-on experience with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash dumps is beneficial.
      • Familiarity with large-scale compute or bare-metal infrastructure is a plus.
      • Contributions to open-source infrastructure or systems software are advantageous.
      • Experience debugging hardware- and driver-level issues, including GPUs, NVLink, or InfiniBand, is a strong plus.
      • Benefits:

        • Competitive compensation.
        • Career growth and continuous learning opportunities.
        • Flexibility and significant ownership in your work.
        • Collaborative and innovative engineering environment.
        • Opportunity to work on impactful AI and cloud infrastructure projects.
        • Exposure to large-scale compute, Linux systems, virtualization, and distributed infrastructure.
        • Opportunity to collaborate with highly skilled international engineering teams.
        • Inclusive workplace committed to equal employment opportunities.
        • Workplace accommodations available during the application process where required.
        • Employment is subject to authorization to work in the country where the position is based.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now
Senior Site Reliability Engineer (SRE, Compute Node Team) at Jobgether — Remote