Jobgether
Jobgether

Senior Site Reliability Engineer — Token Factory (Inference Platform)

engineeringfull-timeGermany
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
    • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
    • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
    • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
    • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
    • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
    • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
    • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
    • Create, maintain, and improve runbooks for incident response and operational procedures.
    • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
    • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
    • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
    • Investigate distributed-system failures and performance issues across infrastructure and application layers.
    • Optimize systems from the kernel and infrastructure layer through to the application layer.
    • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
    • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
    • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
    • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
    • Requirements:

      • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline.
      • Deep practical knowledge of Kubernetes in production environments.
      • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.
      • Advanced experience with Terraform and infrastructure-as-code practices.
      • Strong scripting and automation skills using Python and/or Bash.
      • Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.
      • Experience designing effective alerts, monitoring strategies, and SLOs for high-throughput services or APIs.
      • Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.
      • Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.
      • Hands-on experience with GPU-heavy workloads or accelerator-based infrastructure is highly valuable.
      • Familiarity with GPU inference technologies such as vLLM, Triton, Ray, or comparable accelerator and model-serving stacks.
      • Experience with MLOps, model hosting, AI infrastructure, or machine-learning platforms is advantageous.
      • Strong understanding of infrastructure automation, deployment, configuration management, and operational tooling.
      • Ability to analyze complex performance and reliability problems and translate findings into practical engineering improvements.
      • Strong incident-management and root-cause-analysis capabilities.
      • Ability to collaborate effectively with software engineers and other technical teams to integrate reliability into platform development.
      • Proactive mindset with a strong focus on automation, self-healing systems, and continuous improvement.
      • Comfortable working independently, taking ownership of critical infrastructure, and operating effectively in a fast-paced technical environment.
      • Benefits:

        • Competitive compensation.
        • Career growth and continuous learning opportunities.
        • Flexibility and significant ownership in your work.
        • Collaborative and innovative international working environment.
        • Opportunity to work on high-impact AI infrastructure and inference technologies.
        • Exposure to large-scale GPU infrastructure and complex distributed systems.
        • Opportunity to contribute to infrastructure supporting next-generation multimodal AI applications.
        • Diverse and highly technical international teams.
        • Inclusive workplace committed to equal employment opportunities.
        • Workplace accommodations available throughout the application process where required.
        • Employment is subject to authorization to work in the country where the position is based.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now