Jobgether
Jobgether

Staff Site Reliability Engineer

engineeringfull-timeGermany
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities

    • Architect, deploy, operate, and continuously improve scalable, secure production environments, with a strong preference for AWS-based infrastructure.
    • Lead reliability initiatives across multiple engineering streams and establish consistent SRE practices throughout the organization.
    • Design, evolve, migrate, and optimize Kubernetes-based infrastructure, including production hardening and scaling.
    • Establish and enforce robust Infrastructure-as-Code standards using Terraform or equivalent technologies.
    • Define, implement, and operationalize SLIs, SLOs, error budgets, and other reliability practices.
    • Strengthen observability across applications, infrastructure, data pipelines, and ML systems to improve visibility into system health and performance.
    • Collaborate with product and data teams to incorporate product telemetry, model analytics, and operational data into reliability insights.
    • Design and optimize CI/CD pipelines across the complete software lifecycle, from build and testing through deployment and rollback.
    • Improve release safety, deployment frequency, operational predictability, and adherence to service-level objectives.
    • Lead incident response for complex, cross-system failures and drive thorough post-incident reviews and corrective actions.
    • Reduce operational toil through automation, platform engineering, and improved tooling and processes.
    • Design scalable processes and infrastructure for absorbing, standardizing, monitoring, and troubleshooting customer environments.
    • Support and productionize ML workloads by implementing MLOps practices for model deployment, monitoring, and retraining workflows.
    • Ensure infrastructure and operational practices meet enterprise-grade security, compliance, and regulatory requirements.
    • Mentor engineers, share best practices, and raise the overall reliability and engineering standards across teams.
    • Collaborate with Staff Engineers and Architects to influence global product architecture and long-term technology strategy.
    • Requirements

      • Extensive hands-on experience in Site Reliability Engineering, Production Engineering, or a closely related infrastructure role.
      • Proven experience establishing or scaling SRE practices within high-growth, complex, or highly distributed technical environments.
      • Deep expertise with AWS or Azure cloud infrastructure and modern cloud-native architectures.
      • Strong production experience with Kubernetes, including migration, scaling, optimization, and security hardening.
      • Advanced Infrastructure-as-Code expertise using Terraform or an equivalent technology.
      • Demonstrated experience designing, implementing, and optimizing end-to-end CI/CD pipelines.
      • Strong knowledge of observability practices and tooling across distributed applications and infrastructure.
      • Experience troubleshooting complex multi-tenant, customer-hosted, or enterprise environments.
      • Experience supporting production data platforms and machine learning systems.
      • Practical MLOps experience, including model deployment, monitoring, and operational lifecycle management.
      • Strong understanding of distributed systems, scalability, resilience, fault tolerance, and failure modes.
      • Ability to think across systems and understand the interactions between infrastructure, applications, data, and ML workloads.
      • Strong communication and collaboration skills, with the ability to work effectively across engineering and business functions.
      • Experience with large-scale global B2B or B2C products is desirable.
      • Experience working with AI/ML, NLP, or LLM-based products is a strong advantage.
      • Familiarity with integrating product analytics and model performance metrics into operational monitoring is beneficial.
      • Experience operating in enterprise environments with stringent security, compliance, and regulatory requirements is preferred.
      • Experience implementing regulatory controls within cloud infrastructure is a plus.
      • Experience scaling infrastructure during periods of rapid growth is advantageous.
      • Experience evaluating infrastructure tools, platforms, and vendors is desirable.
      • Experience deploying and operating solutions within large enterprise customer accounts or VPCs is a strong plus.
      • Strong problem-solving skills, high ownership, and accountability.
      • Ability to anticipate failure modes, operate across multiple engineering streams, and influence technical decisions without formal authority.
      • Continuous-learning mindset with a strong commitment to improving systems, processes, and engineering practices.
      • Benefits

        • Full-time, permanent employment.
        • Fully remote position within European time zones.
        • Opportunity to work on infrastructure supporting advanced AI and agent-based workloads.
        • Significant technical ownership and influence over reliability practices and architecture.
        • Opportunity to work across cloud infrastructure, Kubernetes, distributed systems, data platforms, and ML operations.
        • Exposure to complex enterprise environments and large-scale production deployments.
        • Collaboration with highly experienced engineers, architects, product teams, and AI/ML specialists.
        • Opportunity to shape reliability standards and technology strategy as the organization scales.
        • Strong culture of ownership, continuous improvement, and technical excellence.
        • International and distributed working environment.
        • Career growth and opportunities to expand technical leadership and organizational impact.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Staff Site Reliability Engineer at Jobgether — Remote