Jobgether
Jobgether

Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems)

engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Architect, implement, and maintain solutions designed to deliver highly available and reliable application services, with a strong focus on resilient distributed systems.
    • Automate monitoring, deployment, operational processes, and incident-response activities to improve reliability, efficiency, and consistency while reducing manual errors.
    • Lead complex troubleshooting and technical investigations, identify root causes, implement corrective actions, and help prevent recurring incidents.
    • Collaborate with global engineering and cross-functional teams to build, deploy, maintain, and continuously improve application services.
    • Drive the modernization of traditional applications toward cloud-native architectures and contribute to key technical and architectural decisions.
    • Design and manage scalable infrastructure using cloud platforms and infrastructure-as-code practices, particularly AWS, Terraform, and Ansible.
    • Implement and optimize observability solutions covering monitoring, logging, alerting, system performance, and service health.
    • Support and optimize ML model deployment pipelines and their associated monitoring systems.
    • Participate in an on-call rotation, including occasional off-hours support, to maintain operational continuity in a 24x7 environment.
    • Develop and maintain documentation for infrastructure, operational procedures, troubleshooting practices, and knowledge transfer across engineering teams.
    • Promote continuous improvement, automation, reliability engineering practices, and effective collaboration within Agile teams.
    • Requirements:

      • Bachelor’s degree with at least 5 years of relevant experience, or an advanced degree with the corresponding professional experience; candidates with extensive equivalent experience may also be considered.
      • Strong professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, distributed systems, or a closely related engineering discipline.
      • Hands-on experience with cloud platforms such as AWS, with additional exposure to Azure or GCP and hybrid cloud architectures.
      • Proven expertise with infrastructure-as-code technologies, particularly Terraform and/or Ansible.
      • Experience designing, deploying, and supporting containerized microservices and high-performance distributed applications.
      • Strong knowledge of observability and monitoring technologies such as Prometheus, Grafana, Datadog, and ELK.
      • Programming or automation experience with languages such as Python, Go, and Java.
      • Familiarity with modern application and messaging technologies such as Spring Boot, SQS, IBM MQ, Kafka-like streaming systems, Flink, Hazelcast, or comparable platforms.
      • Experience with distributed-system design, system architecture, performance optimization, troubleshooting, and reliability engineering.
      • Experience supporting ML model deployment pipelines and associated monitoring is highly valuable.
      • Strong understanding of automation, incident management, cloud-native architecture, and operational best practices.
      • Experience working in Agile teams and fast-paced 24x7 production environments.
      • Strong communication, collaboration, documentation, and knowledge-sharing skills, with the ability to influence technical decisions across teams.
      • Comfortable adopting emerging technologies, including generative AI tools, to improve productivity and everyday engineering workflows.
      • Benefits:

        • Fully remote working arrangement within India, with occasional office presence potentially required with advance notice.
        • Opportunity to work on large-scale, mission-critical distributed systems and globally significant technology platforms.
        • Exposure to modern cloud, infrastructure-as-code, observability, automation, containerization, and ML technologies.
        • Strong opportunities for technical leadership, architectural influence, continuous learning, and professional development.
        • Collaborative environment with global engineering teams and cross-functional stakeholders.
        • Experience working with modern SRE and cloud-native engineering practices in a fast-paced, 24x7 environment.
        • Inclusive workplace focused on equal opportunity, professional growth, and meaningful technical impact.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems) at Jobgether — Remote