Jobgether
Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems)
engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Architect, implement, and maintain solutions designed to deliver highly available and reliable application services, with a strong focus on resilient distributed systems.
- Automate monitoring, deployment, operational processes, and incident-response activities to improve reliability, efficiency, and consistency while reducing manual errors.
- Lead complex troubleshooting and technical investigations, identify root causes, implement corrective actions, and help prevent recurring incidents.
- Collaborate with global engineering and cross-functional teams to build, deploy, maintain, and continuously improve application services.
- Drive the modernization of traditional applications toward cloud-native architectures and contribute to key technical and architectural decisions.
- Design and manage scalable infrastructure using cloud platforms and infrastructure-as-code practices, particularly AWS, Terraform, and Ansible.
- Implement and optimize observability solutions covering monitoring, logging, alerting, system performance, and service health.
- Support and optimize ML model deployment pipelines and their associated monitoring systems.
- Participate in an on-call rotation, including occasional off-hours support, to maintain operational continuity in a 24x7 environment.
- Develop and maintain documentation for infrastructure, operational procedures, troubleshooting practices, and knowledge transfer across engineering teams.
- Promote continuous improvement, automation, reliability engineering practices, and effective collaboration within Agile teams.
- Bachelor’s degree with at least 5 years of relevant experience, or an advanced degree with the corresponding professional experience; candidates with extensive equivalent experience may also be considered.
- Strong professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, distributed systems, or a closely related engineering discipline.
- Hands-on experience with cloud platforms such as AWS, with additional exposure to Azure or GCP and hybrid cloud architectures.
- Proven expertise with infrastructure-as-code technologies, particularly Terraform and/or Ansible.
- Experience designing, deploying, and supporting containerized microservices and high-performance distributed applications.
- Strong knowledge of observability and monitoring technologies such as Prometheus, Grafana, Datadog, and ELK.
- Programming or automation experience with languages such as Python, Go, and Java.
- Familiarity with modern application and messaging technologies such as Spring Boot, SQS, IBM MQ, Kafka-like streaming systems, Flink, Hazelcast, or comparable platforms.
- Experience with distributed-system design, system architecture, performance optimization, troubleshooting, and reliability engineering.
- Experience supporting ML model deployment pipelines and associated monitoring is highly valuable.
- Strong understanding of automation, incident management, cloud-native architecture, and operational best practices.
- Experience working in Agile teams and fast-paced 24x7 production environments.
- Strong communication, collaboration, documentation, and knowledge-sharing skills, with the ability to influence technical decisions across teams.
- Comfortable adopting emerging technologies, including generative AI tools, to improve productivity and everyday engineering workflows.
- Fully remote working arrangement within India, with occasional office presence potentially required with advance notice.
- Opportunity to work on large-scale, mission-critical distributed systems and globally significant technology platforms.
- Exposure to modern cloud, infrastructure-as-code, observability, automation, containerization, and ML technologies.
- Strong opportunities for technical leadership, architectural influence, continuous learning, and professional development.
- Collaborative environment with global engineering teams and cross-functional stakeholders.
- Experience working with modern SRE and cloud-native engineering practices in a fast-paced, 24x7 environment.
- Inclusive workplace focused on equal opportunity, professional growth, and meaningful technical impact.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply