Jobgether
Jobgether

Sr. Site Reliability Engineer (Azure, IaC, Distributed Systems)

engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Support the deployment, configuration, and maintenance of monitoring and logging solutions across development, staging, and production environments.
    • Maintain and optimize observability technologies, including Splunk, ClickHouse, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, and CloudWatch.
    • Automate repetitive operational activities and identify opportunities to improve workflows, efficiency, and system integration.
    • Contribute to the setup and maintenance of CI/CD pipelines supporting automated build, testing, and deployment processes.
    • Support cloud infrastructure management across AWS and GCP, with a focus on availability, reliability, and security.
    • Use Infrastructure as Code tools such as Terraform, Ansible, and CloudFormation to configure and manage environments.
    • Support the implementation and administration of containerization technologies, including Docker and Kubernetes.
    • Monitor system performance, identify potential issues, and escalate complex incidents appropriately.
    • Participate in production incident troubleshooting, root cause analysis, and the implementation of corrective and preventive actions.
    • Provide first-level support for infrastructure and deployment issues while collaborating with other technical teams on more complex problems.
    • Create and maintain clear documentation covering infrastructure, operational processes, configurations, and procedures.
    • Apply DevOps and SRE best practices while contributing to continuous improvements in platform reliability and operational maturity.
    • Leverage emerging technologies, including Generative AI tools, to support day-to-day engineering activities and improve productivity.
    • Requirements:

      • Bachelor’s degree with 2+ years of relevant experience, or 5+ years of relevant professional experience without the degree requirement.
      • Hands-on experience supporting monitoring and logging tool deployment and configuration in production environments.
      • Practical experience with observability platforms such as Splunk, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, ClickHouse, or CloudWatch.
      • Experience supporting cloud infrastructure, particularly AWS and/or GCP, with an understanding of availability, security, and operational reliability.
      • Experience with Infrastructure as Code technologies such as Terraform, Ansible, and CloudFormation.
      • Experience building or maintaining CI/CD pipelines for automated software delivery and deployment.
      • Familiarity with Docker and Kubernetes and their use in modern infrastructure environments.
      • Experience monitoring system performance and supporting production incident management, troubleshooting, and root cause analysis.
      • Understanding of DevOps and Site Reliability Engineering principles and a willingness to continuously develop technical expertise.
      • Strong analytical and problem-solving skills, with the ability to investigate issues systematically and identify opportunities for automation.
      • Ability to work collaboratively with engineering and infrastructure teams while managing priorities in a dynamic environment.
      • Strong documentation and communication skills, with attention to operational detail and process consistency.
      • Digital fluency and willingness to incorporate modern tools, including Generative AI solutions, into everyday engineering workflows.
      • Benefits:

        • Fully remote work model in India.
        • Full-time position with a standard Monday–Friday work schedule.
        • Opportunity to work with modern cloud, observability, automation, containerization, and Infrastructure as Code technologies.
        • Exposure to large-scale technology environments and distributed systems.
        • Opportunities to develop expertise in DevOps, Site Reliability Engineering, cloud infrastructure, and automation.
        • Collaborative environment with opportunities to work across technical teams and global initiatives.
        • Opportunity to use emerging technologies, including Generative AI, to improve engineering productivity and operational processes.
        • Professional growth through hands-on experience with complex infrastructure and reliability challenges.
        • Inclusive workplace and equal-opportunity employment environment.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Sr. Site Reliability Engineer (Azure, IaC, Distributed Systems) at Jobgether — Remote