Jobgether
Sr. Site Reliability Engineer (Azure, IaC, Distributed Systems)
engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Support the deployment, configuration, and maintenance of monitoring and logging solutions across development, staging, and production environments.
- Maintain and optimize observability technologies, including Splunk, ClickHouse, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, and CloudWatch.
- Automate repetitive operational activities and identify opportunities to improve workflows, efficiency, and system integration.
- Contribute to the setup and maintenance of CI/CD pipelines supporting automated build, testing, and deployment processes.
- Support cloud infrastructure management across AWS and GCP, with a focus on availability, reliability, and security.
- Use Infrastructure as Code tools such as Terraform, Ansible, and CloudFormation to configure and manage environments.
- Support the implementation and administration of containerization technologies, including Docker and Kubernetes.
- Monitor system performance, identify potential issues, and escalate complex incidents appropriately.
- Participate in production incident troubleshooting, root cause analysis, and the implementation of corrective and preventive actions.
- Provide first-level support for infrastructure and deployment issues while collaborating with other technical teams on more complex problems.
- Create and maintain clear documentation covering infrastructure, operational processes, configurations, and procedures.
- Apply DevOps and SRE best practices while contributing to continuous improvements in platform reliability and operational maturity.
- Leverage emerging technologies, including Generative AI tools, to support day-to-day engineering activities and improve productivity.
- Bachelor’s degree with 2+ years of relevant experience, or 5+ years of relevant professional experience without the degree requirement.
- Hands-on experience supporting monitoring and logging tool deployment and configuration in production environments.
- Practical experience with observability platforms such as Splunk, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, ClickHouse, or CloudWatch.
- Experience supporting cloud infrastructure, particularly AWS and/or GCP, with an understanding of availability, security, and operational reliability.
- Experience with Infrastructure as Code technologies such as Terraform, Ansible, and CloudFormation.
- Experience building or maintaining CI/CD pipelines for automated software delivery and deployment.
- Familiarity with Docker and Kubernetes and their use in modern infrastructure environments.
- Experience monitoring system performance and supporting production incident management, troubleshooting, and root cause analysis.
- Understanding of DevOps and Site Reliability Engineering principles and a willingness to continuously develop technical expertise.
- Strong analytical and problem-solving skills, with the ability to investigate issues systematically and identify opportunities for automation.
- Ability to work collaboratively with engineering and infrastructure teams while managing priorities in a dynamic environment.
- Strong documentation and communication skills, with attention to operational detail and process consistency.
- Digital fluency and willingness to incorporate modern tools, including Generative AI solutions, into everyday engineering workflows.
- Fully remote work model in India.
- Full-time position with a standard Monday–Friday work schedule.
- Opportunity to work with modern cloud, observability, automation, containerization, and Infrastructure as Code technologies.
- Exposure to large-scale technology environments and distributed systems.
- Opportunities to develop expertise in DevOps, Site Reliability Engineering, cloud infrastructure, and automation.
- Collaborative environment with opportunities to work across technical teams and global initiatives.
- Opportunity to use emerging technologies, including Generative AI, to improve engineering productivity and operational processes.
- Professional growth through hands-on experience with complex infrastructure and reliability challenges.
- Inclusive workplace and equal-opportunity employment environment.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply