Jobgether
Site Reliability Engineer Technical Lead
engineeringfull-timeUS
SALARY
$100k – $150k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities
- Drive the operational excellence, reliability, scalability, and performance of critical production systems.
- Apply Site Reliability Engineering principles to enterprise-level systems and continuously identify opportunities to improve resilience and availability.
- Lead technical aspects of incident response, troubleshooting complex production issues and developing sustainable solutions to prevent recurrence.
- Design and implement automation that eliminates repetitive manual work, reduces operational toil, and improves engineering efficiency.
- Develop and maintain internal tools and automated workflows that support scalable, reliable, and self-healing infrastructure.
- Work across cloud environments such as AWS, GCP, and Azure, as well as microservices and containerized platforms including Kubernetes.
- Build, maintain, and optimize observability capabilities covering monitoring, alerting, logging, and distributed tracing.
- Use platforms and tools such as Dynatrace, Splunk, ELK Stack, or comparable technologies to analyze system behavior and identify potential issues proactively.
- Analyze operational metrics and system performance data to guide performance tuning, capacity planning, and reliability improvements.
- Identify opportunities to strengthen system architecture, automation, monitoring, and operational processes.
- Collaborate effectively with engineering teams and senior stakeholders to communicate technical risks, recommendations, and solutions.
- Provide technical leadership and influence reliability practices without relying primarily on project-management responsibilities.
- 8+ years of IT experience, including significant experience in a senior-level SRE, infrastructure engineering, systems engineering, or closely related role.
- Deep understanding and hands-on application of SRE principles within enterprise-scale environments.
- Strong experience with cloud platforms such as AWS, Google Cloud Platform (GCP), and/or Microsoft Azure.
- Solid expertise with microservices architectures and container orchestration technologies, particularly Kubernetes.
- Demonstrated ability to write production-quality code, particularly with Python or comparable programming languages.
- Proven experience designing and implementing automation to reduce manual operational work and create scalable, resilient, and self-healing systems.
- Experience building and maintaining internal tooling that improves infrastructure and operational processes.
- Strong expertise in observability, including monitoring, alerting, logging, and distributed tracing.
- Experience with technologies such as Dynatrace, Splunk, ELK Stack, or similar observability platforms.
- Strong analytical skills and the ability to interpret system metrics to proactively identify reliability and performance issues.
- Experience using operational data to support performance optimization and capacity planning.
- Excellent troubleshooting and problem-solving abilities in complex production environments.
- Exceptional communication, leadership, and interpersonal skills, with the ability to influence technical and non-technical stakeholders.
- Ability to collaborate effectively with senior leadership and communicate complex technical concepts clearly.
- Comfortable working independently in a fully remote environment while taking ownership of mission-critical technical challenges.
- Work authorization: U.S. Citizens, Green Card holders, EAD holders, and candidates with H-1B transfer eligibility are encouraged to apply. New H-1B sponsorship is not available for this position.
- Fully remote position within the United States.
- Full-time, direct W-2 employment.
- Competitive annual salary range of $100,000–$150,000, depending on qualifications and experience.
- Opportunity to work on mission-critical enterprise systems and advanced cloud technologies.
- Significant technical ownership and influence over reliability, automation, and operational engineering practices.
- Opportunity to collaborate with experienced technical teams and senior stakeholders.
- Career growth potential within a technology-focused environment.
- Hands-on exposure to cloud platforms, Kubernetes, automation, observability, and modern SRE practices.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply