Jobgether
Jobgether

Site Reliability Engineer Technical Lead

engineeringfull-timeUS
SALARY
$100k – $150k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities

    • Drive the operational excellence, reliability, scalability, and performance of critical production systems.
    • Apply Site Reliability Engineering principles to enterprise-level systems and continuously identify opportunities to improve resilience and availability.
    • Lead technical aspects of incident response, troubleshooting complex production issues and developing sustainable solutions to prevent recurrence.
    • Design and implement automation that eliminates repetitive manual work, reduces operational toil, and improves engineering efficiency.
    • Develop and maintain internal tools and automated workflows that support scalable, reliable, and self-healing infrastructure.
    • Work across cloud environments such as AWS, GCP, and Azure, as well as microservices and containerized platforms including Kubernetes.
    • Build, maintain, and optimize observability capabilities covering monitoring, alerting, logging, and distributed tracing.
    • Use platforms and tools such as Dynatrace, Splunk, ELK Stack, or comparable technologies to analyze system behavior and identify potential issues proactively.
    • Analyze operational metrics and system performance data to guide performance tuning, capacity planning, and reliability improvements.
    • Identify opportunities to strengthen system architecture, automation, monitoring, and operational processes.
    • Collaborate effectively with engineering teams and senior stakeholders to communicate technical risks, recommendations, and solutions.
    • Provide technical leadership and influence reliability practices without relying primarily on project-management responsibilities.
    • Requirements:

      • 8+ years of IT experience, including significant experience in a senior-level SRE, infrastructure engineering, systems engineering, or closely related role.
      • Deep understanding and hands-on application of SRE principles within enterprise-scale environments.
      • Strong experience with cloud platforms such as AWS, Google Cloud Platform (GCP), and/or Microsoft Azure.
      • Solid expertise with microservices architectures and container orchestration technologies, particularly Kubernetes.
      • Demonstrated ability to write production-quality code, particularly with Python or comparable programming languages.
      • Proven experience designing and implementing automation to reduce manual operational work and create scalable, resilient, and self-healing systems.
      • Experience building and maintaining internal tooling that improves infrastructure and operational processes.
      • Strong expertise in observability, including monitoring, alerting, logging, and distributed tracing.
      • Experience with technologies such as Dynatrace, Splunk, ELK Stack, or similar observability platforms.
      • Strong analytical skills and the ability to interpret system metrics to proactively identify reliability and performance issues.
      • Experience using operational data to support performance optimization and capacity planning.
      • Excellent troubleshooting and problem-solving abilities in complex production environments.
      • Exceptional communication, leadership, and interpersonal skills, with the ability to influence technical and non-technical stakeholders.
      • Ability to collaborate effectively with senior leadership and communicate complex technical concepts clearly.
      • Comfortable working independently in a fully remote environment while taking ownership of mission-critical technical challenges.
      • Work authorization: U.S. Citizens, Green Card holders, EAD holders, and candidates with H-1B transfer eligibility are encouraged to apply. New H-1B sponsorship is not available for this position.
      • Benefits:

        • Fully remote position within the United States.
        • Full-time, direct W-2 employment.
        • Competitive annual salary range of $100,000–$150,000, depending on qualifications and experience.
        • Opportunity to work on mission-critical enterprise systems and advanced cloud technologies.
        • Significant technical ownership and influence over reliability, automation, and operational engineering practices.
        • Opportunity to collaborate with experienced technical teams and senior stakeholders.
        • Career growth potential within a technology-focused environment.
        • Hands-on exposure to cloud platforms, Kubernetes, automation, observability, and modern SRE practices.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Site Reliability Engineer Technical Lead at Jobgether — Remote