Jobgether
Jobgether

Site Reliability Engineer - AWS

engineeringfull-timeUS
SALARY
$110k – $140k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Support and continuously improve production applications running in AWS, ensuring high availability, reliability, performance, and operational stability.
    • Monitor system health, availability, latency, performance, logs, metrics, traces, and alerts, proactively identifying and addressing reliability risks.
    • Participate in incident response, troubleshooting, escalation management, root cause analysis, and post-incident improvement activities.
    • Strengthen operational readiness, resiliency, disaster recovery capabilities, and overall production support processes.
    • Engineer and optimize AWS environments using services such as EC2, ECS/EKS, Lambda, S3, RDS, CloudWatch, IAM, VPC, Elastic Beanstalk, Load Balancers, and related cloud technologies.
    • Apply AWS Well-Architected Framework principles across reliability, security, performance efficiency, cost optimization, and operational excellence.
    • Support cloud-native architecture decisions and contribute to infrastructure modernization and optimization initiatives.
    • Build and enhance end-to-end observability through dashboards, monitoring, alerting, logging, metrics, and tracing.
    • Automate recurring operational processes using Python, Bash, PowerShell, CI/CD pipelines, Terraform, and infrastructure-as-code practices.
    • Support applications across both on-premises and AWS environments, contributing to modernization and migration initiatives.
    • Partner with application development teams to identify performance bottlenecks, infrastructure constraints, security concerns, and reliability risks.
    • Manage and improve containerized environments using Kubernetes, Docker, and Amazon EKS, including GitOps-based deployment approaches.
    • Support CI/CD and release management processes using tools and methodologies such as Git, Azure DevOps, ArgoCD, and FluxCD.
    • Troubleshoot APIs, microservices, network connectivity, HTTP-based applications, and cloud infrastructure to improve application performance and uptime.
    • Manage and optimize database environments, including RDS configuration, schemas, users, performance troubleshooting, data integrity, and storage practices.
    • Collaborate across Development, Business, Platform, and Infrastructure teams in Agile environments using Scrum and Kanban practices.
    • Serve as a technical point of contact for infrastructure, reliability, and operational requirements within product teams.
    • Continuously identify opportunities for reengineering, process improvement, efficiency gains, automation, and cloud optimization.
    • Requirements

      • Bachelor's degree or an equivalent combination of education and professional experience.
      • 5+ years of relevant industry or technical experience, with substantial hands-on experience in DevOps, cloud infrastructure, site reliability, or related engineering disciplines.
      • Strong DevOps background, ideally with an emphasis on approximately 70% operations and 30% development activities.
      • Extensive experience with AWS cloud infrastructure and cloud-based production environments.
      • Strong experience with containerization and orchestration technologies, particularly Kubernetes, Docker, and Amazon EKS.
      • Hands-on experience building and managing CI/CD pipelines using GitOps methodologies and tools such as ArgoCD or FluxCD.
      • Strong knowledge of Git-based version control and experience with Azure DevOps, including tickets, releases, CI/CD pipelines, and infrastructure-as-code workflows.
      • Strong Terraform experience and knowledge of infrastructure-as-code best practices.
      • Advanced scripting and automation skills using Python, Bash, and/or PowerShell.
      • Experience with monitoring and observability tools, particularly AWS CloudWatch, including alert configuration, troubleshooting, and escalation management.
      • Strong understanding of AWS services including EC2, AMIs, S3, RDS, Lambda, ECS/EKS, IAM, VPC, Elastic Beanstalk, Load Balancers, and Transfer Family.
      • Experience managing large-file transfers and supporting highly available cloud infrastructure.
      • Strong knowledge of APIs and microservices, including configuration, performance tuning, security, and reliability best practices.
      • Solid understanding of HTTP concepts and protocols, with the ability to analyze and optimize web application performance.
      • Experience using diagnostic tools such as curl and wget to troubleshoot connectivity and network performance issues.
      • Strong SQL and database administration skills, including RDS configuration, schemas, catalogs, users/logins, synonyms, performance troubleshooting, and database hygiene.
      • Strong networking and network management capabilities.
      • Experience with service mesh technologies such as Linkerd or Istio is a plus.
      • Excellent written and verbal communication skills, with the ability to collaborate effectively across technical and business teams.
      • Comfortable working independently in a remote environment and adapting to flexible work hours.
      • Strong troubleshooting, analytical, problem-solving, and continuous-improvement mindset.
      • Benefits

        • Annual full-time base salary range of $110,000–$140,000, with final compensation determined by experience, education, skills, training, geographic location, and market considerations.
        • Potential eligibility for a discretionary bonus based on applicable bonus program guidelines and position eligibility.
        • Fully remote work arrangement.
        • Paid time off and paid holidays in accordance with applicable company policies.
        • Medical, dental, and vision insurance options.
        • Life insurance and disability coverage.
        • 401(k) retirement plan.
        • Opportunity to work on cloud modernization, SaaS reliability, DevOps, automation, and large-scale AWS environments.
        • Exposure to a broad technology stack spanning AWS, Kubernetes, Terraform, GitOps, CI/CD, observability, networking, and database technologies.
        • Collaborative environment with opportunities to work closely with development, platform, infrastructure, and business teams.
        • Flexible work environment designed to support distributed teams and evolving work requirements.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Site Reliability Engineer - AWS at Jobgether — Remote