Jobgether
Jobgether

SRE Engineer

engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities:

    • Define, monitor, and continuously improve reliability indicators, including SLIs, SLOs, SLAs, MTTR, MTTD, and error budgets.
    • Implement and evolve observability, monitoring, alerting, and APM solutions across applications and infrastructure.
    • Monitor latency, traffic, errors, saturation, availability, and overall system performance.
    • Prevent, investigate, and resolve incidents, minimizing their impact on users and business operations.
    • Conduct root cause analyses and define corrective and preventive actions to avoid recurring incidents.
    • Identify operational risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
    • Support the design and evolution of highly available, scalable, resilient, and fault-tolerant solutions.
    • Automate operational activities and reduce repetitive manual tasks through scripting, automation, and Infrastructure as Code.
    • Operate and continuously improve Kubernetes and Docker environments.
    • Support capacity planning, business continuity, disaster recovery, and cloud cost optimization initiatives.
    • Participate in deployments and contribute to application stabilization following releases.
    • Collaborate with engineering, development, infrastructure, and other technical teams to incorporate reliability from the earliest stages of solution design.
    • Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
    • Promote a culture centered on reliability, observability, automation, prevention, and continuous improvement.
    • Requirements:

      • Proven professional experience as a Site Reliability Engineer, SRE, or in an equivalent reliability/platform engineering role.
      • Practical experience with cloud environments, using one or more of GCP, AWS, or Azure.
      • Hands-on knowledge of Kubernetes and Docker.
      • Experience implementing and managing observability, monitoring, alerting, and APM solutions.
      • Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
      • Experience managing, investigating, troubleshooting, and resolving production incidents.
      • Knowledge of application and infrastructure troubleshooting in complex environments.
      • Experience administering Linux environments.
      • Understanding of networking, security, performance, scalability, and high availability.
      • Experience with automation and Infrastructure as Code practices.
      • Hands-on experience with CI/CD pipelines and modern software delivery practices.
      • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
      • Analytical, proactive, collaborative, and prevention-oriented mindset.
      • Experience with GKE, EKS, or AKS is a plus.
      • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable observability platforms is a plus.
      • Experience with ELK Stack, Elasticsearch, and Kibana is desirable.
      • Knowledge of Terraform and Ansible is desirable.
      • Experience supporting critical systems and distributed architectures is an advantage.
      • Experience in financial institutions or other regulated environments is a plus.
      • Experience with cloud capacity management and cost optimization is desirable.
      • Knowledge of disaster recovery and business continuity practices is beneficial.
      • Cloud, Kubernetes, or SRE certifications are considered a plus.
      • Benefits:

        • Meal and food allowance.
        • Home office allowance.
        • Medical insurance.
        • Dental insurance.
        • Life insurance.
        • Birthday Day Off.
        • TotalPass / Wellhub access.
        • Health and wellness support through the Boon Saúde app.
        • Discounts and partnerships with a variety of establishments.
        • Partnerships with educational institutions and other services.
        • Welcome kit.
        • Structured onboarding program.
        • Access to continuous learning and professional development initiatives.
        • Dedicated learning and knowledge-sharing programs.
        • Employee support and engagement initiatives.
        • Fully remote work opportunity.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now
SRE Engineer at Jobgether — Remote