Jobgether
SRE Engineer
engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more
About the role
Accountabilities:
- Define, monitor, and continuously improve reliability indicators, including SLIs, SLOs, SLAs, MTTR, MTTD, and error budgets.
- Implement and evolve observability, monitoring, alerting, and APM solutions across applications and infrastructure.
- Monitor latency, traffic, errors, saturation, availability, and overall system performance.
- Prevent, investigate, and resolve incidents, minimizing their impact on users and business operations.
- Conduct root cause analyses and define corrective and preventive actions to avoid recurring incidents.
- Identify operational risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
- Support the design and evolution of highly available, scalable, resilient, and fault-tolerant solutions.
- Automate operational activities and reduce repetitive manual tasks through scripting, automation, and Infrastructure as Code.
- Operate and continuously improve Kubernetes and Docker environments.
- Support capacity planning, business continuity, disaster recovery, and cloud cost optimization initiatives.
- Participate in deployments and contribute to application stabilization following releases.
- Collaborate with engineering, development, infrastructure, and other technical teams to incorporate reliability from the earliest stages of solution design.
- Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
- Promote a culture centered on reliability, observability, automation, prevention, and continuous improvement.
- Proven professional experience as a Site Reliability Engineer, SRE, or in an equivalent reliability/platform engineering role.
- Practical experience with cloud environments, using one or more of GCP, AWS, or Azure.
- Hands-on knowledge of Kubernetes and Docker.
- Experience implementing and managing observability, monitoring, alerting, and APM solutions.
- Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
- Experience managing, investigating, troubleshooting, and resolving production incidents.
- Knowledge of application and infrastructure troubleshooting in complex environments.
- Experience administering Linux environments.
- Understanding of networking, security, performance, scalability, and high availability.
- Experience with automation and Infrastructure as Code practices.
- Hands-on experience with CI/CD pipelines and modern software delivery practices.
- Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
- Analytical, proactive, collaborative, and prevention-oriented mindset.
- Experience with GKE, EKS, or AKS is a plus.
- Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable observability platforms is a plus.
- Experience with ELK Stack, Elasticsearch, and Kibana is desirable.
- Knowledge of Terraform and Ansible is desirable.
- Experience supporting critical systems and distributed architectures is an advantage.
- Experience in financial institutions or other regulated environments is a plus.
- Experience with cloud capacity management and cost optimization is desirable.
- Knowledge of disaster recovery and business continuity practices is beneficial.
- Cloud, Kubernetes, or SRE certifications are considered a plus.
- Meal and food allowance.
- Home office allowance.
- Medical insurance.
- Dental insurance.
- Life insurance.
- Birthday Day Off.
- TotalPass / Wellhub access.
- Health and wellness support through the Boon Saúde app.
- Discounts and partnerships with a variety of establishments.
- Partnerships with educational institutions and other services.
- Welcome kit.
- Structured onboarding program.
- Access to continuous learning and professional development initiatives.
- Dedicated learning and knowledge-sharing programs.
- Employee support and engagement initiatives.
- Fully remote work opportunity.
Requirements:
Benefits:
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply