Jobgether
Analista de Observabilidade Sênior - Vaga Afirmativa para Mulheres
engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Act as the technical and functional reference for Datadog across the organization, supporting its implementation, administration, governance, and continuous evolution.
- Design, implement, and improve observability solutions using Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Real User Monitoring (RUM), Synthetic Monitoring, and Continuous Testing.
- Define and maintain enterprise observability standards for applications, APIs, microservices, and cloud workloads.
- Build and maintain executive, operational, and analytical dashboards to provide visibility into platform health and performance.
- Develop proactive monitoring strategies based on SLIs, SLOs, SLAs, and relevant business indicators.
- Create, review, and optimize intelligent monitors and alerts to reduce operational noise, false positives, and unnecessary escalations.
- Support development teams with application instrumentation using OpenTelemetry and native Datadog integrations.
- Lead root-cause analysis (RCA) for critical incidents and recommend structural improvements to prevent recurrence.
- Identify opportunities for automation, predictive failure detection, automated incident response, and self-healing.
- Conduct periodic assessments of capacity, performance, availability, reliability, and end-user experience.
- Develop training materials, playbooks, standards, and technical documentation related to observability and monitoring practices.
- Promote a culture of observability, reliability, and operational excellence across technical teams.
- Translate technical indicators and operational events into business impact and actionable recommendations.
- Influence engineering practices and reliability standards across multiple teams and contribute to the organization's broader observability maturity.
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
- Advanced hands-on experience with Datadog, including implementation, administration, configuration, and platform evolution.
- Strong practical experience with Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Monitors, and Service Catalog.
- Advanced understanding of the core observability pillars: logs, metrics, and distributed tracing.
- Experience monitoring APIs, microservices, and distributed architectures.
- Professional experience working with AWS and cloud environments.
- Strong troubleshooting and complex incident investigation capabilities.
- Experience with OpenTelemetry and application instrumentation.
- Knowledge of automation using Python, Shell Script, or PowerShell.
- Experience with Kubernetes, Docker, and cloud-native ecosystems.
- Solid understanding of system availability, performance, scalability, and reliability.
- Ability to connect technical metrics and operational events with business outcomes.
- Excellent communication skills and the ability to collaborate with multiple technical and business stakeholders.
- Investigative, analytical, and problem-solving mindset with a strong focus on continuous improvement.
- Strong sense of ownership and operational responsibility.
- Ability to influence technical teams and encourage the adoption of observability and reliability best practices.
- Collaborative and consultative approach, with the ability to train and enable other teams.
- Experience with SRE practices is desirable.
- Datadog Certified Associate certification or higher is a plus.
- Experience designing observability strategies for large-scale distributed environments is an advantage.
- Knowledge of CI/CD and DevSecOps practices is desirable.
- Experience with automated incident response and self-healing processes is a plus.
- Familiarity with ITIL, incident management, Problem Management, and operational governance is desirable.
- Experience with tools such as Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix is an advantage.
- Senior-level opportunity with significant influence over enterprise observability and reliability strategy.
- High-visibility work impacting multiple engineering and technology teams.
- Opportunity to lead automation, intelligent observability, and incident-reduction initiatives.
- Professional development through exposure to advanced cloud-native technologies and modern observability practices.
- Opportunity to work with Datadog, OpenTelemetry, AWS, Kubernetes, Docker, and related technologies.
- Collaborative environment involving Engineering, Architecture, Development, SRE, and Operations teams.
- Inclusive workplace culture focused on diversity, equity, and professional growth.
- Position specifically designed as an affirmative opportunity for women, supporting greater gender equity in technology leadership.
- Access to initiatives and communities focused on women's development, leadership, networking, and career advancement.
- Remote work arrangement in Brazil.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply