Jobgether
Jobgether

Analista de Observabilidade Sênior - Vaga Afirmativa para Mulheres

engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Act as the technical and functional reference for Datadog across the organization, supporting its implementation, administration, governance, and continuous evolution.
    • Design, implement, and improve observability solutions using Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Real User Monitoring (RUM), Synthetic Monitoring, and Continuous Testing.
    • Define and maintain enterprise observability standards for applications, APIs, microservices, and cloud workloads.
    • Build and maintain executive, operational, and analytical dashboards to provide visibility into platform health and performance.
    • Develop proactive monitoring strategies based on SLIs, SLOs, SLAs, and relevant business indicators.
    • Create, review, and optimize intelligent monitors and alerts to reduce operational noise, false positives, and unnecessary escalations.
    • Support development teams with application instrumentation using OpenTelemetry and native Datadog integrations.
    • Lead root-cause analysis (RCA) for critical incidents and recommend structural improvements to prevent recurrence.
    • Identify opportunities for automation, predictive failure detection, automated incident response, and self-healing.
    • Conduct periodic assessments of capacity, performance, availability, reliability, and end-user experience.
    • Develop training materials, playbooks, standards, and technical documentation related to observability and monitoring practices.
    • Promote a culture of observability, reliability, and operational excellence across technical teams.
    • Translate technical indicators and operational events into business impact and actionable recommendations.
    • Influence engineering practices and reliability standards across multiple teams and contribute to the organization's broader observability maturity.
    • Requirements:

      • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
      • Advanced hands-on experience with Datadog, including implementation, administration, configuration, and platform evolution.
      • Strong practical experience with Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Monitors, and Service Catalog.
      • Advanced understanding of the core observability pillars: logs, metrics, and distributed tracing.
      • Experience monitoring APIs, microservices, and distributed architectures.
      • Professional experience working with AWS and cloud environments.
      • Strong troubleshooting and complex incident investigation capabilities.
      • Experience with OpenTelemetry and application instrumentation.
      • Knowledge of automation using Python, Shell Script, or PowerShell.
      • Experience with Kubernetes, Docker, and cloud-native ecosystems.
      • Solid understanding of system availability, performance, scalability, and reliability.
      • Ability to connect technical metrics and operational events with business outcomes.
      • Excellent communication skills and the ability to collaborate with multiple technical and business stakeholders.
      • Investigative, analytical, and problem-solving mindset with a strong focus on continuous improvement.
      • Strong sense of ownership and operational responsibility.
      • Ability to influence technical teams and encourage the adoption of observability and reliability best practices.
      • Collaborative and consultative approach, with the ability to train and enable other teams.
      • Experience with SRE practices is desirable.
      • Datadog Certified Associate certification or higher is a plus.
      • Experience designing observability strategies for large-scale distributed environments is an advantage.
      • Knowledge of CI/CD and DevSecOps practices is desirable.
      • Experience with automated incident response and self-healing processes is a plus.
      • Familiarity with ITIL, incident management, Problem Management, and operational governance is desirable.
      • Experience with tools such as Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix is an advantage.
      • Benefits:

        • Senior-level opportunity with significant influence over enterprise observability and reliability strategy.
        • High-visibility work impacting multiple engineering and technology teams.
        • Opportunity to lead automation, intelligent observability, and incident-reduction initiatives.
        • Professional development through exposure to advanced cloud-native technologies and modern observability practices.
        • Opportunity to work with Datadog, OpenTelemetry, AWS, Kubernetes, Docker, and related technologies.
        • Collaborative environment involving Engineering, Architecture, Development, SRE, and Operations teams.
        • Inclusive workplace culture focused on diversity, equity, and professional growth.
        • Position specifically designed as an affirmative opportunity for women, supporting greater gender equity in technology leadership.
        • Access to initiatives and communities focused on women's development, leadership, networking, and career advancement.
        • Remote work arrangement in Brazil.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now