Jobgether
Technical Lead, Platform Engineering (Observability)
engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Observability Platform Engineering: Design, build, operate, and continuously improve an observability platform covering metrics, logs, distributed traces, dashboards, and alerting at scale.
- Technical Leadership: Drive the technical direction of observability capabilities, make architectural decisions, establish engineering standards, and lead initiatives from design through implementation and operation.
- Reliability Improvement: Deliver measurable improvements in Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM), while improving service reliability and operational resilience.
- AI-Powered Observability: Design and implement intelligent solutions for automated anomaly detection, alert correlation, root cause analysis, and AI-assisted incident response.
- Self-Service Tooling: Build developer-focused tooling that enables product engineering teams to independently instrument services, create dashboards, configure alerts, and adopt observability best practices.
- Standards & SLOs: Define and champion observability standards, instrumentation practices, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error-budget frameworks across the engineering organization.
- Platform Optimization: Improve data collection, processing pipelines, alert quality, and operational workflows while reducing noise, unnecessary toil, and observability-related costs.
- Cross-Functional Collaboration: Partner with SRE, Security, platform teams, and product engineering teams to ensure comprehensive system visibility and consistent reliability practices.
- Automation: Automate recurring operational processes and workflows to increase engineering efficiency and improve the team's ability to respond to incidents.
- Mentorship & Culture: Lead by example, mentor engineers, contribute to technical discussions, and help foster a strong engineering culture centered on collaboration, ownership, and continuous improvement.
- Documentation: Produce high-quality technical and architectural documentation and facilitate design discussions with relevant engineering stakeholders.
- Professional Experience: 9+ years of experience building, operating, and maintaining scalable production systems, with significant experience in platform engineering, infrastructure, SRE, or related disciplines.
- Observability Expertise: Strong hands-on experience with production observability and monitoring platforms such as Datadog, Prometheus, Grafana, or equivalent technologies.
- Distributed Systems: Deep understanding of metrics, logging, distributed tracing, instrumentation patterns, and observability data pipeline architecture.
- Kubernetes: Proven experience deploying, operating, and troubleshooting Kubernetes-based production environments and containerized workloads.
- Programming: Strong proficiency in Go or Python for developing infrastructure tooling, platform services, and automation.
- Cloud & Infrastructure as Code: Experience with GCP and/or AWS and Infrastructure as Code technologies such as Terraform.
- Alerting & Incident Detection: Demonstrated ability to design and tune alerting systems to improve signal quality, reduce alert fatigue, and accelerate incident detection.
- Reliability Engineering: Strong understanding of SLIs, SLOs, error budgets, and modern reliability engineering practices.
- Developer Platforms: Experience building internal platforms, tools, or self-service capabilities that improve developer productivity and engineering workflows.
- Communication: Strong written and verbal communication skills, with the ability to produce design documentation, explain complex technical concepts, and lead technical discussions.
- Technical Leadership: Proven ability to influence technical direction, mentor engineers, make sound architectural decisions, and drive initiatives across teams.
- AI & Advanced Observability: Experience applying AI or machine learning to observability use cases such as anomaly detection, alert correlation, or root cause analysis is preferred.
- Large-Scale Systems: Experience supporting observability across large distributed environments, ideally involving hundreds of microservices.
- OpenTelemetry: Hands-on experience with OpenTelemetry for instrumentation and telemetry collection is an advantage.
- Observability Cost Optimization: Experience with sampling strategies, data tiering, pipeline optimization, or other approaches to controlling observability costs at scale is preferred.
- Developer Experience: Passion for improving engineering productivity through better platform tooling, automation, and self-service capabilities.
- Open Source: Contributions to or active participation in open-source observability communities are a plus.
- Incident Management: Experience with incident management processes, tooling, and operational response practices is desirable.
- Full-time employment with the opportunity to work on large-scale platform and observability challenges.
- Hybrid work model with 2 days per week working from the office and 3 days remotely.
- Bengaluru-based office environment with opportunities for collaboration and team engagement.
- Opportunity to shape observability practices across a complex, distributed engineering environment.
- Exposure to advanced technologies spanning cloud infrastructure, Kubernetes, AI-powered operations, distributed systems, and developer platforms.
- Strong opportunities for technical leadership, mentoring, and career development.
- Collaborative environment involving platform engineering, SRE, security, and product engineering teams.
- Opportunity to directly improve engineering productivity, system reliability, and incident response at scale.
- Professional environment focused on innovation, continuous improvement, collaboration, and high engineering standards.
Requirements
Benefits
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply