Datavant2
Senior Site Reliability Engineer
engineeringfull-timeRemote - United States
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
healthcare
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
What We’re Looking For
We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission-critical data and ML workloads across our organization.
This role is ideal for someone who combines a strong SRE mindset with deep cloud infrastructure and data platform experience. You're comfortable operating at scale in a complex, hybrid cloud environment and can architect systems that balance velocity, safety, and cost. You’ll work closely with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to build a modern data platform that is secure, self-service, and production-grade.
What You Will Do
- Operate and Improve Databricks and Snowflake: Own Databricks & Snowflake platforms lifecycle—including automation, workspace governance, job orchestration, and cost optimization.
- Design for Reliability: Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.
- Advance Observability: Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.
- Drive CI/CD for Data & ML: Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
- Enable Data Flow Across Platforms: Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
- Champion Event-Driven Architectures: Leverage cloud-native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
- Collaborate Across Teams: Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.
- Contribute to Strategy: Influence engineering-wide decisions on data platform architecture, ML enablement, and data product strategy.
What You Need to Succeed
- 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
- AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
- Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
- Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven p
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply