Senior Site Reliability Engineer

Datavant
$168,000 - $200,000

About The Position

Datavant is seeking a Senior Site Reliability Engineer to join their Data & ML Platform team. The role involves building and operating a resilient, observable, and scalable platform for mission-critical data and ML workloads. This position is suited for individuals with a strong SRE mindset, deep cloud infrastructure experience, and expertise in data platforms. The ideal candidate will be comfortable operating at scale in a complex, hybrid cloud environment and can design systems that balance velocity, safety, and cost. The Senior SRE will collaborate with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to develop a modern, secure, self-service, and production-grade data platform.

Requirements

  • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
  • AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
  • Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.
  • Proven expertise with observability tools (especially Datadog) and architecting platform-wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, infrastructure-as-code (Terraform), and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault-tolerant systems.
  • Excellent communication and collaboration skills; able to work effectively across teams.

Nice To Haves

  • DevSecOps mindset: Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.
  • Experience with ML infrastructure tooling such as MLflow, Feature Stores, and GPU workload orchestration.
  • Strong experience in both Databricks and Snowflake in a large scale production lakehouse with cross-warehouse interoperability, e.g. Iceberg v3, Glue, etc.
  • Background in compliance-aware architecture (e.g., HIPAA, SOC 2) or regulated industries.
  • Familiarity with multi-cloud or hybrid cloud data environments; experience with Azure.
  • Contributions to open-source infrastructure, SRE, or observability tools.

Responsibilities

  • Operate and improve Databricks and Snowflake platforms, including lifecycle management, automation, workspace governance, job orchestration, and cost optimization.
  • Design resilient, scalable, and secure infrastructure across cloud environments, driving initiatives in failover, autoscaling, chaos testing, and capacity planning.
  • Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling, defining and enforcing SLOs/SLAs.
  • Automate deployments of data pipelines, ML workflows, and infrastructure components using GitHub Actions, Terraform, and related IaC tooling.
  • Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
  • Leverage cloud-native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
  • Act as the SRE and platform partner for various teams, ensuring the platform meets the needs of analytics, data science, and product use cases.
  • Influence engineering-wide decisions on data platform architecture, ML enablement, and data product strategy.

Benefits

  • Total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service