Data Platform Infrastructure Manager

SailPoint
$124,700 - $210,158

About The Position

As the Data Infrastructure Platform Manager, you will lead the team responsible for the runtime and platform layer beneath SailPoint’s batch, streaming, analytics, and machine learning workloads. Your team owns the availability, scalability, security, performance, lifecycle, and cost efficiency of shared services such as Airflow, Flink, Spark on AWS EMR, Kafka, Snowflake, and Iceberg. Data and product engineering teams own the workload-specific pipelines and processing logic that run on those services. Your day will span people leadership, production operations, technical strategy, and cross-functional execution. You will review service health and incidents, set priorities across operational and roadmap work, coach and unblock engineers, make architectural and investment tradeoffs, and partner with data engineering, Developer Platform, SRE, Observability, Security, and Infrastructure teams. You will make the platform easier to consume through paved roads, CI/CD, configuration as code, observability, automation, and self-service—all while protecting reliability, quality, and cost efficiency in a fast-moving environment. The Data Infrastructure Platform team designs, builds, and operates the production-grade data processing infrastructure that powers SailPoint Identity Security. We provide reliable, scalable, and secure data platforms as services so data and product engineers can focus on DAGs, streaming jobs, pipelines, models, and business logic rather than provisioning and operating the underlying infrastructure. The team values engineering and operations excellence, service ownership, practical automation, constructive debate, continuous learning, and a “strong opinions, loosely held” mindset.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • Proven experience managing and leading a cloud infrastructure, platform engineering, SRE, or data infrastructure team responsible for business-critical production services.
  • Deep understanding of cloud infrastructure, distributed systems, and production operations.
  • 5+ years of production experience with AWS.
  • 3+ years of experience with Kubernetes and containerized workloads.
  • Ability to read, review, and troubleshoot software written in Python, Go, Java, or a comparable language.
  • Experience with infrastructure as code, preferably Terraform, and modern CI/CD or GitOps practices.
  • Strong systems, networking, security, and distributed-systems troubleshooting fundamentals.
  • Substantial hands-on experience operating production data platforms at scale, with depth in several of the following: Apache Airflow, Apache Kafka, Apache Flink, Apache Spark, AWS EMR, Snowflake, and Apache Iceberg. Experience operating these technologies as shared services is strongly preferred.
  • Experience with machine learning systems and their operational lifecycle, including practical experience training, evaluating, or deploying ML models and familiarity with AWS SageMaker or an equivalent ML platform.
  • Working knowledge of modern LLM capabilities and AI-assisted engineering practices, with sound judgment about responsible use, validation, and production guardrails.
  • Demonstrated success building internal platforms as products, including self-service developer experiences, standardized delivery workflows, CI/CD, observability, and clear service ownership.
  • Strong production-operations discipline, including metrics and observability, SLOs, incident response, capacity planning, disaster recovery, and continuous reliability improvement.
  • Experience managing cloud infrastructure cost, capacity, and performance, with a track record of making measurable efficiency improvements.
  • Excellent leadership, communication, negotiation, and cross-functional collaboration skills.
  • Ability to thrive in a fast-paced environment with shifting priorities and incomplete information while preserving engineering rigor, production uptime, and quality.
  • Experience with Agile or similar iterative planning and delivery practices, applied pragmatically to a team that balances roadmap work with operational demand.
  • Passion for engineering excellence, reliability, continuous learning, and developing people.

Nice To Haves

  • Databricks experience is a plus.
  • Familiarity with AWS SageMaker or an equivalent ML platform.

Responsibilities

  • Lead the team responsible for the runtime and platform layer beneath SailPoint’s batch, streaming, analytics, and machine learning workloads.
  • Own the availability, scalability, security, performance, lifecycle, and cost efficiency of shared services such as Airflow, Flink, Spark on AWS EMR, Kafka, Snowflake, and Iceberg.
  • Manage people leadership, production operations, technical strategy, and cross-functional execution.
  • Review service health and incidents, set priorities across operational and roadmap work, coach and unblock engineers, make architectural and investment tradeoffs, and partner with various teams.
  • Make the platform easier to consume through paved roads, CI/CD, configuration as code, observability, automation, and self-service.
  • Design, build, and operate production-grade data processing infrastructure.
  • Provide reliable, scalable, and secure data platforms as services.
  • Establish working relationships with team members and key partners.
  • Document and align stakeholders on the team charter, service catalog, ownership boundaries, escalation paths, and the distinction between platform ownership and workload-specific pipeline ownership.
  • Complete an initial assessment of the team’s people, platforms, roadmap, on-call load, incidents, operational risks, capacity, and cloud costs.
  • Establish a regular operating cadence for team priorities, service health, incidents, roadmap delivery, and cross-team dependencies.
  • Publish an outcome-oriented 12-month platform roadmap.
  • Define the operating model for the team’s highest-criticality services.
  • Select and begin delivery of the first high-value paved-road or self-service improvement.
  • Set clear performance expectations and development goals for each team member.
  • Baseline the platform’s key reliability, delivery, toil, utilization, and cost measures.
  • Ensure service-level objectives, actionable dashboards, alerts, and recurring service reviews are in place for the highest-criticality data platform services.
  • Deliver at least one production self-service or standardized delivery capability.
  • Implement a cost and capacity management program.
  • Strengthen incident response, change management, disaster recovery, vulnerability remediation, and operational runbooks.
  • Establish a clear platform approach for supporting machine learning workloads.
  • Operate the data infrastructure platform as a mature internal product.
  • Deliver the highest-priority roadmap outcomes and demonstrate measurable year-over-year improvement.
  • Build a healthy, high-performing team.
  • Establish a durable multi-year strategy for key data infrastructure components.
  • Be recognized by partner teams as a responsive, reliable platform organization.

Benefits

  • Health and wellness coverage: Medical, dental, and vision insurance
  • Disability coverage: Short-term and long-term disability
  • Life protection: Life insurance and Accidental Death & Dismemberment (AD&D)
  • Additional life coverage options: Supplemental life insurance for employees, spouses, and children
  • Flexible spending accounts for health care, and dependent care; limited purpose flexible spending account
  • Financial security: 401(k) Savings and Investment Plan with company matching
  • Time off benefits: Flexible vacation policy
  • Holidays: 8 paid holidays annually
  • Sick leave
  • Parental support: Paid parental leave
  • Employee Assistance Program (EAP) and Care Counselors
  • Voluntary benefits: Legal Assistance, Critical Illness, Accident, Hospital Indemnity and Pet Insurance options
  • Health Savings Account (HSA) with employer contribution
  • SailPoint Corporate Bonus Plan or a role-specific commission
  • Potential eligibility for equity participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service