Principal Platform Engineer, AI Engineering

RxSense
$190,000 - $235,000

About The Position

RxSense is building a new cloud platform that it will own end to end. This platform will host the next generation of RxSense products, from established pharmacy benefit services to AI-native applications. The Principal Platform Engineer will lead the technical build of this platform, setting the direction for how services are built, deployed, secured, observed, and paid for. This is a hands-on role focused on building the infrastructure as code foundation, running Kubernetes at production quality, building and defending the deploy pipeline, and ensuring the platform is the fastest path to production. The role also involves hardening security and compliance, managing cloud spend, building observability, setting standards, and partnering across engineering teams. The engineer will be embedded with the AI Engineering team and partner closely with data engineering.

Requirements

  • 8 + years building and operating production platform infrastructure (strong candidates with less experience can still be considered).
  • Proven, hands-on experience operating production Kubernetes end to end, including cluster lifecycle, autoscaling, ingress, workload identity, secrets delivery, and hardening (EKS preferred).
  • Proven, hands-on experience owning infrastructure as code in Terraform at scale, including module design, state layout across multiple environments, and provider upgrades.
  • A track record of building or substantially rebuilding a CI/CD system yourself (e.g., GitHub Actions, GitLab CI, Argo, Jenkins), with clear positions on artifact immutability, build-once and promote-everywhere delivery, and keeping application pipelines thin.
  • Experience running self-hosted GitHub Actions runners at scale.
  • Hands-on depth in AWS: IAM, VPC networking and DNS, secrets management (e.g., Secrets Manager, External Secrets), container registries, and managed compute.
  • Hands-on experience with Helm at scale, including shared chart libraries, templating boundaries, and per environment configuration, alongside GitOps or push-based deployment workflows.
  • Proven experience building the developer-facing side of a platform: service templates, golden paths, self-service tooling, and documentation.
  • Hands-on experience implementing observability, including structured logging, metrics, and distributed tracing (e.g., OpenTelemetry, Prometheus and Grafana, Datadog), with correlation that holds across service boundaries.
  • Practical security experience in a regulated or security-sensitive environment: least privilege IAM, secrets hygiene, network isolation, image provenance and scanning, and rigorous PHI/PII handling, built so audit evidence comes out of the platform rather than getting assembled by hand.
  • Experience supporting data workloads on Kubernetes (e.g., Spark, Kafka, or orchestration tools such as Airflow or Dagster).
  • Demonstrated cloud cost ownership, including tagging and allocation, right sizing, and measurable spend reduction that did not degrade reliability.
  • A track record of writing and shipping production code yourself, not just producing diagrams and design documents.
  • Working fluency in at least one backend language (e.g., Python, Go, C# / .NET) and comfort in the shell.
  • Excellent communication and collaboration skills. You translate infrastructure and deployment decisions into terms engineers, architects, and non-technical leadership can act on, and you write things down so decisions outlive the conversation.
  • Experience mentoring engineers on infrastructure, deployment, and platform thinking, and setting standards a team can extend safely without you in the room.
  • Comfort working in a small, fast moving team where you will wear multiple hats, and a bias toward directness over ceremony: minimal dependency sprawl, skepticism of abstractions that do not earn their cost, and a preference for clear, traceable systems over fashionable patterns.

Nice To Haves

  • Experience in healthcare, pharmacy benefits, or another regulated data environment.
  • Direct experience preparing infrastructure evidence for HIPAA, SOC 2, or comparable audits.
  • Experience standing up platforms from scratch (greenfield), not just extending or migrating existing systems.
  • Experience supporting latency-sensitive or high-throughput services, including workloads that call large language models, with attention to cost and latency.

Responsibilities

  • Build the infrastructure as code foundation, designing and maintaining a Terraform monorepo across dev, QA, staging, and production, covering Kubernetes clusters, networking, IAM, and per-application platform stacks.
  • Keep state layout, module boundaries, and provider baselines clean and current.
  • Run Kubernetes at production quality, operating EKS clusters end to end, including node lifecycle, autoscaling, ingress, workload identity, secrets delivery, and cluster security.
  • Keep clusters hardened and appropriately isolated.
  • Build and defend the deploy pipeline, creating push-based CI/CD on self-hosted GitHub Actions runners with build-once, promote-everywhere artifact immutability across environments.
  • Enforce a promotion flow so no environment is ever skipped and production always mirrors a released artifact.
  • Make the platform the fastest path to production by maintaining a shared Helm chart library and per-service charts, building golden paths for new services, and pushing per-application behavior into configuration.
  • Harden the security and compliance posture by setting least-privilege IAM, secrets management, network boundaries, image provenance, and production guardrails.
  • Make controls automatic where possible and auditable where not, so evidence for security reviews falls out of the platform.
  • Keep cloud spend predictable by establishing tagging and allocation, right-sizing compute, and keeping spend predictable as traffic, data, and model inference grow.
  • Build observability in, not on, by establishing structured logging, metrics, tracing, and correlation across service hops as a default property of the platform.
  • Treat telemetry contracts as published, versioned schemas rather than debug output.
  • Set platform conventions (tagging, naming, DNS, versioning, security posture) and document the reasoning behind them.
  • Review infrastructure and deploy changes, mentor engineers, and make the platform something the team can extend.
  • Partner with application, data, and AI teams so the platform fits how services actually run, including the contracts they deploy against and the environments they promote through.

Benefits

  • Equal Opportunity and Affirmative Action employer
  • Recruitment process free from discriminatory hiring practices
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service