Principal Platform Engineer

ELLKAY, LLC US,
$160,000 - $180,000Hybrid

About The Position

ELLKAY started out providing connectivity solutions to laboratories and within a few years, grew to also provide data management solutions to ambulatory organizations. ELLKAY is now a trusted data management partner in five healthcare segments. ELLKAY’s solutions continue to serve laboratories and ambulatory practices and have expanded to empower hospitals and health systems, healthcare IT vendors, ambulatory practices, health plans, and other healthcare organizations with cutting-edge technologies and solutions that drive their growth and interoperability strategies. Today, ELLKAY remains true to our core values, building strong partner relationships and offering unparalleled service and support while providing innovative, scalable solutions to the challenges our customers face in today’s data-rich world. ELLKAY's experience, customer-focused approach, and reputation for innovation, speed, and accuracy differentiate ELLKAY as a premier partner for your interoperability needs and data management strategy. We're looking for a Principal Platform Engineer to own the design, build, and operational health of our infrastructure across AWS, Azure, and on-premises environments. This is a hands-on, high-ownership role for someone who thinks in systems, codifies everything, and is equally comfortable writing Terraform modules, debugging a production incident at 2 AM, and advising leadership on cost and security posture. You will be the primary technical bridge between Platform Engineering and our Site Reliability Engineering (SRE) teams, setting the standards for how services are deployed, observed, secured, and operated at scale across hybrid cloud and on-prem infrastructure.

Requirements

  • 15+ years in platform engineering, infrastructure engineering, DevOps, or SRE roles, with demonstrated staff-level scope and impact
  • Deep, production-grade experience with both AWS and Azure, plus experience managing on-premises infrastructure in a hybrid model
  • Expert-level Terraform experience — module design, state management, workspace/environment strategy at scale
  • Strong experience building CI/CD pipelines for Kubernetes and Docker-based workloads
  • Hands-on experience with configuration management tooling (Ansible, Chef, Puppet, or similar)
  • Proven track record implementing observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry, ELK/Splunk) and defining meaningful SLIs/SLOs
  • Demonstrated experience leading production incident response and driving reliability improvements
  • Experience designing cloud cost governance and security/compliance frameworks in a multi-cloud or hybrid environment
  • Strong scripting/programming ability (Python, Go, or Bash) for automation and tooling
  • Excellent cross-functional communication — able to work directly with SRE, security, and engineering leadership

Nice To Haves

  • Experience with policy-as-code frameworks (OPA/Gatekeeper, Sentinel)
  • Relevant certifications (AWS, Azure, CKA/CKAD)
  • Experience operating infrastructure under regulatory or compliance requirements (SOC 2, HIPAA, PCI-DSS, ISO 27001)
  • Prior experience in an on-call rotation for critical production systems
  • History of mentoring engineers or leading infrastructure initiatives across multiple teams

Responsibilities

  • Design, build, and maintain reusable Terraform modules to provision and manage infrastructure across AWS, Azure, and on-prem environments
  • Establish IaC standards, module versioning strategy, state management practices, and review processes across engineering teams
  • Drive migration of manually managed infrastructure to fully codified, version-controlled definitions
  • Architect and implement CI/CD pipelines for containerized and Kubernetes-based services, from build through progressive production rollout
  • Define deployment strategies (blue/green, canary, rolling) and the automation that supports them
  • Partner with development teams to streamline the path from commit to production while maintaining safety and auditability
  • Own configuration management tooling and practices across hybrid environments (e.g., Ansible, Chef, Puppet, or equivalent) to ensure consistency, repeatability, and drift detection
  • Standardize secrets management, environment configuration, and golden image/baseline practices across cloud and on-prem fleets
  • Design and implement observability patterns (metrics, logging, tracing) that provide actionable signal across distributed, hybrid-cloud services
  • Define SLIs/SLOs in partnership with SRE and product teams, and build the dashboards and alerting that make them actionable
  • Reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR) through better instrumentation, not just more of it
  • Act as a senior escalation point for complex production incidents, driving root cause analysis and durable remediation
  • Lead or contribute to postmortems and translate findings into infrastructure, process, or tooling improvements
  • Proactively identify and remediate reliability risk before it becomes an incident
  • Design and implement cost governance practices — tagging standards, budget alerting, rightsizing, and reserved capacity strategy across AWS, Azure, and on-prem infrastructure
  • Partner with Security to define and enforce infrastructure security guardrails (IAM least-privilege, network segmentation, secrets handling, compliance controls)
  • Build automated policy enforcement (e.g., policy-as-code) so governance scales with infrastructure rather than depending on manual review
  • Serve as the primary point of contact between Platform Engineering and SRE teams, aligning on standards, priorities, and shared tooling
  • Mentor senior and mid-level engineers on infrastructure design, operational excellence, and IaC best practices
  • Influence infrastructure architecture and technical roadmap at the organizational level

Benefits

  • Medical, Dental, and Vision benefits
  • Employer-paid Life and LTD
  • 401k w/ matching – once eligibility is met
  • Work/life balance
  • Paid Volunteer Program
  • Flexible working hours
  • Generous FTO
  • Remote work options
  • Employee Discounts
  • Parental Leave
  • Gym membership / Exercise class stipends
  • On site in HQ Free daily lunches
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service