Lead Kubernetes Platform Engineer

EverOps•San Francisco, CA
•Remote

About The Position

EverOps is seeking a Lead Kubernetes Platform Engineer with deep Amazon EKS experience at very large scale to lead a modernization discovery and the subsequent engineering program for a high-scale consumer mobile platform. The role involves stepping into a complex, multi-cluster production EKS environment with significant scale and performance demands. The engineer will be responsible for assessing the current state, identifying areas for improvement, and leading the execution of a modernization program focused on reducing blast radius, automating upgrades, and optimizing compute costs, including a migration to Graviton. Success will be measured by improved cost efficiency and reduced platform engineering maintenance time. This is a player-coach role, leading an embedded TechPod and serving as the primary technical contact for customer infrastructure leadership.

Requirements

  • 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering.
  • 4+ years operating production Kubernetes.
  • Prior experience in a technical lead, staff, or principal-level role.
  • Deep production experience with Amazon EKS at large scale (multiple production clusters, thousands of nodes, or tens of thousands of pods).
  • Working understanding of control plane, scheduling, networking, and API limits at scale.
  • Hands-on ownership of EKS version upgrades across multiple production clusters, including API deprecations, add-on compatibility, node rotation, and rollback planning.
  • Working knowledge of the EKS standard and extended support lifecycle.
  • Advanced production experience with Karpenter (v1+), including NodePools, disruption and consolidation, weighting, and instance-type flexibility.
  • Strong understanding of EC2 instance families and generations, CPU architecture differences, network performance limits.
  • Experience migrating production workloads to Graviton or other ARM64 platforms, including multi-arch container builds and performance validation.
  • Experience modeling compute costs and commitments (Savings Plans, Reserved Instances, Spot) and building savings cases.
  • Solid knowledge of the AWS VPC CNI, ingress controllers, Gateway API, service mesh, and load balancing on EKS.
  • Advanced proficiency with Terraform.
  • Experience with Helm and GitOps tooling such as Argo CD or Flux.
  • Comfortable using Datadog, Prometheus, Grafana, or comparable tooling.
  • Strong scripting ability using Python, Go, or Bash.
  • Demonstrated ability to enter an unfamiliar environment, surface undocumented information, and produce a defensible current-state picture and roadmap.
  • Ability to translate technical findings into cost, risk, and capacity terms for engineering leadership and executives.

Nice To Haves

  • Experience with high-traffic B2C platforms with sharp daily peaks, or with real-time, safety-critical services.
  • Experience designing cell-based or multi-cluster architectures.
  • Experience running performance-per-dollar comparisons across instance generations and architectures.
  • Experience operating Spot capacity for large production fleets, including interruption handling for stateful or stream-processing workloads.
  • Experience tuning JVM heap, garbage collection, and Kubernetes requests/limits.
  • Experience running Bottlerocket or similar, including node boot-time optimization.
  • Experience building AI agents or LLM-driven tooling for infrastructure operations.
  • Experience leading assessment or discovery engagements that end in a committed, funded program of work.
  • CKA, CKS, AWS Certified Solutions Architect – Professional, AWS Certified DevOps Engineer – Professional, or similar advanced certifications.

Responsibilities

  • Lead a two-month EKS modernization discovery, including baselining the production estate, mapping Kubernetes versions, analyzing failure domains, and delivering an upgrade automation design, compute economics model, and prioritized roadmap.
  • Lead the execution of the modernization program across three phases: reducing blast radius with a multi-cluster architecture, automating the upgrade cycle, and right-sizing compute with a move to Graviton (ARM64).
  • Build a complete baseline of a multi-cluster production EKS estate, covering cluster inventory, topology, workload placement, ownership, cost, and Kubernetes version/support status.
  • Analyze failure domains and design a multi-cluster target architecture with clear workload placement and tenancy models.
  • Design and implement an upgrade approach that automates the EKS upgrade cycle.
  • Lead instance sizing and workload-fit analysis across EC2 instance families, generations, and node sizes.
  • Plan and drive a phased ARM64 migration, including multi-architecture builds and performance validation.
  • Configure Karpenter for optimal price-performance and design safe Spot instance patterns.
  • Model compute, support, and commitment economics and quantify savings.
  • Assess and improve ingress, service mesh, CNI, GitOps, and infrastructure-as-code posture.
  • Quantify maintenance and toil burden and build automation to reduce it.
  • Define scope and success criteria for AI-assisted DevOps agent workstreams.
  • Lead the TechPod as a player-coach, manage the working cadence with customer infrastructure owner, and partner with AWS specialists.
  • Produce documentation including estate baselines, architecture designs, upgrade plans, roadmaps, and executive readouts.
  • Present findings and recommendations to engineering leadership.

Benefits

  • 100% Remote Workplace
  • Unlimited Paid Time Off
  • Equity
  • 401K with company contribution
  • Sponsored healthcare
  • Professional Growth: Access to training and certification programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service