This role involves supporting the day-to-day operations of an enterprise Kubernetes platform, which includes over 100 clusters, with approximately 50% in production. The engineer will perform routine operational tasks such as cluster maintenance, upgrades, patching, health checks, and capacity management. A key part of the role is troubleshooting and resolving Kubernetes platform issues that affect cluster or application availability, participating in incident response, root-cause analysis, and post-incident reviews. The position also requires acting as a backup platform engineer to support an on-call rotation and reduce key-person dependency, including providing after-hours support. The engineer will serve as a secondary escalation point for critical Production issues and assist internal application teams with Kubernetes-related questions and issues. This includes supporting common Kubernetes constructs and helping teams troubleshoot networking, DNS, ingress, certificate, and resource-related issues. The role also involves reviewing application configurations for Kubernetes best practices and platform alignment, and working with integrated enterprise tools like ingress controllers, logging platforms, monitoring/observability tools, and container registries. Additionally, the engineer will help document operational procedures, runbooks, and troubleshooting guides, share Kubernetes knowledge, and assist in improving platform resiliency, operational maturity, and supportability.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed