Full Stack Cloud Engineer

Agama SolutionsChicago, IL
Onsite

About The Position

This role involves supporting the day-to-day operations of an enterprise Kubernetes platform, which includes over 100 clusters, with approximately 50% in production. The engineer will perform routine operational tasks such as cluster maintenance, upgrades, patching, health checks, and capacity management. A key part of the role is troubleshooting and resolving Kubernetes platform issues that affect cluster or application availability, participating in incident response, root-cause analysis, and post-incident reviews. The position also requires acting as a backup platform engineer to support an on-call rotation and reduce key-person dependency, including providing after-hours support. The engineer will serve as a secondary escalation point for critical Production issues and assist internal application teams with Kubernetes-related questions and issues. This includes supporting common Kubernetes constructs and helping teams troubleshoot networking, DNS, ingress, certificate, and resource-related issues. The role also involves reviewing application configurations for Kubernetes best practices and platform alignment, and working with integrated enterprise tools like ingress controllers, logging platforms, monitoring/observability tools, and container registries. Additionally, the engineer will help document operational procedures, runbooks, and troubleshooting guides, share Kubernetes knowledge, and assist in improving platform resiliency, operational maturity, and supportability.

Requirements

  • Experience with Kubernetes platform operations
  • Experience with cluster maintenance, upgrades, patching, health checks, and capacity management
  • Experience with incident response, root-cause analysis, and post-incident reviews
  • Experience providing after-hours support and participating in on-call rotations
  • Experience supporting common Kubernetes constructs (Pods, Deployments, Services, Ingress, ConfigMaps, Secrets)
  • Experience troubleshooting networking, DNS, ingress, certificate, and resource-related issues
  • Experience with integrated enterprise tools such as Ingress controllers (e.g., Contour / Envoy)
  • Experience with Logging platforms (e.g., Fluent Bit, centralized log aggregation)
  • Experience with Monitoring/observability tools (e.g., Dynatrace or similar)
  • Experience with Container registries (e.g., Harbor, JFrog, etc)

Nice To Haves

  • Serve as a backup platform engineer to enable on-call rotation and reduce key-person dependency
  • Serve as a secondary escalation point for critical Production issues
  • Review application configurations for Kubernetes best practices and platform alignment
  • Help document operational procedures, runbooks, and troubleshooting guides
  • Share Kubernetes knowledge and best practices with internal teams
  • Assist in improving platform resiliency, operational maturity, and supportability

Responsibilities

  • Support day-to-day operations of an enterprise Kubernetes platform (100+ clusters, ~50% production)
  • Perform routine operational tasks including cluster maintenance, upgrades, patching, health checks, and capacity management
  • Troubleshoot and resolve Kubernetes platform issues impacting cluster or application availability
  • Participate in incident response, root-cause analysis, and post-incident reviews
  • Provide after-hours support as part of a shared on-call rotation
  • Serve as a secondary escalation point for critical Production issues
  • Assist internal application teams with Kubernetes-related questions and issues
  • Support common Kubernetes constructs such as Pods, Deployments, Services, Ingress, ConfigMaps, and Secrets
  • Help teams troubleshoot networking, DNS, ingress, certificate, and resource-related issues
  • Review application configurations for Kubernetes best practices and platform alignment
  • Work with integrated enterprise tools such as Ingress controllers (e.g., Contour / Envoy), Logging platforms (e.g., Fluent Bit, centralized log aggregation), Monitoring/observability tools (e.g., Dynatrace or similar), and Container registries (e.g., Harbor, JFrog, etc)
  • Help document operational procedures, runbooks, and troubleshooting guides
  • Share Kubernetes knowledge and best practices with internal teams
  • Assist in improving platform resiliency, operational maturity, and supportability
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service