Remote | Kubernetes Engineer — $60–$80/hour

24-MagNew York, NY
$60 - $80Remote

About The Position

We are sharing a specialised part-time consulting opportunity for experienced Kubernetes Engineers with strong hands-on expertise in production cluster operations, infrastructure troubleshooting, manifests, Helm, networking, storage, and platform reliability. This role focuses on evaluating Kubernetes engineering tasks for technical correctness, production readiness, and operational realism. Selected engineers will review cluster scenarios, manifests, troubleshooting workflows, and failure modes while providing clear, rubric-based technical feedback.

Requirements

  • 3+ years of hands-on production Kubernetes experience
  • Experience operating EKS, GKE, AKS, or self-managed Kubernetes clusters
  • Deep understanding of CNI, DNS, ingress, PV/PVC storage, RBAC, scheduling, and cluster failure modes
  • Strong experience authoring and reviewing Kubernetes manifests and Helm charts
  • Demonstrated experience debugging live production cluster incidents
  • Proficiency in Go, Python, or TypeScript
  • Strong understanding of workload resource management and production reliability
  • Strong written communication and ability to provide precise technical feedback

Nice To Haves

  • CKA or CKAD certification is preferred
  • Experience with Istio, HPA, Prometheus, Grafana , or comparable technologies is advantageous
  • Previous SRE, platform engineering, code-review, or technical task-grading experience is advantageous

Responsibilities

  • Review production-oriented Kubernetes scenarios across managed or self-hosted environments
  • Assess cluster configuration, workload behaviour, and operational readiness
  • Evaluate whether proposed approaches reflect sound Kubernetes practices
  • Identify configuration errors, unsafe assumptions, or production risks
  • Apply practical judgement grounded in real-world cluster operations
  • Evaluate Kubernetes networking and service-discovery configurations
  • Review CNI, DNS, ingress, and service connectivity
  • Diagnose connectivity and routing issues across workloads
  • Identify incorrect networking assumptions or configuration problems
  • Assess whether proposed fixes address the underlying cluster issue
  • Review PersistentVolume and PersistentVolumeClaim (PV/PVC) configurations
  • Assess storage classes, volume lifecycle behaviour, and workload dependencies
  • Identify provisioning, mounting, or persistence issues
  • Evaluate storage-related failure scenarios
  • Review proposed remediation for technical correctness
  • Evaluate Kubernetes RBAC configurations
  • Review roles, cluster roles, bindings, and service-account permissions
  • Identify missing or excessive access
  • Assess whether permissions appropriately support workload requirements
  • Apply least-privilege principles when evaluating access-control decisions
  • Diagnose common Kubernetes failures such as CrashLoopBackOff, OOMKilled, scheduling failures, and pod eviction
  • Review logs, events, workload configuration, and resource behaviour
  • Evaluate troubleshooting sequences for efficiency and technical accuracy
  • Identify root causes rather than superficial symptoms
  • Assess whether proposed corrective actions are production appropriate
  • Author and review Kubernetes manifests
  • Evaluate Helm charts for correctness, maintainability, and deployment safety
  • Review resource specifications, probes, requests, limits, selectors, and dependencies
  • Identify templating or configuration problems
  • Assess whether manifests accurately represent the intended workload
  • Evaluate live-cluster troubleshooting scenarios
  • Review diagnostic reasoning and incident-response decisions
  • Assess prioritisation, escalation, and remediation approaches
  • Identify missed evidence or ineffective debugging paths
  • Determine whether proposed solutions reduce recurrence risk
  • Evaluate scaling, availability, resilience, and operational-readiness considerations
  • Review workloads involving autoscaling or service-mesh technologies where relevant
  • Assess observability and operational visibility
  • Identify reliability weaknesses in workload or cluster design
  • Evaluate recommendations for production readiness
  • Review monitoring and troubleshooting workflows using tools such as Prometheus and Grafana
  • Evaluate alerting, metrics, logs, and operational signals
  • Assess HPA configurations and scaling behaviour where relevant
  • Review service-mesh scenarios involving technologies such as Istio
  • Identify gaps affecting diagnosis or performance management
  • Assess Kubernetes tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific configuration, runtime behaviour, or diagnostic evidence
  • Apply grading standards consistently across assignments
  • Distinguish valid alternative approaches from genuinely incorrect solutions

Benefits

  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service