Senior AI-Enabled Platform / SRE Engineer

ZipStaffDallas, TX
Hybrid

About The Position

ZipStaff is seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills to support a major healthcare organization. This hybrid role (Dallas, TX or Scottsdale, AZ) focuses on platform reliability at scale and using AI/LLMs to automate operations. The initial assignment is approximately one year, with potential extension.

Requirements

  • 5+ years of strong hands-on Kubernetes platform experience with GKE and Rancher RKE2, including multi-cluster management, troubleshooting, and performance optimization.
  • 5+ years of advanced programming in Python and Java (Node.js preferred for integrations and automation).
  • Strong SRE background: reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
  • Hands-on observability experience with Splunk, Grafana, Datadog, and/or AppDynamics.
  • Experience with API and microservices reliability (Apigee / Apigee X, REST, GraphQL, traffic routing, canary, failover).
  • Experience applying LLMs (Gemini, Llama, Mistral, Qwen, or similar) to alert analysis, incident triage, automation, or operational workflows.
  • Ability to work hybrid in Dallas, TX or Scottsdale, AZ.
  • Must be legally authorized to work in the United States without sponsorship now or in the future.

Nice To Haves

  • Experience supporting highly available, multi-datacenter production platforms.
  • Prior healthcare or regulated-enterprise SRE experience.
  • Demonstrated AIOps / GenAI-for-operations implementations in production.

Responsibilities

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage Generative AI (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee / Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Support active-active deployments, disaster-recovery readiness, and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring using Splunk, Grafana, Datadog, and AppDynamics.
  • Drive SRE best practices by partnering with cross-functional teams on reliability, security, incident management, and continuous improvement.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service