Site Reliability Engineer

MSPPlano, TX
Hybrid

About The Position

Genesis10 is seeking a Site Reliability Engineer for a hybrid contract position with a Global Financial Institution. This role involves a 12+ month contract and requires 3 days onsite per week. The position is responsible for the reliability and support of the enterprise Container Platform, which spans on-premise and external clouds like Azure, AWS, and Google. The engineer will monitor and troubleshoot the platform, conduct in-depth analysis of systemic issues, and identify opportunities for automation and operational excellence.

Requirements

  • BS/MS degree in Computer Science or a related technical field
  • Minimum 5+ years of hands-on experience supporting Kubernetes, OpenShift, RKE, or EKS Container platforms
  • Experience with Python, Ansible, Golang, and shell scripting
  • Strong experience in major services related to Compute, Storage, Network, and Security
  • Experience with monitoring tools like Prometheus and Dynatrace, as well as cloud-native tools like Azure Monitor and Log Analytics
  • Strong understanding of complex IAM infrastructure, including Active Directory, Azure AD, and SSO solutions
  • Advanced knowledge of Linux OS, DNS, DHCP, Kerberos, and Windows Authentication
  • Experience with CI/CD tools such as Git/Jenkins and GitOps models
  • Excellent understanding of Linux/Windows operating systems administration
  • Experience in Container security and vulnerability remediation
  • Systematic problem-solving approach, sense of ownership, and drive
  • Excellent interpersonal, organizational, and communication skills

Nice To Haves

  • Kubernetes, OpenShift, or Terraform certifications are a plus
  • Experience in OpenShift, RKE, and CSP Kubernetes services such as AKS and EKS
  • Experience in Terraform, ArgoCD, Tekton, and K-native technologies
  • Experience in agile deployment methodologies (GitOps)
  • Knowledge of various container runtimes
  • Familiarity with the operator deployment pattern
  • Experience working in a highly available multi-datacenter environment
  • Experience working with monitoring tools such as Prometheus, Splunk, Dynatrace, or Sysdig
  • Understanding of cost management, inventory management, and FinOps models

Responsibilities

  • Monitor and troubleshoot Container platform (OpenShift), Rancher (RKE) and Azure (AKS) environment performance issues, connectivity issues, and security issues
  • Perform deep dives into systemic and latent reliability issues, including incident and problem management
  • Identify, analyze, and resolve infrastructure vulnerabilities and application deployment issues
  • Perform blameless RCA and partner with engineering and operation teams to roll out fixes
  • Manage application onboarding and provide troubleshooting support through the lifecycle of applications
  • Identify and drive opportunities to improve automation to reduce TOIL and improve operational excellence
  • Partner with risk and compliance teams to implement controls and remediate vulnerabilities
  • Ensure resiliency during implementation and identify/fix resiliency problems by collaborating with engineering teams
  • Act as a key stakeholder in the design of cloud services, working with Architecture, engineering, and product teams
  • Participate in 24x7 on-call coverage in a follow-the-sun model

Benefits

  • Access to hundreds of clients, most who have been working with Genesis10 for 5-20+ years.
  • The opportunity to have a career-home in Genesis10; many of our consultants have been working exclusively with Genesis10 for years.
  • Access to an experienced, caring recruiting team (more than 7 years of experience, on average.)
  • Behavioral Health Platform
  • Medical, Dental, Vision
  • Health Savings Account
  • Voluntary Hospital Indemnity (Critical Illness & Accident)
  • Voluntary Term Life Insurance
  • 401K
  • Sick Pay (for applicable states/municipalities)
  • Commuter Benefits (Dallas, NYC, SF, and Illinois)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service