Senior Site Reliability Engineer

Wesco•Dacula, GA
•Hybrid

About The Position

As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities. Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and long-term success.

Requirements

  • Associate degree (U.S.)/College Diploma (Canada) or equivalent combination of education and technical experience required.
  • 5+ years of experience in Infrastructure Engineering, Platform Engineering, Site Reliability Engineering (SRE), Systems Administration, or related technical roles.
  • Experience deploying, supporting, or validating AI, GPU, HPC, or large-scale enterprise infrastructure environments.
  • Experience with Kubernetes, container platforms, and enterprise Linux administration.
  • Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
  • Experience with VMware ESXi, Hyper-V, KVM, or related virtualization technologies.
  • Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.
  • Knowledge of Infrastructure-as-Code and automated deployment methodologies.
  • Demonstrated troubleshooting, root-cause analysis, and problem-solving skills.
  • Possess a customer-centric mindset and strong written and verbal communication skills.
  • Possess intermediate computer skills, including proficiency with Microsoft Office applications.
  • Ability to travel up to 25%.

Nice To Haves

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline preferred.
  • Experience with NVIDIA GPU technologies, CUDA, AI Enterprise, or related AI infrastructure platforms preferred.
  • Knowledge of InfiniBand, RDMA, RoCE, or high-performance networking technologies preferred.
  • Certified Kubernetes Administrator (CKA)
  • Red Hat Certified System Administrator (RHCSA) or equivalent Linux certification
  • NVIDIA certifications related to AI, GPU, or DGX platforms
  • VMware Certified Professional (VCP) or equivalent

Responsibilities

  • Provide technical expertise and engagement to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
  • Deploy, configure, and validate AI, GPU, and High Performance Computing (HPC) infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
  • Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, AI Enterprise prerequisites, and telemetry solutions.
  • Validate accelerated networking technologies including InfiniBand, RoCE, RDMA, and GPU-to-GPU communications.
  • Perform infrastructure readiness assessments, burn-in testing, operational acceptance testing, and performance validation activities.
  • Configure and support server infrastructure including iDRAC, iLO, BMC, firmware, storage, and networking components.
  • Deploy and administer Windows, Linux, VMware ESXi, Hyper-V, and KVM-based environments.
  • Apply security hardening standards, compliance requirements, and operational best practices throughout deployment and validation activities.
  • Develop and maintain automation workflows utilizing PowerShell, Python, Bash, and Infrastructure-as-Code methodologies.
  • Create customer-facing deployment documentation, technical reports, readiness assessments, and operational validation deliverables.
  • Troubleshoot complex hardware, operating system, virtualization, containerization, networking, and AI platform issues.
  • Participate in advanced technical training and continued education to maintain expertise in cloud, infrastructure, AI, and platform technologies.
  • Support technical engagements across customer environments and collaborate with internal engineering, architecture, and service delivery teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service