Senior AI Infrastructure Engineer

Tencent•Palo Alto, CA
•$124,800 - $283,800•Onsite

About The Position

We are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical evaluation and end-to-end execution of AI infrastructure deployments — from data center due diligence and solution review, through cross- functional delivery coordination, to ongoing operations of production AI environments. You will act as a key technical owner bridging internal stakeholders and external partners, ensuring infrastructure is delivered on time, to spec, and operated reliably at scale.

Requirements

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • 5+ years of experience in data center infrastructure, infrastructure deployment, or infrastructure operations.
  • Solid understanding of high-performance compute server hardware, high-speed networking (e.g., InfiniBand/RoCE), storage systems, and data center power & cooling fundamentals.
  • Proven experience evaluating data center facilities and reviewing technical proposals for compute infrastructure.
  • Strong track record coordinating complex, multi-party technical projects to closure, including working effectively with external partners and cross-regional teams.
  • Hands-on experience with cluster orchestration and management tools (e.g., Slurm, Kubernetes) and Linux system administration.
  • Proficiency in scripting/automation (Python, Bash; Ansible/Terraform a plus).
  • Familiarity with observability/monitoring stacks (Prometheus, Grafana, ELK) and DCIM tooling.
  • Bilingual proficiency in English and Mandarin highly preferred
  • Strong ownership mentality, able to independently drive projects with minimal supervision.

Nice To Haves

  • Experience with high-performance storage / parallel file systems (e.g., Lustre, GPFS/Spectrum Scale, WekaFS, VAST, Ceph) in AI infrastructure environments.
  • Experience deploying/operating large-scale AI compute environments in a hyperscale or cloud environment.
  • Data center or infrastructure certifications.

Responsibilities

  • Conduct on-site data center assessments to evaluate whether candidate facilities meet AI infrastructure requirements.
  • Review and validate AI infrastructure solution designs, identifying technical risks, gaps, and cost/performance trade-offs before sign-off.
  • Coordinate with external partners, driving timelines, resolving technical issues, and ensuring deliverables meet internal requirements.
  • Partner with internal business and engineering teams to translate requirements into deliverable technical specifications.
  • Participate in and eventually take ownership of day-2 operations for live AI environments — monitoring, incident response, capacity/health checks, firmware and lifecycle management, coordinating hardware maintenance as needed.
  • Build and improve operational standards, runbooks, and documentation for infrastructure delivery and operations to enable consistent execution across regions.
  • Provide technical guidance and mentorship to junior engineers/interns on hardware diagnostics, cluster tooling, and best practices.
  • Track and report on project status and risks to management.

Benefits

  • sign on payment
  • relocation package
  • restricted stock units
  • medical
  • dental
  • vision
  • life and disability benefits
  • 401(k) plan
  • vacation
  • holidays
  • paid sick leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service