AI Cloud Senior DevOps Engineer

Bitdeer Technologies GroupSan Jose, CA

About The Position

We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role, you will be the backbone of our deployment and infrastructure operations, ensuring that our AI-powered products and platforms are delivered with speed, security, and exceptional reliability. You will act as a crucial bridge between our research/development teams and real-world deployment, driving automation, optimizing cloud-native architectures, and establishing best practices for MLOps and traditional DevOps workflows.

Requirements

  • Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
  • Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
  • Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices.
  • Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies.
  • Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development.
  • Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles.
  • Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment.

Nice To Haves

  • Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads.
  • Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms).
  • Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity.
  • Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001).
  • Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams.

Responsibilities

  • Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models.
  • Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
  • Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker.
  • Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
  • Take ownership of high-availability design in production environments.
  • Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
  • Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
  • Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics.
  • Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency through Internal Developer Platforms (IDP) and Platform Engineering initiatives.
  • Establish and enforce robust system stability and security standards.
  • Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance with industry frameworks (e.g., SOC2, ISO27001).
  • Act as the technical lead during complex system anomalies and major incidents.
  • Spearhead rapid troubleshooting, conduct thorough root cause analysis (RCA), and implement preventative remediation plans.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service