Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies GroupSan Jose, CA

About The Position

NeoCloud is building an AI-operated GPU cloud — and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery. You are the SRE who stands between raw infrastructure and the customer's experience — designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.

Requirements

  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer-facing cloud service experience — defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude — you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

Nice To Haves

  • What Success Looks Like in Year 1: Customer-facing GPU cloud service SLAs published and met — availability, job completion, provisioning latency.
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
  • BMaaS live for external tenants with self-service onboarding.
  • MTTD and MTTR for customer-impacting incidents reduced through automation.
  • Tenant self-service observability live — customers can see their own job health, quota, and status.

Responsibilities

  • End-to-end reliability of the customer-facing GPU cloud service — availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity — placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty — tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling — minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.
  • Customer-Facing Ownership: You are accountable for the customer's reliability experience — when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
  • Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
  • Build self-service observability that lets customers answer their own questions — status, quota, job health — reducing support load.
  • Feed the AIOps Substrate: The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter — turning customer-impacting incidents into self-healing events.

Benefits

  • equal employment opportunities in accordance with country, state, and local laws.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service