Staff DevOps Engineer

Nexxa.AISunnyvale, CA

About The Position

Nexxa is seeking a Senior/Staff DevOps Engineer to build and operate the infrastructure for AI and industrial systems. This role involves ensuring the reliability, observability, and resilience of production ML and data workloads, from GPU-backed clusters to pipelines connecting to industrial environments. The ideal candidate will have deep infrastructure ownership experience, where uptime and reliability directly impact physical operations. You will collaborate with AI, data, and product engineering teams to ensure their systems run safely and at scale in production.

Requirements

  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
  • Deep hands-on experience with Cloud platforms (AWS, GCP, or Azure) at production scale
  • Kubernetes in production, including GPU workload scheduling
  • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)
  • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
  • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
  • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
  • Proven ability to independently scope and lead infrastructure projects from design through production rollout
  • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause

Nice To Haves

  • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines
  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)
  • History of building internal developer platforms or self-service infrastructure tooling
  • Experience scaling infrastructure teams or setting technical direction at a Staff level

Responsibilities

  • Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end
  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
  • Partner with data and AI teams to support the infrastructure behind: Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks), Feature stores, embedding indices, and retrieval pipelines, Model training, evaluation, and serving infrastructure
  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity
  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy
  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

Benefits

  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service