Cloud Platform Architect

SambaNovaSan Jose, CA
$245,000 - $325,000

About The Position

The Cloud Operations team is seeking an experienced engineering leader to scale the platform our internal and external customers use to access SambaNova RDUs. In this role, you'll be architecting our next-generation system from the ground up, running the Kubernetes infrastructure that powers some of the most advanced AI workloads in the industry, and bridging multi-cloud and on-prem environments in ways no generic SaaS company can offer. Your work will directly impact the productivity of every engineer at SambaNova and by extension, the speed at which we ship the future of AI computing.

Requirements

  • 7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
  • Proficiency in at least one programming language (e.g., Python, Go, Rust)
  • Expertise with Kubernetes (EKS, GKE, or self-managed) in production environments - pods, operators, CRDs, CNIs, etc.
  • Expertise with Infrastructure as Code with the ability to manage complex, multi-cloud environments
  • Strong proficiency with at least one major cloud provider (AWS, GCP, or Azure), with a solid understanding of the core services (compute, storage, networking, IAM)
  • Networking fundamentals (TCP/IP, DNS, HTTP, load balancing) and security best practices in the cloud

Nice To Haves

  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure
  • Experience managing infrastructure for data-intensive or ML/AI workloads
  • Knowledge of building and maintaining CI/CD pipelines (e.g., GitLab CI, Jenkins, ArgoCD)
  • Experience with service mesh technologies (e.g., Istio, Linkerd)
  • Contributions to open-source projects or a public portfolio of code (GitHub)

Responsibilities

  • Architect, build, and maintain our next-generation internal developer platform, automating and streamlining our cloud and on-prem infrastructure
  • Design, write, and manage Terraform modules to provision and manage resources across AWS, GCP, and Azure, ensuring consistency and reproducibility
  • Build and manage highly available, secure, and performant Kubernetes clusters that serve as the primary runtime for our diverse AI workloads
  • Design and implement robust networking solutions (VPCs, load balancers, firewalls, service meshes) that seamlessly connect our multi-cloud and hybrid environments
  • Collaborate with AI and software engineering teams to understand their needs, provide golden paths to production, and build internal tools that accelerate their development cycles
  • Implement best practices for observability (monitoring, logging, tracing) to ensure system reliability and performance, and participate in on-call rotation

Benefits

  • 95% premium coverage for employee medical insurance
  • 77% premium coverage for dependents
  • Health Savings Account (HSA) with employer contribution
  • Dental insurance
  • Vision insurance
  • Short/Long term Disability insurance
  • Basic Life insurance
  • Voluntary Life insurance
  • AD&D insurance
  • Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care
  • Headspace subscription
  • Gympass+ membership with access to physical gyms
  • One Medical membership
  • Counseling services with an Employee Assistance Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service