Senior Software Engineer - Cloud Infrastructure - 1817

PlacingITRedwood City, CA
Hybrid

About The Position

We're looking for a Senior Software Engineer – Cloud Infrastructure to help design, build, and scale the cloud platform that powers a rapidly growing AI organization. This is a highly technical, hands-on engineering role where you'll architect production infrastructure, write production-quality code, and build systems that support large-scale compute and machine learning workloads. You'll join a small, collaborative engineering team with significant ownership and visibility, working across cloud infrastructure, Kubernetes, automation, and developer tooling. This role is ideal for engineers who enjoy building infrastructure from the ground up—not simply maintaining existing systems.

Requirements

  • 4+ years of experience in Cloud Infrastructure Engineering.
  • Experience designing and building cloud infrastructure at production scale.
  • Strong experience architecting highly available, fault-tolerant distributed systems.
  • Hands-on experience operating Kubernetes in production environments.
  • Proven experience implementing Infrastructure-as-Code across multiple environments.
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related engineering discipline.
  • Programming Python (production-level software development)
  • Bash
  • AWS (required)
  • Kubernetes
  • Docker
  • Helm
  • Terraform
  • Pulumi
  • CloudFormation
  • CI/CD pipelines
  • Infrastructure monitoring
  • Observability
  • Systems debugging
  • Production incident response
  • Performance optimization

Nice To Haves

  • Go (preferred)
  • Google Cloud Platform (GCP)
  • Microsoft Azure
  • Oracle Cloud Infrastructure (OCI)
  • On-premises infrastructure
  • Experience supporting AI or machine learning infrastructure.
  • Experience designing infrastructure for distributed compute workloads.
  • Multi-cloud architecture experience.
  • Strong systems design background.
  • Experience working in fast-paced startup environments.
  • Proven ability to mentor engineers while remaining deeply hands-on.
  • Exceptional Python programming skills beyond scripting or automation.
  • Deep expertise in cloud infrastructure and distributed systems.
  • Strong Infrastructure-as-Code experience.
  • High ownership and accountability.
  • Excellent systems design and troubleshooting abilities.
  • Startup mentality with the ability to move quickly and independently.
  • Strong communication and collaboration skills.

Responsibilities

  • Design and build highly available, scalable cloud infrastructure supporting production workloads.
  • Architect and manage multi-cloud environments, with AWS as the primary cloud platform.
  • Develop Infrastructure-as-Code using Terraform or Pulumi.
  • Design, deploy, and manage Kubernetes clusters supporting compute-intensive applications.
  • Write production-quality Python code to build infrastructure services and automation.
  • Improve CI/CD pipelines and streamline software deployments.
  • Build monitoring, alerting, logging, and observability solutions to ensure platform reliability.
  • Collaborate closely with engineering teams to support large-scale distributed systems and machine learning infrastructure.
  • Optimize infrastructure performance, scalability, reliability, and cloud costs.
  • Respond to production incidents and implement long-term solutions that improve operational stability.
  • Mentor engineers and contribute to infrastructure best practices while remaining a hands-on individual contributor.

Benefits

  • Equity
  • Base Salary
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service