Manager, Cloud Engineering & Operations

SS8 Networks•Milpitas, CA
•$175,000 - $210,000•Hybrid

About The Position

SS8 is seeking an experienced Manager, Cloud Engineering & Operations to lead the architecture, engineering, deployment, and continuous improvement of SS8's cloud and cloud-native infrastructure. This is a hands-on technical leadership and people-management role with end-to-end accountability for both Cloud Engineering and Cloud Operations. You will lead a team responsible for building and operating highly available, secure, scalable, and resilient infrastructure supporting SS8's mission-critical solutions across public, private, hybrid, sovereign, customer-managed, and restricted or air-gapped environments. The ideal candidate is a player-coach who combines strong people leadership with deep technical credibility and operational accountability — equally comfortable conducting an architecture review with senior engineers, leading a major production incident, reviewing operational KPIs, and discussing cloud strategy with executive leadership. You should be able to create a culture in which reliability, security, automation, and operational ownership are engineering responsibilities from day one, not activities that begin after development is complete.

Requirements

  • 10+ years of experience in cloud infrastructure, platform engineering, software engineering, distributed systems, DevOps, SRE, or Cloud Operations.
  • 5+ years of engineering management or significant technical leadership experience.
  • Demonstrated ability to lead engineers while remaining technically engaged and hands-on when needed.
  • Experience owning production infrastructure and mission-critical services.
  • Strong architecture and design experience with large-scale distributed systems, including Kubernetes and container platforms.
  • Experience with Infrastructure as Code, automated provisioning, and CI/CD pipelines.
  • Strong knowledge of Linux, networking, compute, storage, databases, and distributed infrastructure.
  • Experience implementing observability, monitoring, logging, alerting, and telemetry platforms.
  • Strong production incident-management and troubleshooting experience, including high availability, fault tolerance, disaster recovery, and capacity planning.
  • Experience with at least one major cloud platform (AWS, Azure, GCP, or OCI).
  • Strong understanding of cloud and infrastructure security principles and best practices.
  • Ability to operate effectively across Engineering, Security, Product, Operations, and customer-facing teams.

Nice To Haves

  • Multi-cloud, private/hybrid, or sovereign cloud architecture experience.
  • Experience supporting government or national-security environments and air-gapped deployments.
  • Large-scale data platforms, distributed databases, and real-time data processing.
  • FedRAMP or comparable government compliance experience.
  • AI infrastructure, LLM/RAG infrastructure, AI Ops, or intelligent operations experience.

Responsibilities

  • Owning the architecture, engineering strategy, and technical roadmap for SS8's cloud and cloud-native infrastructure, including public/private/hybrid cloud, Kubernetes, distributed infrastructure, and disaster recovery.
  • Establishing a disciplined Cloud Operations function covering 24x7 availability, incident response, root-cause analysis, change management, and capacity/performance management.
  • Driving DevOps and SRE practices across the organization, including Infrastructure as Code, CI/CD, and automation of provisioning, monitoring, and remediation.
  • Building comprehensive observability across infrastructure, platforms, and applications — metrics, logging, tracing, and security telemetry.
  • Developing infrastructure and operational models that let SS8 products run consistently across AWS, Azure, GCP, OCI, private cloud, customer data centers, sovereign, and air-gapped environments.
  • Identifying opportunities to bring AI and automation into Cloud Operations, including intelligent alert correlation, automated triage, and anomaly detection.
  • Set technical direction and remain technically engaged through architecture reviews, design reviews, proof-of-concept work, and critical engineering initiatives — following the full lifecycle from architecture through build, automation, deployment, operation, monitoring, recovery, and continuous improvement.
  • Define operational ownership across Cloud Engineering, Software Engineering, Support, Professional Services, and Security; establish SLIs, SLOs, and KPIs, and own measurable outcomes including platform availability and reliability, MTTD/MTTA/MTTR improvement, change failure rate, deployment consistency across customer environments, and reduction in customer-impacting incidents.
  • Build, lead, mentor, and develop a high-performing Cloud Engineering & Operations team, including hiring, performance management, resource planning, and succession planning — with success measured in part by the team's growth, retention, and ability to independently own major elements of the platform.
  • Partner with Security and Product teams to embed encryption, IAM, secrets/certificate management, hardening, and vulnerability management into engineering and operational processes, tracking security and vulnerability remediation as a core outcome.
  • Establish proactive capacity planning and forecasting and optimize cloud consumption and cost without compromising reliability or performance.
  • Partner with Product Management, Software Engineering, Security, QA, Professional Services, and Technical Program Management, translating requirements into cloud engineering priorities and communicating infrastructure strategy, operational risk, and major incidents to leadership — while ensuring new product releases launch with strong operational readiness and the overall cloud technology roadmap executes successfully.

Benefits

  • medical
  • dental
  • vision
  • 401(k) with company match
  • paid time off
  • corporate bonus program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service