Principal DevOps Engineer - Azure

TENEX.AISan Jose, CA
Hybrid

About The Position

TENEX is seeking a Principal DevOps Engineer to be a key technical leader responsible for the architecture, evolution, and operation of our Azure infrastructure, CI/CD pipelines, and Site Reliability Engineering (SRE) practices. The goal is to keep the platform highly available, secure, and performant as it scales to handle petabytes of security data and billions of daily events. This role involves close collaboration with Software Engineering, AI/ML, and Security Operations teams to define the technical vision and architecture for production systems, driving automation and operational excellence. The ideal candidate will have deep, hands-on Azure platform expertise, a strong software engineering foundation, and fluency in DevSecOps principles. As an early employee in a fast-growing, well-funded startup, this role offers the opportunity to play a meaningful role in shaping the company culture and culture is highly prioritized, with a preference for in-person collaboration.

Requirements

  • 8+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering roles.
  • Deep, hands-on expertise building production Azure platforms, including AKS, Entra ID and workload identity federation, VNet design and Private Link, Key Vault, Azure Policy, and subscription or landing zone architecture. This is a platform engineering role rather than a Microsoft 365, Intune, or Windows administration role.
  • Experience standing up or moving production workloads across cloud environments, with the ability to describe the design, the data path, the cutover, and what broke.
  • Experience building secure and compliant (e.g., SOC 2, ISO 27001) environments.
  • Deep understanding of microservices architecture, containerization (Docker, Kubernetes), and event-driven systems.
  • Extensive experience with Infrastructure-as-Code tools (e.g., Terraform, Bicep) and CI/CD best practices.
  • Production experience writing and shipping software in Go or Python, beyond scripting and configuration.
  • Experience with monitoring and observability tools (Prometheus, Grafana, Azure Monitor, ELK stack, or similar).
  • Familiarity with real-time data pipelines and stream processing (e.g., Kafka, Event Hubs, Service Bus, Pub/Sub).
  • Proven track record of architecting, building, and operating highly scalable, distributed, and secure enterprise-grade SaaS platforms.
  • 10-12 years of experience, Bachelor's or Master's degree in Computer Science, Engineering, and or years of relative experience

Nice To Haves

  • Working knowledge of more than one major cloud provider, deep enough to judge where the provider models differ rather than assume they match.
  • Prior experience in cybersecurity (SIEM, EDR, SOAR, or MDR) or an MSSP environment.
  • Experience with large-scale data warehousing/lakehouse technologies (e.g., Azure Data Explorer, Microsoft Fabric, Snowflake, BigQuery).
  • Background leading technical initiatives in high-growth startups or enterprise SaaS.
  • Familiarity with the underlying infrastructure to support AI/ML model deployment and monitoring (MLOps).
  • Relevant certifications (Azure Solutions Architect Expert, Kubernetes, or security-related credentials) are a plus. Certifications complement production depth and do not substitute for it.

Responsibilities

  • Own the architecture of our Azure platform as it scales to petabytes of security data and billions of daily events.
  • Own our Azure governance and environment model, including subscription and management group structure, Azure Policy, network topology, and identity.
  • Lead Site Reliability Engineering (SRE) initiatives, defining and driving adherence to critical Service Level Objectives (SLOs) and Service Level Indicators (SLIs), and managing on-call rotations.
  • Drive operational excellence by implementing advanced monitoring, observability (logs, metrics, tracing), automated provisioning, and disaster recovery strategies.
  • Establish and enforce DevSecOps practices, standardizing CI/CD pipelines, infrastructure-as-code (IaC), security testing, and deployment mechanisms for rapid, secure, and reliable software delivery.
  • Automate deployment, scaling, and management of microservices and event-driven systems using containerization and orchestration technologies (Docker, Kubernetes, AKS).
  • Maintain workload portability across the platform so that infrastructure decisions remain reversible.
  • Partner with engineering teams to optimize application performance, resource utilization, and cloud cost efficiency.
  • Mentor and influence engineering teams on best practices in Azure architecture, reliability, and security-first development.
  • Collaborate with Product Management and Security Operations to translate new product requirements and operational needs into scalable and cost-effective platform solutions.
  • Evaluate and drive the adoption of new infrastructure technologies and engineering methodologies to maintain a competitive advantage.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service