About The Position

Alkira is transforming networking through its Network Infrastructure as a Service (NIaaS) platform, enabling enterprises to build and operate global cloud networks with simplicity, scale, and agility. We are seeking a Senior Platform Infrastructure Software Engineer to own the foundational layer every other engineering team at Alkira depends on, including Platform Services. Platform Infrastructure does not own customer-facing features. We build the primitives other engineering teams need in order to work at all: the global Kubernetes cluster footprint, the internal cluster running our CI/CD runners and build platforms, the shared infrastructure teams consume (databases, object storage, container registries), and the lifecycle of the shared services everything else is built on. We also build the Terraform modules other teams use to ship their own infrastructure quickly and safely. Our customers are engineers, and the output is code, Terraform configuration/modules, Go and Python tooling, cluster and service automation. You will own complex projects, drive architectural decisions, and set the standards other engineers work within.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field.
  • 5+ years of professional experience in infrastructure, platform, DevOps, or SRE engineering, including production ownership.
  • Deep production Kubernetes experience operating clusters, not only deploying to them.
  • Strong Terraform or OpenTofu experience at scale: module design intended for other teams to consume, state management, and policy-as-code.
  • Strong programming ability in Go, Python, or a comparable language.
  • Experience running stateful shared services/infrastructure in production (message broker, databases, certificate authority, etc...) or comparable.
  • Deep Linux fundamentals: process and resource behavior, and system-level debugging under load.
  • Understanding of networking fundamentals including TCP/IP, DNS, HTTP, load balancing, and service-to-service communication.
  • Hands-on experience with at least one major cloud provider and the ability to work across others.
  • Sound judgment about blast radius, and the discipline to make risky changes reversible.
  • Strong technical judgment, problem-solving, collaboration, and communication skills.

Nice To Haves

  • Experience on an internal platform or developer platform team, where the users were other engineers.
  • Experience with Scalr or other IaC platforms, including policy-as-code with OPA.
  • GitOps with ArgoCD or Flux; Helm, Kustomize, or operator development.
  • Managing self-hosted CI runners.
  • Deep understanding of Prometheus/Thanos or equivalent, including writing exporters and instrumentation standards.
  • Experience with observability platforms, telemetry systems, logging, distributed tracing, and monitoring architectures.
  • Experience building highly available, multi-tenant SaaS platforms operating at significant scale.
  • Proven ability to influence engineering strategy across organizations and mentor technical talent.
  • Carrying production on-call for what you built.

Responsibilities

  • Own the global Kubernetes cluster footprint across regions and cloud providers: provisioning, upgrades, multi-tenancy, resource and network policy, and failure isolation.
  • Own the internal cluster running CI/CD runners and build platforms: capacity, isolation, scaling, and the build performance every team feels daily.
  • Own the shared infrastructure other teams consume - managed databases, object storage buckets, container registries, and build platforms including provisioning paths, access model, and cost.
  • Own the full lifecycle of shared platform services such as internal PKI, vaults, Kubernetes operators.
  • Own disaster recovery for the platform: backup and restore, replication and failover design, documented RTO and RPO targets, and regular DR exercises that prove recovery works.
  • Design and publish reusable Terraform/OpenTofu modules that let other teams deliver their own infrastructure faster, with correctness and policy built in rather than reviewed in. Treat adoption by other teams as the measure of whether a module succeeded.
  • Own the observability stack and set the standard other teams instrument against.
  • Build production tooling in Go or Python. A significant share of this role is software engineering, not configuration.
  • Mentor engineers through code review, design review, and technical guidance.
  • Drive down cost and eliminate drift: orphaned resources, stale state, pipelines outliving the services they built.
  • Participate in the on-call rotation for the Platform Infrastructure.

Benefits

  • The salary range for this position is $ 84,629- $ 155,152 USD.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service