About The Position

Nscale is seeking a Staff Cloud Native Software Engineer to design, build, operate, and enhance the cloud-native software integrations that connect AI applications and networking components at scale. This role involves working with shared Kubernetes-based platforms, deployment patterns, observability foundations, infrastructure architecture, and operational tooling to ensure internal teams can run services safely and efficiently on GPU-backed infrastructure. The engineer will collaborate with platform engineering, infrastructure, and product teams to meet developer and operational needs, contributing significantly to the reliability, scalability, and usability of Nscale's software integrations. As a Staff engineer, responsibilities include taking ownership of major components, setting technical direction across teams, delivering complex technical work independently, and improving operations and engineering quality through practical improvements, sound technical judgment, and mentoring.

Requirements

  • At least 8 years of experience in production-level software development.
  • Deep hands-on experience building and operating Kubernetes-native software: custom controllers, operators, CRDs, or admission webhooks, using controller-runtime, client-go, or equivalent.
  • Strong understanding of Kubernetes internals: the API server, informer/lister patterns, reconciliation loops, and the object model.
  • Strong networking fundamentals: CNI, service mesh, kube-proxy/eBPF datapaths, DNS, load balancing, and experience building software that integrates with these systems.
  • Proficiency in Go (strongly preferred) or a similar language, with a track record of shipping well-tested, production-quality code at scale.
  • Experience with observability practices: metrics, tracing, structured logging, built into software rather than added afterward.
  • Comfortable owning components independently end-to-end, from design through operation, while setting direction for adjacent teams.
  • Track record of leading technical design at a staff level and mentoring engineers through practical technical guidance.

Nice To Haves

  • Experience with or strong interest in GPU-backed infrastructure and AI workload patterns.

Responsibilities

  • Design, build, and operate Kubernetes-native software, including controllers, operators, custom resources (CRDs), and admission webhooks, to connect AI applications with core networking components on GPU-backed infrastructure.
  • Extend Kubernetes control-plane capabilities to support AI workload requirements, such as network policy controllers, CNI/service-mesh integrations, and resource/scheduling extensions.
  • Own significant components end-to-end and set the technical direction for their design, deployment, and operation across the team.
  • Build reconciliation loops, informers, and client-go–based tooling to maintain infrastructure state consistency between the API server, networking systems, and AI runtime components.
  • Develop operational tooling and automation to simplify the deployment, operation, and support of Kubernetes-native services for internal teams.
  • Drive infrastructure architecture decisions regarding the integration of AI applications and networking components across the platform, considering cross-team trade-offs.
  • Build observability foundations for controller and operator software, including metrics, structured events, tracing, and status reporting via the Kubernetes API and platform dashboards.
  • Design systems that degrade gracefully and self-heal using controller patterns like reconciliation and backoff.
  • Debug and resolve complex issues spanning the Kubernetes control plane, networking (CNI, service mesh, kube-proxy/eBPF datapaths), and workload runtime behavior on GPU-backed infrastructure.
  • Define standards for the safe rollout of controller and platform changes, including versioning, compatibility, and staged deployment.
  • Set technical direction for the team's approach to building Kubernetes-native software, establishing patterns for controller design, CRD schema evolution, and testing.
  • Lead design discussions and code reviews, maintaining high standards for Kubernetes API conventions and idiomatic client-go usage.
  • Partner with platform engineering, infrastructure, and product teams to translate needs into CRDs, APIs, and controller-managed abstractions.
  • Define reusable patterns, shared libraries, and scaffolding to enable other teams to build correctly on the platform.
  • Mentor engineers in Kubernetes internals, controller-runtime patterns, and operational judgment.

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service