About The Position

Nscale is hiring a Staff Software Engineer to build the source of truth for how customers consume Nscale’s GPU cloud. You’ll turn usage of provisioned compute across Kubernetes, Slurm, bare metal, and storage into reliable, auditable records that power usage-based billing, credit enforcement, customer visibility, and internal cost intelligence. This role sits at the intersection of cloud infrastructure, distributed systems, billing, and platform reliability. You’ll own the domain-level architecture for metering provisioned resources such as GPU compute, Kubernetes nodes, bare-metal capacity, storage, and future platform services. Your work will ensure that every billable resource has a trustworthy usage trail: accurate enough for billing, timely enough for credit enforcement, and explainable enough for customers, finance, and engineering teams. This is an opportunity to define a foundational platform domain early, setting the metering architecture and standards that future Nscale services will build on. Cloud metering powers Nscale’s usage-based billing platform by producing accurate, deduplicated, auditable usage records for rating, credit burn-down, entitlements, cost attribution, margin analysis, and customer-facing usage dashboards.

Requirements

  • Extensive experience designing, building, and operating distributed systems in production, ideally in cloud infrastructure, data platforms, billing, control planes, or platform engineering.
  • Experience with resource lifecycle, capacity, or usage tracking systems such as compute instances, Kubernetes nodes, storage volumes, jobs, workloads, quotas, or entitlements.
  • Strong understanding of event-driven architecture, including reliable delivery, idempotency, aggregation, replay, and failure handling.
  • Strong operational discipline, including monitoring, alerting, incident response, reconciliation, data-quality checks, and post-incident improvement.
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • Strong software engineering fundamentals, with proficiency in typed backend or systems languages. Our primary stack is Go, with some services in Rust and Python.
  • Comfortable working in a fast-paced, ambiguous environment with high ownership, pragmatic judgement, and a bias toward measurable business impact.
  • You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.

Nice To Haves

  • Experience with cloud billing, chargeback/showback, prepaid credit systems, entitlements, quota enforcement, or customer-facing usage dashboards.
  • Experience with GPU cloud infrastructure, Kubernetes, bare-metal provisioning, workload scheduling, storage platforms, or AI/ML inference and training workloads.
  • Experience with billing, usage, ledger-style, event-sourced, or replayable data systems that support auditability, reconciliation, backfills, and dispute investigation.
  • Experience with Kubernetes controllers/operators, controller-runtime, CRDs, admission webhooks, or multi-cluster resource watchers.
  • Experience with infrastructure-as-code, cloud providers, regional control planes, and service catalogs or flavor/rate-card models.
  • Strong product sense for making usage data understandable to customers, finance, and support teams, especially when investigating billing disputes or consumption anomalies.

Responsibilities

  • Domain-level technical direction. Set the technical direction for platform metering across a defined domain, influencing engineering squads across Nscale that produce or consume usage data.
  • Design accurate metering models. Define how Nscale measures resources over time, including instance-hours, GPU-hours, GiB-hours, readiness states, allocation burn-down, and future shared-pool or pod-level usage models.
  • Build reliable event producers. Design and implement controllers and resource watchers that emit self-contained usage quanta for compute, Kubernetes, bare metal, storage, and other platform resources.
  • Engineer for idempotency and auditability. Ensure metering events can survive retries, redelivery, backfills, partial outages, and customer disputes through deterministic transaction IDs, provenance, traceability, and reconciliation.
  • Integrate with billing systems. Partner with billing, product, finance, and platform teams to ensure usage events can be rated, aggregated, credited, and surfaced to customers and internal stakeholders.
  • Create leverage through standards. Establish shared schemas, libraries, conventions, dashboards, and operational runbooks so that new platform services can add metering consistently.
  • Own production outcomes. Operate the metering platform with strong observability, alerting, incident response, data-quality checks, and reconciliation against billing records and infrastructure state.
  • Mentor and influence. Guide other engineers through architecture reviews, implementation choices, and operational best practices for high-integrity usage systems.

Benefits

  • medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service