Staff MaaS Backend Engineer

Bitdeer Technologies GroupSan Jose, CA

About The Position

We are seeking a Staff Backend Engineer to take our Model-as-a-Service (MaaS) working system and re-architect it into a commercial, globally distributed, multi-tenant token service — one that sustains at least millions of monthly active users, high sustained token throughput per GPU, and invoice-grade accounting, with scalability, reliability, and observability engineered in deliberately rather than absorbed under load. This role works directly with the principal architect: co-owning the MaaS system design as the deepest backend voice in that conversation, and owning the implementation end to end — the code that ships, the migrations that land, and the service that stays up. Design authority is shared; delivery accountability is not. The mandate is explicit: measure what exists, find where it breaks before it breaks in front of a paying customer, and carry the platform there incrementally — with each step independently shippable, reversible, and non-disruptive to the tenants already on it. The role is deeply hands-on: it reads and rewrites the existing Go services, owns SLOs and on-call for a revenue-bearing service, and sets the backend engineering standard for the MaaS team.

Requirements

  • 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers.
  • Act as a senior technical voice who can collaborate closely with architects, write rigorous design documents, and commit to executing architectural decisions effectively.
  • Demonstrate expertise in scaling distributed systems through multi-region active-active deployments, caching, backpressure, and targeted performance engineering that measurably lowers unit costs.
  • Proven track record of safely executing zero-downtime brownfield migrations for stateful subsystems—like metering or ledgers—without regressions or accounting gaps.
  • Deep hands-on proficiency with Go-based services, production Kubernetes (including Envoy and GPU-aware scheduling), and the architectural trade-offs of datastores like PostgreSQL, Redis, and Kafka.
  • Drive operational visibility by owning end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores.
  • Apply a systems-level understanding of LLM serving to manage complexities like server-sent-event streaming, KV/prefix caching, and the trade-offs between time-to-first-token (TTFT) and throughput.
  • Leverage this foundation to build highly reliable, exactly-once metering and billing systems that accurately reconcile billions of events under partial failure conditions.
  • Enforce strict multi-tenant security disciplines by designing fail-closed authorization, mandating verified identities, and guaranteeing absolute cross-tenant isolation.
  • Bring operational maturity to a revenue-bearing platform by carrying on-call responsibilities, running blameless incident reviews, and translating outages into structural improvements.

Responsibilities

  • Co-own the end-to-end MaaS system design with the Principal Architect, authoring decision records and defending technical trade-offs.
  • Drive the platform through its maturity roadmap by delivering operable, measurable capabilities rather than mere demos.
  • Lead technical execution by setting stringent Go and API standards, mentoring engineers, and aligning cross-functional teams.
  • Own the wire compatibility contract for major formats (OpenAI, Anthropic), supporting advanced features like streaming, tool calling, and structured output.
  • Evolve the routing tier to handle load-aware, model-aware, and prefix-cache-aware endpoint selection with robust circuit breaking and fallback mechanisms.
  • Run versioning and deprecation as a published contract to guarantee external customer code stability across underlying changes.
  • Maximize platform economics and performance by optimizing token throughput, KV cache tiering, and time-to-first-token (TTFT) latency at the p95/p99 levels.
  • Mature the model serving control plane by integrating deployment tooling, LoRA multiplexing, and cold-start-aware autoscaling directly with the Kubernetes fleet.
  • Treat regressions in cost-per-million-tokens or latency metrics as critical system incidents.
  • Scale the platform to a globally distributed architecture featuring regional inference pools, capacity-aware failovers, and an active-active control plane.
  • Define, publish, and rigorously defend strict Service Level Objectives (SLOs) baselined against actual system performance rather than aspirations.
  • Ensure operational resilience through peak-concurrency load testing, robust on-call runbooks, and predictable load-shedding during overloads.
  • Enforce fail-closed authorization, robust multi-tenant isolation, and zero-retention data paths across the network, cache, and storage layers.
  • Manage the complete lifecycle of API keys and OAuth credentials while distributedly enforcing rate limits and quotas without relying on client-supplied identifiers.
  • Design strict abuse, rate, and prompt-injection controls, treating all user and model-generated content as untrusted data.
  • Build an idempotent, exactly-once metering system to capture uncached, cached, output, and reasoning tokens accurately across all requests.
  • Maintain a high-volume usage ledger that enforces prepaid spend caps asynchronously and reconciles perfectly with the invoicing system.
  • Safeguard business integrity by treating any metering or billing defect as a critical revenue and trust incident.
  • Execute zero-regression, incremental platform upgrades using strangler-style replacements, shadow traffic, and stateful dual-writes.
  • Deliver end-to-end request tracing and cost telemetry, defining internal schemas for token and quality attributes.
  • Enforce absolute log hygiene by ensuring prompts, completions, and PII are never retained outside of explicit, consented policies.

Benefits

  • Equal employment opportunities in accordance with country, state, and local laws.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service