SRE Platform Software Engineer (Early Career / Temporary)

Bitdeer Technologies GroupSan Jose, CA
$105,000 - $155,000

About The Position

Bitdeer is building an AI-operated GPU cloud, a global fleet of self-built and OEM-rented data centers running the world's most valuable compute. This platform is designed to observe, protect, and operate the fleet. The SRE Platform team is responsible for building the monitoring and automation infrastructure that supports various other teams, including storage, network, GPU, K8S, and L1 operators. As an entry-level Software Engineer on the SRE / Monitoring Platform team, you will contribute to the NeoCloud SRE platform, a multi-region system that observes, protects, and operates a GPU rental fleet across various data centers. You will work within a team led by a senior engineer, taking well-scoped components from design to production, ensuring they meet SLOs and maintain stability. This role is focused on both building and learning, where you will write code, tests, and operate your work under guidance. You will also participate in on-call duties, initially as a shadow, with the goal of becoming independent in your assigned area within 12 months.

Requirements

  • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
  • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. Ability to write clean, tested, readable code and explain design choices.
  • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / HTTP), and operating-system concepts (processes, threads, I/O). Ability to reason about correctness and performance.
  • Distributed systems basics — understanding of idempotency, retries, back-pressure, caching, and eventual consistency. Eagerness to go deep.
  • Monitoring / observability exposure — hands-on experience with Prometheus, Grafana, Loki, or similar; ability to write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
  • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
  • Kubernetes basics — understanding of Pods, Services, Deployments; experience running something on K8s (project, lab, or internship).
  • Git + CI basics — branching, pull requests, and experience with a CI pipeline (GitHub Actions, GitLab CI, or similar).
  • Test discipline — habit of writing unit and integration tests.
  • Communication — clear written and verbal English; ability to write a good PR description and ask good questions.
  • Curiosity and a learning mindset — excitement to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice To Haves

  • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
  • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
  • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
  • Contributions to open-source observability or cloud-native projects.

Responsibilities

  • Build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, and collection-monitor.
  • Write ingestion, query, and storage-path code.
  • Contribute to alert-engine-framework, alert-correlation, and slo-framework; implement and tune default alert rules.
  • Contribute to topology-service, cluster-health-rollup, and OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
  • Help build remediation-actuator, orchestration / workflow components, inspection probes, and job-scheduler.
  • Instrument services with metrics, logs, and traces via OpenTelemetry.
  • Build dashboards.
  • Write runbooks for on-call procedures.
  • Write unit / integration / contract tests for all shipped code.
  • Participate in chaos and soak tests.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service