SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer Technologies GroupSan Jose, CA
Onsite

About The Position

NeoCloud is building an AI-operated GPU cloud. This role is the first human in the loop, serving as the escalation target when the AIOps system needs a decision and the source of ground truth that turns novel incidents into new automations. In this L1 role, you will cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8 AM–8 PM PST shift. You will execute SOPs, escalate complex cases, and provide the AIOps system with the ground truth it needs to learn from novel incidents. This role is different from a traditional NOC job because the platform is the last line of defense, and you serve as the training signal. Strong L1s have a growth path into SME roles or the platform team as automation authors.

Requirements

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8 AM-8 PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8 AM and 8 PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
  • Tag and describe every novel incident resolved to provide data for the platform team, enabling future automation.
  • Ensure runbooks evolve towards platform executability rather than manual execution.
  • Provide structured signal through handoff notes, not free-form email.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service