SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer Technologies GroupSan Jose, CA
Onsite

About The Position

NeoCloud is building an AI-operated GPU cloud. This L1 role involves front-line monitoring and incident response for NeoCloud's US GPU Data Centers during the 8 AM–8 PM PST shift. You will execute Standard Operating Procedures (SOPs), escalate complex issues, and provide ground truth data to the AIOps system to enable learning and automation. This role is distinct from a traditional NOC job, as the platform is the first line of defense, and your primary function is to train and improve the platform's automation capabilities. There is a real growth path into SME roles or the platform team as automation authors.

Requirements

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8 AM-8 PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8 AM and 8 PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
  • Tag and describe novel incidents to feed the AIOps substrate for automation.
  • Contribute to making runbooks executable by the platform.
  • Provide structured signal through handoff notes.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service