About The Position

Lawrence Livermore National Laboratory (LLNL) is seeking a High-Performance Computing (HPC) Data Center Operator to monitor, diagnose, troubleshoot, and repair system faults on a large number of HPC systems, storage systems, and networks with minimal supervision. The role involves interacting with other Livermore Computing (LC) staff to remediate problems and provide advanced technical support in a complex HPC computing and networking environment. This position requires full-time on-site presence and will be filled at either an intermediate or advanced level based on experience. The intermediate level (525.2) focuses on providing intermediate technical support and operational monitoring, while the advanced level (525.3) requires advanced technical support, independent troubleshooting of complex issues, and serving as a resource to team members.

Requirements

  • Ability to obtain and maintain a U.S. DOE Q-level security clearance which requires U.S. Citizenship.
  • Associate’s degree in a computer-related field or equivalent combination of technical training and experience.
  • General working knowledge of networking concepts, protocols, and connectivity troubleshooting.
  • Demonstrated experience using Linux command-line utilities for system monitoring, routine troubleshooting, and routine administration tasks.
  • General working knowledge of enterprise storage concepts such as NAS, RAID, and file system operations.
  • Experience installing, racking, inspecting, replacing, and performing routing troubleshooting and repair of server, storage, or network hardware components in a data center environment.
  • Proficient verbal and written communication skills necessary to interact with customers and team members with the ability to work independently and as a member of a team.
  • Experience and knowledge of the skills needed for a customer support role to include a focus on listening, rapport-building, friendly and approachable nature, and courtesy and patience.
  • Ability to work all shifts, including weekends and holidays.
  • Significant experience developing, refining, and documenting work processes, operational procedures, technical instructions, or troubleshooting workflows (at 525.3 level).
  • Significant experience administering, maintaining, repairing, securing, and troubleshooting systems in large scale production environments involving Linux, clustered computing, or networked storage, with minimal supervision (at 525.3 level).
  • Demonstrated experience independently diagnosing and resolving moderately complex hardware, software, operating system, storage, and network issues (at 525.3 level).

Nice To Haves

  • Advanced training, certifications and experience in Linux operating systems.
  • Advanced knowledge and experience working in a datacenter.
  • Computer Support experience or advanced experience in desktop support.

Responsibilities

  • Provide intermediate technical support and operational monitoring for HPC systems including large Linux clusters, file systems, storage systems, and associated infrastructure.
  • Apply working knowledge of Linux/Unix systems and use in-house and vendor-supplied tools to monitor systems, diagnose issues, and perform routine repairs and recovery actions.
  • Utilize the Laboratory’s trouble ticketing system, ServiceNow, for problem ticket tracking.
  • Receive, document, triage, and respond to customer issues during business and off-hours, resolving routine problems or escalating to appropriate technical staff.
  • Perform data center facilities monitoring, problem remediation, and emergency event response during normal daily operation and off-hours.
  • Participate in system installation, hardware swaps, system relocation, and decommissioning activities in support of ongoing data center operations.
  • Promote the use of inter-departmental resources for tools, metrics, and common solutions to team members via email and presentations.
  • Perform other duties as assigned.
  • Provide advanced technical support and monitoring capabilities for the HPC systems clusters, file systems, and storage systems under minimal supervision (at 525.3 level).
  • Independently troubleshoot and resolve moderately complex hardware, software, operating system, network, and infrastructure issues; analyze symptoms, determine root cause, implement corrective actions, and coordinate escalation when required (at 525.3 level).
  • Perform advanced technical tasks including installation, diagnosis, repair and maintenance of clustered computer systems and related file systems and networks (at 525.3 level).
  • Analyze system events, alarms, logs, and monitoring data to identify patterns, isolate faults, and recommend improvements to procedures, tools, or operational practices (at 525.3 level).
  • Serve as a knowledgeable resource to team members during off-hours operations and contribute to the development, refinement, and documentation of operating procedures and response practices (at 525.3 level).

Benefits

  • Flexible Benefits Package
  • 401(k)
  • Relocation Assistance
  • Education Reimbursement Program
  • Flexible schedules (depending on project needs)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service