Cloud Systems Engineer

Alarm.com•Tysons, VA
•$100,000 - $135,000•Hybrid

About The Position

We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads. The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure.

Requirements

  • Bachelor’s degree required
  • 3-5 years of Linux systems administration experience in production environments.
  • 3-5 years of experience supporting enterprise server infrastructure.
  • Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.
  • Experience performing hardware diagnostics, firmware management, and lifecycle maintenance.
  • Experience working within datacenter operations environments.
  • Bash, Python, PowerShell, or similar scripting languages
  • Operating system performance tuning and monitoring
  • Storage and networking fundamentals
  • Experience with infrastructure monitoring and observability platforms, ex Grafana.
  • Hardware and firmware lifecycle management

Responsibilities

  • Deploy, configure, and maintain GPU-accelerated compute infrastructure.
  • Manage operating system, firmware, BIOS, BMC, driver, and software lifecycle updates.
  • Monitor system health, performance, utilization, and capacity across AI infrastructure environments.
  • Support infrastructure utilized for AI model training, inference, and data processing workloads.
  • Develop and maintain operational standards, runbooks, and maintenance procedures.
  • Participate in on-call support and incident response activities.
  • Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions.
  • Perform system patching, hardening, and operating system lifecycle management.
  • Troubleshoot operating system, kernel, storage, networking, and application-level issues.
  • Develop automation to streamline deployment, monitoring, and operational processes.
  • Support security and compliance initiatives across AI infrastructure platforms.
  • Install, configure, maintain, and troubleshoot enterprise compute hardware.
  • Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components.
  • Perform firmware upgrades and hardware lifecycle management activities.
  • Coordinate hardware replacements, vendor support engagements, and warranty services.
  • Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects.
  • Maintain accurate asset inventories and operational documentation.
  • Support high-performance networking technologies, including Ethernet and InfiniBand environments.
  • Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations.
  • Assist with scalability, resiliency, and performance optimization initiatives.
  • Perform root-cause analysis of infrastructure failures and develop preventative measures.
  • Other duties as assigned.

Benefits

  • medical plans with company subsidies
  • Health Savings Account (HSA) with a company contribution
  • 401(k) with an employer match
  • paid vacation that increases with tenure
  • paid holidays
  • wellness time
  • paid maternity and bonding leave
  • company-paid disability and life insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service