Data Center Infrastructure Software Engineer

Designworks TalentBellevue, WA
Hybrid

About The Position

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking Data Center Software Engineers to lead the design, development, configuration, and automation of AI infrastructure clusters. Your responsibility begins once servers and racks are installed in the data center and extends through software deployment, networking, configuration, cluster bring-up, and automation, ensuring the platform is fully operational and ready for customer workloads.

Requirements

  • 5+ years of experience designing, building, or operating large-scale Linux-based infrastructure.
  • Hands-on experience with Kubernetes, containerization, and distributed systems in production environments.
  • Experience with infrastructure-as-code and automation tools such as Terraform, Ansible, or similar frameworks.
  • Strong experience operating cloud or datacenter-scale infrastructure.

Nice To Haves

  • Experience with bare-metal provisioning and hardware lifecycle management platforms (e.g., MAAS, Ironic, xCAT, Foreman, or similar).
  • Experience with IPMI, Redfish, PXE boot, and automated operating system deployment at scale.
  • Experience managing GPU clusters in datacenter or cloud environments.

Responsibilities

  • Develop infrastructure-as-code, automation, and provisioning systems for compute, networking, and storage.
  • Deploy and optimize Kubernetes, container, and distributed computing platforms.
  • Optimize GPU, networking, storage, and system performance for large-scale AI workloads.
  • Troubleshoot complex issues across hardware, operating systems, networking, storage, and software stacks.
  • Build reliability, observability, and operational excellence practices for mission-critical infrastructure

Benefits

  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives.
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service