Data Center Operations Engineer - Chicago ORD

LambdaElk Grove Village, IL
56dOnsite

About The Position

Lambda, The Superintelligence Cloud, builds Gigawatt-scale AI Factories for Training and Inference. Lambda’s mission is to make compute as ubiquitous as electricity and give every person access to artificial intelligence. One person, one GPU. If you'd like to build the world's best deep learning cloud, join us. Note: This position requires presence in our Chicago/Elk Grove Village Data Center location 5 days per week. The Operations team is at the heart of keeping our AI-IaaS infrastructure running smoothly from start to finish. They handle everything from sourcing the right hardware and components to keeping our data centers performing at their best day in and day out. The team also works closely across the company, making sure our operational capabilities stay in sync with product goals and overall strategy. By managing the entire lifecycle — from procurement through deployment and ongoing efficiency — the Operations team ensures our AI infrastructure stays reliable, scalable, and ready to support the business as it grows.

Requirements

  • Sourcing & Procurement
  • Researching, evaluating, and securing the right hardware and infrastructure components.
  • Building relationships with peers and supply chain to ensure cost-effective and timely supply.
  • Data Center Operations
  • Monitoring day-to-day performance of data centers to maintain uptime and efficiency.
  • Troubleshooting and resolving hardware or infrastructure issues quickly.
  • Performing regular maintenance and upgrades to keep systems running at peak performance.
  • Deployment & Lifecycle Management
  • Overseeing the full lifecycle of infrastructure, from initial setup to ongoing optimization.
  • Coordinating deployments of new hardware and ensuring seamless integration with existing systems.
  • Managing capacity planning to make sure infrastructure can scale with business growth.
  • Cross-Team Collaboration
  • Working with product management, support, and other teams to align operational capabilities with company goals.
  • Translating business priorities into technical and operational requirements.
  • Supporting cross-functional projects where infrastructure plays a critical role.
  • Reliability & Scalability
  • Ensuring infrastructure remains stable, secure, and scalable as demand increases.
  • Continuously improving processes to boost efficiency and reduce downtime risks.

Nice To Haves

  • Certifications: Any Linux or project management.
  • Military background.
  • Experience in the machine learning or computer hardware industry

Responsibilities

  • Make sure new servers, storage, and networking gear are racked, labeled, cabled, and configured the right way.
  • Keep data center layouts and network topologies up to date in our DCIM software.
  • Coordinate with supply chain and manufacturing teams so systems are deployed on time, especially for large-scale projects.
  • Evaluate current and future data center needs based on growth and technology trends.
  • Manage parts depot inventory and track equipment as it moves from delivery → storage → staging → deployment → handoff.
  • Work closely with hardware support teams to get tickets resolved quickly.
  • Create and manage RMA tickets when needed, making sure faulty parts are replaced and reinstalled without delay.
  • Develop and maintain installation standards (placement, labeling, cabling) to ensure consistency across all data centers.
  • Act as a subject matter expert on data center deployments, supporting sales engagements for major deployments in our facilities or at customer sites.

Benefits

  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and Commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible Paid Time Off Plan that we all actually use
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service