Server Architect

Base Power Company•Austin, TX
•Onsite

About The Position

Base is America’s next-generation power company, focused on deploying distributed batteries to create a resilient and abundant energy system. The hardest problem in AI compute today is power – permitted, interconnected, and available now. Base has this power at thousands of sites. This role involves building a distributed GPU fleet that connects datacenter-grade compute to Base's power network, serving the AI industry through hosted hardware, bare-metal nodes, and an OpenAI-compatible inference API. This is a founding-team role where you will help build the entire system, including the control plane for managing remote nodes, the provisioning path for servers, a secure agent for node management, and a dispatch layer that integrates with energy planning systems.

Requirements

  • 8+ years building systems software close to hardware: platform management, firmware, provisioning, fleet orchestration, or backend services that operate physical machines.
  • Deep experience with server platforms — BMC/iDRAC, Redfish or IPMI, PXE/virtual-media boot, hardware inventory, and remote recovery.
  • Strong C, C++, Rust, or Go, and comfort moving between a boot log and a distributed service in the same day.
  • Experience operating fleets you couldn’t walk up to: servers, network gear, vehicles, telecom, or energy hardware.
  • Judgment about what to build versus adopt — this project leans on industry standards and existing supply chains, not one-off inventions.
  • Ownership. Small team, new business line, real revenue targets. You’ll set patterns others follow.

Nice To Haves

  • Datacenter or cloud platform background: hypervisors, bare-metal clouds, provisioning at scale.
  • GPU serving experience: vLLM, inference routing, KV-cache-aware scheduling, or GPU fleet operations.
  • Compiler, toolchain, or OS-image build depth.
  • Exposure to power systems or energy markets.

Responsibilities

  • Build the fleet control plane: inventory, identity, activation, telemetry, and operator tooling for a growing network of GPU nodes in the field.
  • Own out-of-band provisioning and recovery — Redfish/iDRAC automation, network boot, immutable OS images, and the rescue paths that make ad hoc SSH sessions unnecessary.
  • Develop the node agent: a secure, durable local substrate that establishes trust with the cloud, reports state, receives work, and survives reboots and bad networks.
  • Design overlay networking and secure connectivity across consumer internet links, and make it boring.
  • Connect compute to power: integrate dispatch with our energy planning systems so nodes run when power is available and cheap, and back off when the home or the grid needs it.
  • Work with hardware and deployment teams on enclosures, thermals, and the realities of running servers outdoors in a Texas summer.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service