Staff Platform Engineer - AI Platform

Facility Grid LLC
$180,000 - $225,000Remote

About The Position

Facility Grid builds commissioning and turnover software for the teams that bring large buildings online: data centers, airports, hospitals, and commercial real estate. Our platform is the system of record proving every piece of equipment was installed, tested, and accepted. Founded in 2012 and backed by Nexa Equity, we are about 30 people with a small, onshore engineering team that is growing. We are building an agentic product suite on top of our system of record, and we run engineering the same way: agents that qualify changes, watch production, and share context across the team. This role owns the platform that makes both possible. You take a modern production platform (ECS Fargate, Aurora, GitOps with ArgoCD and Crossplane, SigNoz) and turn it into an AI platform for running and observing agents in production. You will set the direction for how we operate software in an agentic world and you will build it. You also keep production reliable and cheap. This is a hands-on staff role with a lot of autonomy. You decide how the work gets done. You write code every day, mostly through agents you direct.

Requirements

  • Ownership. You take problems end to end without waiting to be asked. You own what you ship all the way to production and you are accountable for it there.
  • Curiosity about code. You read application code as well as infrastructure, and you want to understand how the whole system works and how it fails.
  • Forward-looking engineering. You already work agent-first with tools like Claude Code, and you design the platform for agents to operate as well as people. You expect this to keep changing every few months and you keep up. Using ChatGPT for lookups or Copilot for autocomplete does not meet this bar. Informed skepticism about where these tools fail is welcome.
  • Technical fundamentals. Distributed systems, networking, Linux, and AWS. You can explain how autoscaling, queues, and deploys behave under load and what a change does to a running system.

Responsibilities

  • Build the internal AI platform: harnesses that qualify changes, progressive rollout, and shared context for every engineer’s agents.
  • Move operations from reactive to proactive: agents that baseline the system and flag anomalies before anyone gets paged.
  • Own the production platform and the path to production, including the ephemeral environment every merge request gets.
  • Own reliability, cloud cost, and compliance (SOC 2 today, FedRAMP in progress) with measured results.
  • Evaluate new models, runtimes, and tooling as they ship and adopt what is better. We stay model-independent.

Benefits

  • Health insurance
  • Paid time off
  • Vision insurance
  • 401(k) matching
  • Bonus based on performance
  • Dental insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service