Staff AI Platform Engineer

Code Metal•Boston, MA
•Hybrid

About The Position

Code Metal is a leader in automated software engineering, focusing on AI-driven code verification. As AI generates more code, the challenge shifts to ensuring its correctness. Code Metal addresses this by using AI for code generation and formal methods for verification, keeping engineers involved in critical decisions. Their platform ensures provably correct code with auditable proof. Customers like the U.S. Air Force, L3Harris, RTX, and Toshiba utilize Code Metal for modernizing legacy code, optimizing performance, and accelerating prototype-to-production timelines. Founded in 2023, Code Metal has offices in Boston and San Francisco and is backed by prominent investors. The Role: Code Metal's engineering teams are developing AI-driven code transpilation and AI-enabled mission planning and wargaming systems. These initiatives require a robust AI enablement stack, including model serving, agent execution, context management, and results measurement. The AI Platform team is responsible for building these foundational components. As a Staff AI Platform Engineer, you will be the technical lead for this new four-person team. Your primary responsibility will be to architect and build the AI enablement stack that engineers rely on, encompassing GPU inference serving, a model gateway, agent harnesses, context engineering, observability, and AI experimentation management. While initially an internal platform, it is being developed to product standards. This is fundamentally an engineering role, with the majority of your time dedicated to designing, building, and operating production systems. You will also need strong data science and AI research fundamentals, as you will collaborate closely with the Applied AI Research team and occasionally conduct experiments to support platform decisions.

Requirements

  • Production-grade Python and strong platform engineering fundamentals: API and service design, distributed systems, containers and Kubernetes, CI/CD, and testing.
  • Shipped production agentic systems, with a clear understanding of failure modes and reliability strategies.
  • Experience with context engineering: retrieval-augmented generation, embeddings, vector or hybrid search, and memory for agents, ideally over code or large technical corpora.
  • Experience instrumenting services (e.g., with OpenTelemetry tracing and metrics) and operating AI services against SLOs.
  • Solid data science and AI research fundamentals: understanding of transformers and LLM inference, experiment design, benchmarking, and model evaluation.
  • Working familiarity with PyTorch and Hugging Face, and experience fine-tuning, evaluating, or serving language models.
  • Staff-level technical leadership: experience owning architecture across multiple systems or teams, writing design docs and RFCs, translating ambiguous customer needs into a roadmap, and mentoring engineers.

Nice To Haves

  • Production experience running an LLM gateway or proxy such as SMG or Bifrost, or equivalent experience building an API gateway, including routing, auth, rate limiting, quotas, failover, and cost attribution.
  • Hands-on experience deploying and tuning LLM inference engines such as vLLM, SGLang, or TensorRT-LLM on GPU infrastructure, with measurable gains in throughput, latency, or cost per token.
  • Familiarity with inference optimization techniques: speculative decoding, prefix caching, tensor/pipeline/expert parallelism, disaggregated prefill and decode, GPU profiling.
  • Experience building evaluation harnesses for LLMs and agents, and experiment-tracking or artifact systems like MLflow or Weights & Biases.
  • Experience taking an internal platform to an external product: multi-tenancy, SDKs, versioned APIs, and documentation.
  • Experience deploying AI systems on-prem or in air-gapped or classified environments, or in regulated domains such as defense or aerospace.

Responsibilities

  • Set the technical direction and architecture for Code Metal's AI platform and lead the team building it.
  • Own the design docs and RFCs, assist with build-vs-buy decisions, and mentor the team.
  • Deploy, benchmark, and tune production inference for open-weight models on vLLM, SGLang, and TensorRT-LLM.
  • Own the model gateway that teams use to access self-hosted and commercial models, ensuring consistent authentication, routing, failover, quotas, and cost attribution.
  • Design reusable agent harnesses and orchestration primitives that product teams can compose into reliable, verifiable workflows.
  • Build context-engineering services for memory, retrieval, and data discovery to provide agents with the right information within their context and cost budgets.
  • Instrument the stack end-to-end with OpenTelemetry traces and service metrics, and build the experiment-tracking and artifact layer for reproducibility and comparison of results.
  • Design for productization from day one, including multi-tenancy, versioned APIs, security, and deployment in customer and air-gapped environments.
  • Partner with Applied AI Research, product teams, and DevOps to ensure the platform aligns with their needs.

Benefits

  • Pay depends on experience, but we strive to be at the upper end of the salary range
  • Health care plan with 100% premium coverage, including medical, dental, and vision
  • 401k with 5% matching
  • Paid Time Off (uncapped vacation, plus sick and public holidays)
  • Flexible hybrid or remote work arrangement
  • Relocation assistance for qualifying employees
  • Wage Transparency
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service