Senior AI Platform Engineer

Code Metal•Boston, MA
•$170,000 - $210,000•Hybrid

About The Position

Code Metal's engineering teams are building AI-driven code transpilation, modernization, optimization, modeling and simulation solutions. They all need the same foundations: models to serve, agents to run, context to manage, and results to measure. Our AI Platform team builds those foundations. As a Senior Software Engineer on the AI Platform team, you'll design, build, and operate parts of the AI enablement stack our engineers depend on: GPU inference serving, model gateways, agent harnesses, context engineering, observability, and AI experimentation management. It is an internal platform, but we're building it to product standard. Over time you will own 1 area, such as model serving, agentic infrastructure, or experimentation and telemetry, while you also contribute across the rest of the stack. We don't expect you to arrive with every skill on this page. This is an engineering role first. Most of your time goes to designing, building, and operating production systems. You'll also need solid data science and AI research fundamentals: you'll work closely with our Applied AI Research team, and you'll sometimes run experiments when a platform decision needs evidence.

Requirements

  • Production-grade Python and solid platform engineering fundamentals: API and service design, distributed systems, containers and orchestration (such as Kubernetes), CI/CD, and testing.
  • Production experience in at least 1 focus area: Model serving: deploying and tuning LLM inference, or running an LLM gateway or API gateway. Agentic infrastructure: shipping agentic systems, or building context-engineering services such as retrieval-augmented generation, embeddings, and vector or hybrid search. Observability and experimentation: instrumenting and operating services against SLOs, or building evaluation and experiment-tracking systems.
  • Solid data science and AI research fundamentals: how transformers and LLM inference work, experiment design, benchmarking, and model evaluation. Working familiarity with PyTorch and Hugging Face.
  • Experience owning a service or component in production, including debugging and improving its reliability.
  • Experience writing design docs for focused projects, reviewing code, and mentoring or onboarding teammates.

Nice To Haves

  • Experience in more than 1 focus area.
  • Production experience running an LLM gateway or proxy such as SMG or Bifrost, or building an API gateway, including routing, auth, rate limiting, quotas, failover, and cost attribution.
  • Familiarity with inference optimization: speculative decoding, prefix caching, tensor/pipeline/expert parallelism, GPU profiling.
  • Experience with durable workflows, agent frameworks, or tool protocols such as MCP.
  • Experience fine-tuning language models, including distributed training across multiple GPU nodes.
  • Experience with experiment-tracking or artifact management systems such as MLflow or Weights & Biases.
  • Experience deploying AI systems on-prem or in regulated domains such as defense or aerospace.

Responsibilities

  • Own 1 or more platform components end to end, from design doc to production operation.
  • Build and operate the services in your focus area, and contribute across the rest of the stack.
  • Write clean, well-tested code, and hold AI-generated code to the same bar: correct, modular, and fully tested.
  • Debug hard production issues in the components you own, such as tail latency, GPU memory pressure, or agent runs that fail or loop.
  • Scope and estimate efforts with internal customers, and raise risks early.
  • Review teammates' code and designs, and help onboard new engineers.
  • Run focused benchmarks and experiments when a platform decision needs evidence.

Benefits

  • Pay depends on experience, but we strive to be at the upper end of the salary range
  • Health care plan with 100% premium coverage, including medical, dental, and vision
  • 401k with 5% matching
  • Paid Time Off (uncapped vacation, plus sick and public holidays)
  • Flexible hybrid or remote work arrangement
  • Relocation assistance for qualifying employees
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service