Principal AI Platform Engineer

PEMCOSeattle, WA
$113,879 - $215,063Remote

About The Position

The Principal AI Platform Engineer role will build and run PEMCO's AI platform: the operational, economic, and governance layer under every AI agent and model the company uses. You own the stack end to end, from GPU compute through model serving, retrieval, and orchestration up to observability and governance, and you are expected to be capable of building and operating it on-premise when the economics justify it, not only consuming managed cloud services. This is a hands-on principal role with organizational influence: you build the systems yourself while setting the standards others build against.

Requirements

  • 6+ years in platform engineering, SRE, or technical operations at senior/lead scope
  • 2+ years running LLM/AI systems in production (or 4+ years ML platform operations)
  • Hands-on open-weight model serving (vLLM, Triton/TensorRT-LLM or equivalent), including quantization and GPU right-sizing
  • LLM observability and evaluation experience (LangFuse, LangSmith, Arize class), or the demonstrated ability to stand it up
  • Experience building RAG systems: vector stores, embedding pipelines, document processing
  • FinOps/cost management for cloud consumption, ideally GPU or AI workloads
  • Working knowledge of AI security risks: prompt injection, data leakage, model abuse

Nice To Haves

  • on-prem GPU infrastructure design
  • fine-tuning/LoRA
  • agent frameworks (MCP, LangChain/LangGraph, Semantic Kernel)
  • regulated-industry experience

Responsibilities

  • Own AI economics and observability. Build the consumption telemetry for the entire AI estate (cost per agent, per workflow, per business outcome; token usage; GPU utilization) and the LLM observability layer under it (LangFuse-class tracing, evaluation, drift monitoring). Leadership decisions about AI spend run on your data.
  • Run the AI gateway. A single gateway fronting every provider and model (LiteLLM-class or Azure APIM GenAI): routing, fallback chains, quotas, and per-agent cost capture. This is the enforcement point for the economics.
  • Stand up open-weight serving. High-throughput serving of open-weight models on Azure GPU capacity (vLLM, Triton/TensorRT-LLM), quantization, right-sizing. You build the sizing evidence that justifies or kills any future on-prem investment.
  • Build the retrieval layer. Vector stores, embedding pipelines, and document processing for in-house AI builds, on a governed platform rather than one-off deployments.
  • Operate production agents. Deployment, monitoring, incident response, and retirement for the agent fleet (MCP-based orchestration, per-agent least-privilege identity). When an agent supporting a business workflow fails, you own recovery.
  • Shape the architecture. Represent platform reliability, security, and economics in AI solution reviews across teams; constructively challenge designs with data and propose alternatives.
  • Hold the governance line. RBAC for models and agents, prompt/output guardrails, a seat on the AI Governance Working Group with authority to block deployments that do not meet the bar.
  • Compute and acceleration stack: Azure GPU VMs and AKS GPU pools first; on-premise GPU build-out (hardware selection, CUDA stack, Kubernetes GPU scheduling) when the sizing data says so
  • Models and serving stack: open-weight model families with high-throughput serving (vLLM, NVIDIA Triton/TensorRT-LLM), quantization, fine-tuning and LoRA adaptation; managed frontier APIs (Azure OpenAI / AI Foundry) as the other half of the portfolio
  • Gateway and routing stack: a single AI gateway fronting every provider and model; routing and fallback chains, quotas, per-agent cost capture
  • Retrieval and data stack: vector stores, embedding and chunking pipelines, document processing, knowledge-source governance
  • Orchestration and agents stack: agent frameworks and protocols (MCP, LangChain/LangGraph, Semantic Kernel), tool registration, per-agent least-privilege identity
  • Observability, evaluation, governance stack: LangFuse-class tracing and cost telemetry, evaluation tooling with regression testing before prompt or model changes ship, guardrails, RBAC
  • Across all layers: model lifecycle from evaluation to retirement, and cost-against-capability optimization (model selection and routing, caching, batching, token budgets).

Benefits

  • medical
  • dental
  • vision
  • employer-paid basic life and accidental death & dismemberment insurance policies
  • long- and short-term disability benefit coverages
  • 401(k) plan with a generous employer match (2 for 1 on the first 6% employee pre-tax and/or Roth deferral, up to federal maximums)
  • Vacation
  • eight (8) paid holidays
  • four (4) floating holidays
  • up to ten (10) days of sick leave
  • paid time off for bereavement
  • paid time off for jury duty
  • paid time off for employee volunteering in the community
  • Education Assistance Program after one year of service
  • Scholarship program for children of PEMCO employees after one year of service
  • children’s birthday gift program
  • Flexible Spending Accounts
  • Employee Assistance Program
  • charitable gift matching
  • Discretionary bonuses
  • Tiered sales commissions and/or incentives
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service