Senior DevOps Engineer, AI Platform

FloQastSan Jose, CA

About The Position

FloQast is seeking a Senior DevOps Engineer to specialize in AI infrastructure. This role is crucial because FloQast's AI products have outgrown existing infrastructure patterns. The engineer will embed with the Transform and Close AI pods, owning the AI runtime for products like Transform, AI Matching, and AutoBuilder. This is a hands-on DevOps role focused on AI infrastructure, not research or modeling. The primary goal is to ensure the systems serving AI models are fast, observable, multi-region, cost-bounded, and auditable. The role involves managing AWS Bedrock, sandboxed execution environments, infrastructure as code (Terraform), observability for AI workloads, cost engineering, CI/CD pipelines, reliability, on-call duties, and ensuring security and compliance for the AI stack.

Requirements

  • 5+ years in DevOps, SRE, platform, or infrastructure engineering, with experience in production systems and on-call rotations.
  • Deep, hands-on AWS experience with ECS/Fargate, Lambda, SQS, S3, IAM, VPC and networking, ALB/NLB.
  • Experience with Terraform at production scale, including modules, state management, multi-region, and multi-account setups.
  • Experience with CI/CD and container ownership, preferably GitHub Actions, Docker, image supply chain, and scaling policies.
  • Production AI infrastructure experience, having operated at least one LLM-backed or ML-serving workload in production and understanding concepts like tokens, latency, throttling, and cost.
  • Experience with managed foundation-model APIs (e.g., Bedrock, Vertex, Azure OpenAI, Anthropic), serving platforms (e.g., SageMaker, KServe, Ray Serve), agent or tool-execution runtimes, and eval harnesses in CI.
  • Observability practice with tools like Grafana, Prometheus, or OpenTelemetry, distributed tracing, and the ability to define user-experience-focused SLOs.
  • Working fluency in Python or TypeScript/Node.js, at a level sufficient for reading services, debugging, and submitting PRs, with the ability to read the other language.
  • Understanding of multi-region infrastructure under data-residency constraints.

Nice To Haves

  • AI-specific cost and performance work, including token accounting, prompt caching, batching, model routing, and right-sizing serving capacity.
  • Experience with sandboxed execution of untrusted or generated code.
  • Experience with Terraform orchestration layers (e.g., Atmos) and monorepo build systems (e.g., NX, Turborepo, Bazel).
  • Experience with progressive delivery using feature flags (e.g., Harness).
  • Experience with FinOps tooling (e.g., CloudZero) and per-tenant cost attribution.
  • Exposure to data infrastructure such as MongoDB, PostgreSQL, Snowflake, EMR/Spark.
  • Experience producing audit evidence for SOC 2 / ISO 27001, or delivering and explaining cloud-cost reductions.
  • Prior work in a regulated or audited SaaS domain like fintech, accounting, or healthcare.

Responsibilities

  • Own the AI runtime infrastructure across AWS Bedrock and Bedrock AgentCore in US, EU, and AU regions, including model access, throughput, cross-region inference, quotas, throttles, and region-appropriate model availability.
  • Define standards for how product teams call, retry, budget, log, and trace models.
  • Operate execution environments for AI-generated code, managing session lifecycle, network controls, and least-privilege IAM.
  • Write and review Terraform for a multi-account, multi-region AWS estate, ensuring all AI resources are managed as code.
  • Extend the Grafana platform with AI-specific signals such as token consumption, latency distributions, throttle and retry rates, tool-call failure taxonomy, sandbox session outcomes, generation success rate, and end-to-end agent traces.
  • Define SLOs against critical user journeys for AI systems.
  • Manage AI spend as a first-class cost line, including model inference, serving capacity, sandbox compute, and telemetry volume, with detailed tagging for attribution.
  • Build GitHub Actions pipelines for Node/TypeScript and Python services in NX monorepos, ensuring model and prompt changes are versioned, gated, feature-flagged, and reversible.
  • Join the DevOps on-call rotation, developing runbooks for AI-specific failure modes.
  • Ensure the AI stack meets SOC 2 and ISO 27001/42001 evidence requirements, including audit logging, encryption, least-privilege access, patching, and asset inventory for model endpoints and sandboxes.
  • Enforce tenant isolation on all AI paths, including prompts, retrieved context, generated code, and logs.

Benefits

  • Medical, Dental, Vision
  • Family Forming benefits
  • Life & Disability Insurance
  • Unlimited Vacation
  • Employee Stock Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service