Staff Software Engineer - AI/ML Platform

Fanatics CommerceRedwood City, CA

About The Position

Fanatics Commerce is the global leader in licensed sports merchandise, operating a vertically integrated platform that designs, manufactures, and delivers officially licensed apparel, jerseys, headwear, and collectibles for major leagues, teams, and events worldwide. With more than 900 e-commerce sites and a global omnichannel presence across digital, in-venue, and retail, Fanatics Commerce reaches fans in over 180 countries and powers official fan experiences for many of the world's most iconic sports properties. At Fanatics, we bring our BOLD Leadership Principles to life every day - building championship teams, obsessing over fans, acting with entrepreneurial speed, and delivering with a determined and relentless mindset. About the team The AI/ML Platform pod builds Fanatics' internal AI platform: the paved road for every team building with LLMs, MCPs, and agents. Gateway, agent runtime, registry, retrieval, evaluation, and cost analytics, all running on a petabyte-scale lakehouse with enterprise-grade governance. You'll join a pod that owns this platform end to end. This is a Staff-level role, and a hands-on one. You will choose the architecture, define the standards other teams build to, and stay in the code, shipping the hardest parts yourself. The decisions you make are ones the organization lives with for years.

Requirements

  • 8 to 12 years building production software, including 2+ years on AI/ML platform.
  • Bachelor's or Master's in Computer Science, Engineering, or a related field.
  • Staff-level impact: system design across multiple services, decisions that span teams, and platforms other engineers build on.
  • Strong Python and FastAPI, plus one of Java, GoLang, Scala, or TypeScript.
  • Hands-on with Terraform, Kubernetes, Airflow, Postgres, Grafana, and Prometheus, and familiar enough with the AWS data science toolkit and libraries like PyTorch to partner credibly with data scientists.
  • Hands-on with LLMs and generative AI on Bedrock or Vertex: model routing, RAG, tool calling, structured outputs, and context engineering.
  • Agentic systems and MCP server and client integrations.
  • Agent runtimes such as AWS Bedrock AgentCore matter here, and because they are new we look for depth rather than years.
  • The data engineering behind retrieval: chunking, embedding, and indexing pipelines, vector or search backends such as OpenSearch, open lakehouse table formats, and a working method for measuring retrieval quality.
  • LLM evaluation and observability: eval dataset design, tracing, LLM-as-judge scoring, and regression testing, with tools such as Langfuse or Arize Phoenix.
  • Security and governance for AI platforms: RBAC, OAuth2/OIDC, least privilege, cost optimization, and usage metering.
  • Writes clearly and holds the room with executives.

Nice To Haves

  • Chat UI or internal developer portal work is welcome.
  • Exposure to self-hosted serving with vLLM, SGLang, or llama.cpp is welcome.
  • Familiarity with governance systems like DataHub and semantic layers, and with administering enterprise AI tools such as ChatGPT Enterprise, Claude Enterprise, or Glean, is welcome.
  • Open-source contributions or published work in agentic AI or LLMOps is welcome.

Responsibilities

  • Own the end-to-end architecture of the AI control plane: gateway, agent runtime, registry, retrieval, evaluation, and cost. Set technical direction, sequence the roadmap, and make the build-versus-buy calls.
  • Build the enterprise LLM gateway: multi-provider routing, failover, caching, rate limiting, token optimization, and key management, with a path to self-hosted open-weight models where cost, latency, or data residency call for it.
  • Build the Agent/MCP ecosystem on a managed runtime such as AWS Bedrock AgentCore: enterprise systems exposed as governed tools over MCP, a control-plane registry for discovery, ownership, versioning, and lifecycle, and a no-code path from prototype to governed production agent, surfaced through a company-wide chat portal.
  • Own the unified retrieval layer and the data platform behind it: the governed front door plus the ingestion, chunking, embedding, and indexing pipelines and the vector and search backends that feed it, kept reproducible, observable, and quality-gated at petabyte scale.
  • Build the evaluation and feedback loop: offline evals, tracing, and regression gates in CI/CD, with production failures and user edits turning into new eval cases and root-cause attribution across prompts, retrieval, tools, and models.
  • Own governance and cost: RBAC, agent identity, least-privilege access to enterprise data, audit logging, use-case-level cost attribution, and the enterprise AI tool portfolio, standing up to finance, executive, and EU AI Act scrutiny.
  • Raise the bar: set the engineering standards, mentor engineers, and represent the architecture and its trade-offs directly to senior leadership

Benefits

  • For information about our benefits, please visit https://benefitsatfanatics.com/
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service