Senior Machine Learning Engineer, Infrastructure

PatreonSan Francisco, CA
$212,000 - $318,000Hybrid

About The Position

Patreon is seeking a Senior Machine Learning Engineer, Infrastructure to join their Relevance team. This role will focus on building and scaling ML systems that power creator and content discovery on the Patreon platform, including search, feed ranking, and creator-fan matching. The engineer will work closely with a collaborative team and cross-functional partners to deliver impactful infrastructure solutions. The position is based in San Francisco or New York and operates on a hybrid work model, requiring 3 days per week in the office.

Requirements

  • Deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, specifically with low-latency live inference pipelines and feature store architectures.
  • Strong background in distributed systems and backend engineering.
  • Ability to write robust, maintainable code in Python.
  • Systematic approach to debugging complex, high-throughput systems and performance bottlenecks.
  • Strong communication skills and effective at creating clear documentation for system architectures and infrastructure strategies.
  • Growth mindset, keen eye for detail in code reviews, and passion for empowering teammates by improving developer velocity.

Nice To Haves

  • Building '0 to 1' infrastructure systems that stand the test of time and provide a reliable foundation for the team.

Responsibilities

  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems.
  • Own the end-to-end feature store lifecycle—from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features.
  • Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift.
  • Collaborate with cross-functional partners, such as product, data engineering, and trust and safety, to translate product requirements into robust, scalable infrastructure solutions.
  • Automate model deployment and reliability testing to improve developer velocity and ensure system stability.
  • Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues.

Benefits

  • salary
  • equity plans
  • healthcare
  • flexible time off
  • company holidays
  • recharge days
  • commuter benefits
  • lifestyle stipends
  • learning and development stipends
  • patronage
  • parental leave
  • 401k plan with matching
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service