Senior Principal Software Engineer, Machine Learning

Toast
$222,000 - $453,000Hybrid

About The Position

The Machine Learning Platform team builds and operates the core infrastructure that powers AI and ML across Toast — the feature store, model hosting and serving, the experimentation platform, training pipelines, and the tooling ML engineers and data scientists rely on every day. Our work directly enables the models that drive personalization, forecasting, fraud detection, search, and the growing set of AI-powered experiences shipping to restaurants. Toast is seeking a Senior Principal Software Engineer to act as the technical leader on the ML Platform team, shaping the systems that will carry Toast's AI and ML capabilities into the next decade. The role involves driving architectural direction across the platform, delivering foundational infrastructure that other teams build on, and elevating fellow engineers. The ideal candidate is a domain expert with considerable ML platform experience, who partners with ML engineers, data scientists, product, and infrastructure leadership on high-leverage opportunities. This position suits an engineer comfortable writing production code, leading technical design for distributed systems, and influencing organizational decisions about how Toast builds and deploys ML.

Requirements

  • 10+ years delivering complex backend or infrastructure systems at scale
  • Direct experience building or operating core ML infrastructure — feature stores, model serving, experimentation platforms, training orchestration, or equivalent
  • Mastery of a modern backend language, ideally Java or Kotlin
  • Deep proficiency with distributed systems concepts: consistency, latency, throughput, fault tolerance, and observability
  • Strong understanding of data modeling, query languages, and the online/offline data patterns that underpin ML systems
  • Demonstrated technical leadership, with ability to drive cross-team alignment and influence engineering, product, and business stakeholders
  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience

Nice To Haves

  • Hands-on experience with open-source or commercial ML platform components (e.g. Tecton, MLflow, SageMaker, Databricks)
  • Experience building or operating experimentation / A-B testing platforms at scale
  • Familiarity with real-time streaming systems (Kafka, Flink, Spark Streaming) and their use in feature computation
  • Experience serving LLMs or deep-learning models in production, including GPU capacity planning and inference optimization
  • Prior work supporting internal-developer-facing platforms with a product mindset

Responsibilities

  • Own technical direction of the ML Platform — feature store, model hosting and serving, experimentation, training infrastructure — driving architectural decisions around scalability, reliability, latency, and cost
  • Lead design and delivery of large-scope platform initiatives from conception through production, coordinating across ML, data, and infrastructure teams
  • Identify and resolve systemic technical challenges: online/offline feature parity, model deployment friction, experimentation velocity, GPU utilization, cross-team dependencies
  • Set and maintain a high engineering quality bar through hands-on code contributions, design reviews, and mentorship of platform and ML-adjacent engineers
  • Partner with ML engineering, data science, product, and platform leadership to translate ML strategy into technical roadmaps
  • Define the paved paths ML teams use to ship models safely — from feature registration through canary rollout, monitoring, and rollback
  • Leverage AI-augmented development tools to increase development velocity and code quality

Benefits

  • Competitive compensation and benefits programs
  • Healthy lifestyle with flexibility
  • Cash compensation (overtime, bonus/commissions if eligible)
  • Equity
  • Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service