Staff ML Engineer

Press Ganey
•$130,000 - $190,000

About The Position

We are seeking a Staff ML Engineer to lead the design, delivery, and operation of production-grade AI systems. You will build the services and high-throughput pipelines behind our text analytics platform and a growing portfolio of AI products, processing feedback and conversations from every channel our customers use to engage with patients and their customers. The central challenge is sustaining service reliability and measurable AI quality as models, inputs, and volumes change. As a technical leader, you will guide engineers through design and implementation, contribute directly to the codebase, and lead architectural decisions across teams. Working primarily with Python and cloud-based data and AI platforms, with a growing emphasis on Databricks, you will shape the shared infrastructure supporting both established products and new AI applications.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline, or equivalent practical experience.
  • 8+ years of professional software engineering experience, including at least 3 years owning ML or LLM systems in production and their operational support.
  • Proven track record of independently leading complex technical initiatives from requirements through production, making architectural decisions and coordinating delivery across stakeholders.
  • Advanced proficiency in Python for production services and data processing, with strong SQL skills.
  • Experience designing and operating high-throughput distributed systems, with a strong understanding of failure recovery, multi-tenancy, and capacity planning.
  • Hands-on experience deploying and operating LLM-based applications, including evaluation, output validation, observability, and cost management.
  • Strong production engineering practices across automated testing, CI/CD, monitoring, incident response, and root-cause analysis.
  • Demonstrated technical leadership through system design, hands-on implementation, code review, and mentorship.
  • Ability to communicate technical decisions and tradeoffs clearly to engineering, research, product, and governance stakeholders.

Nice To Haves

  • Experience with Databricks or comparable cloud-based data and AI platforms for workflow orchestration, scalable processing, model deployment, and evaluation.
  • Experience with NLP, text analytics, or large-scale processing of unstructured data.
  • Experience building shared infrastructure for inference, evaluation, and model lifecycle management.
  • Familiarity with retrieval-augmented generation, semantic search, and LLM orchestration frameworks.
  • Experience with speech-to-text, speaker diarization, or processing conversational audio data.
  • Experience deploying and operating cloud-native services on AWS or Azure.
  • Experience with healthcare or other regulated environments, including sensitive data handling, auditability, and model governance.

Responsibilities

  • Own technical delivery from prototype through deployment and ongoing production support, partnering with AI Scientists and Product to define requirements, plan implementation, and resolve cross-team dependencies.
  • Evaluate proposed AI solutions for production suitability, identify technical risks, and choose architectures that meet quality, reliability, and cost requirements without unnecessary complexity.
  • Architect and build high-volume AI services and processing pipelines, including fault tolerance, backpressure, retries, idempotency, and recovery from partial failures.
  • Lead the evolution of our Python services and Databricks-based platform for distributed processing, model serving, and integration of traditional ML and LLM-based components.
  • Establish production engineering standards for automated testing, CI/CD, model and prompt versioning, load testing, controlled rollouts, and rollback.
  • Build evaluation and monitoring capabilities to detect AI quality regressions and track service reliability, throughput, latency, and inference cost.
  • Partner with Product and Responsible AI teams to define release criteria and implement requirements for model validation, data privacy, security, and governance.
  • Optimize processing and inference workloads, balancing model quality, throughput, latency, capacity, and cost.
  • Mentor engineers and lead architecture and code reviews, maintaining consistent standards for software quality and maintainability.

Benefits

  • competitive benefits package
  • discretionary bonus or commission tied to achieved results
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service