About The Position

Machine Learning/Artificial Intelligence powers innovation in all areas of the business, from helping members choose the right title for them through personalization, to better understanding our audience and our content slate, to optimizing our payment processing and other revenue-focused initiatives. Building highly scalable and differentiated AI infrastructure is key to accelerating this innovation. More recently, fast-paced innovation in large language models (LLMs) has greatly helped advance state-of-the-art technology in many areas of personalization, including search and recommendation experiences. The Model Serving Systems team provides the computational platform on which we build nearly all our consumer and studio-facing AI/ML applications. We provide all the building blocks to serve AI/ML models at scale, including a real-time model inference and serving platform, foundational abstractions that ensure consistency between online and offline systems, and more. Additionally, as we expand to enable LLM innovation in numerous areas of personalization, we’re building model serving infrastructure for LLMs and other large foundation models. We're expanding our model serving systems to meet the evolving needs of the evolving AI landscape. We are looking for strong engineers to develop and expand our compute infrastructure to support the growing AI needs, enable the application of ML in new business areas, and drive AI/ML innovation across Netflix. Our systems power some of Netflix's most business-critical models, and we need you to take our AI/ML initiatives to the next level. You will play a highly cross-functional role, partnering with other engineers, product managers, machine learning engineers, and data/research scientists.

Requirements

  • Experience building high-traffic distributed services and infrastructure for online ML model inference.
  • Familiarity with supporting large-scale ML models focusing on high availability and performance.
  • Understanding of scalable model-serving solutions for generative models and LLMs, with skills in reducing latency and costs, and ability to solve bottlenecks to streamline research-to-production workflows.
  • Proficiency in object-oriented programming (preferably Java) and engineering excellence in production hosting, including performance tuning, deployment management, and capacity planning.
  • Familiarity with deploying ML models using tools like Triton Inference Server, TensorRT, Docker.
  • Experience working with public cloud like AWS, Azure, or GCP.
  • Proactive communicator who promotes best practices in observability and logging.
  • BS/MS in Computer Science, Applied Math, Engineering, or a related field.

Responsibilities

  • Develop and expand compute infrastructure to support growing AI needs.
  • Enable the application of ML in new business areas.
  • Drive AI/ML innovation across Netflix.
  • Partner with other engineers, product managers, machine learning engineers, and data/research scientists in a cross-functional role.

Benefits

  • Health Plans
  • Mental Health support
  • 401(k) Retirement Plan with employer match
  • Stock Option Program
  • Disability Programs
  • Health Savings and Flexible Spending Accounts
  • Family-forming benefits
  • Life and Serious Injury Benefits
  • Paid leave of absence programs
  • Flexible time off (for full-time salaried employees)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service