About The Position

We are looking for a Senior Software Engineer, Machine Learning to help build and operate the infrastructure that trains, hosts, and serves Etsy's machine learning models — including our internal serving platform, and a growing portfolio of predictive and generative (open source LLM) models. You won't just be deploying models; you'll be architecting the systems that make model serving fast, reliable, and scalable for millions of inferences per second. Your work directly enables Applied Scientists and Engineers across Etsy to train, host, and iterate on models with confidence, from classic ML predictors to large open-source language models. This is a full-time position in ML Enablement team, reporting to its Engineering Manager.

Requirements

  • Bachelor's degree in Computer Science, Applied Statistics, Mathematics, Electrical Engineering, or a related quantitative field, or equivalent professional experience.
  • 5+ years of professional experience building, iterating on, and troubleshooting complex backend and infrastructure systems.
  • Strong software engineering fundamentals, including solid command of algorithms and data structures, with the ability to write production-ready code in Python.
  • Hands-on experience with cloud infrastructure (Google Cloud preferred) and Kubernetes, including deploying and operating production workloads.
  • Familiarity with observability tooling (metrics, logging, tracing) for monitoring and debugging distributed systems.
  • Working knowledge of machine learning fundamentals and concepts, with awareness of LLM serving infrastructure and hosting open-source models.
  • Basic understanding of transformer architectures, predictors, and how model design choices affect serving infrastructure.
  • Comfort with system design for large-scale, high-availability infrastructure.

Responsibilities

  • Write high-quality, production-grade Python code (additional languages a plus); participate in code reviews and pair programming; contribute to and help drive architectural decisions.
  • Build, operate, and improve infrastructure for training and serving ML models on GCP and Kubernetes, with a focus on scalability, reliability, and observability.
  • Design and maintain serving infrastructure for open-source and internally hosted LLMs, including model deployment, resource management, and performance tuning.
  • Apply working knowledge of ML fundamentals — including neural network deep learning as well as latest transformer architectures along with prediction and inference systems — to make sound infrastructure and design decisions.
  • Partner cross-functionally with Applied Scientists to understand model training and serving needs, and translate that understanding into infrastructure that removes friction from their workflows.
  • Contribute to system design discussions, weighing tradeoffs across performance, cost, and reliability for infrastructure serving 100M+ users.
  • Thoughtfully use generative AI and other productivity tools to work efficiently, with a focus on learning and intentional contribution.

Benefits

  • equity package
  • annual performance bonus
  • competitive benefits that support you and your family
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service