Senior Machine Learning Engineer, Trust

AirbnbSan Francisco, CA
$200,000 - $235,000Remote

About The Position

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Trust team is responsible for developing the technology that helps protect our community and platform from fraud while also ensuring our hosts, guests, homes, and experiences meet our high standards. We constantly work to fight against online fraud (such as monetary loss, compromised accounts, spam and scam in messages, fake inventory, etc.) as well as offline fraud (theft, property damage, personal safety, etc.). We also work on onboarding and screening of users, and think about complex topics like identity and reputation to ensure that every interaction with Airbnb helps build trust in us and our community. The Trust Frontier AI team is where new AI technology for Trust gets invented and proven. We build specialized models for mission-critical trust and safety problems, develop the AI agents and agentic capabilities that automate trust decisions, and create the benchmarks and evaluation harnesses that keep decision quality high as those agents take on more autonomy. We work on problems before the answer is known — prototyping, experimenting, and iterating with our partner teams until a solution proves itself against real business and top line metrics. You'll work side-by-side with talented product managers, data scientists, software engineers, fraud intelligence, and operations teams. Together, you'll design and build ML solutions that have direct, meaningful impact on user trust, business success, and the global Airbnb community.

Requirements

  • 5-10 years of industry experience in applied Machine Learning, with a track record of building and productionizing models at scale.
  • 1-2+ years of hands-on experience with LLMs and GenAI technologies, including building with agentic frameworks, orchestration, and evaluation.
  • Strong programming skills in Python (required) and familiarity with Scala, Java, or equivalent.
  • Solid understanding of Machine Learning best practices — e.g., training/serving skew minimization, A/B testing, feature engineering, model selection — and algorithms such as gradient boosted trees, neural networks, transformers, and deep learning.
  • Experience with ML frameworks and tooling such as TensorFlow, PyTorch, or equivalent.
  • Experience with data engineering and building end-to-end ML pipelines, including both batch and real-time systems.
  • Experience designing evaluation methodology for ML or LLM systems — benchmarks, ground truth, offline/online metrics, calibration.
  • Comfort with ambiguity and a bias toward action: you can take a loosely defined problem, scope it, prototype quickly, and drive it to a measurable outcome.
  • Exposure to architectural patterns of large, high-scale software applications (e.g., well-designed APIs, high-volume data pipelines, efficient algorithms).
  • Experience with test-driven development, incremental delivery, and deployment practices.
  • A Bachelor's, Master's, or PhD in CS/ML or a related field.

Nice To Haves

  • Experience with multimodal models (vision, document, or speech) is a plus.
  • Exposure to the Trust and Risk domain (e.g., fraud detection, anomaly detection, identity, account integrity) is a plus.

Responsibilities

  • Frame and prototype ML and agentic solutions for problems that do not yet have an established approach, in partnership with product managers, data scientists, and front line defense teams.
  • Design, build, and productionize end-to-end Machine Learning pipelines — including feature engineering, model training, evaluation, and deployment — for both batch and real-time use cases.
  • Build and improve abuse behavior detection that generalizes across defenses.
  • Design, launch, and iterate on AI agents that automate trust decisions, including orchestration, tool interfaces, and the guardrails that hold quality steady as autonomy increases.
  • Build benchmarks, evaluation harnesses, and instrumentation that let us measure agentic and model decision quality objectively, and use them to drive real improvements.
  • Develop specialized models for trust and safety use cases, and use LLMs and AI agents to accelerate how we build models.
  • Write, review, and ship clean, testable code — whether training a new model, improving an existing pipeline, or optimizing a feature for scalability and reliability.
  • Work with large-scale structured and unstructured data to continuously improve ML models for Airbnb product, business, and operational use cases.
  • Partner with front line defense teams to validate solutions through experiments and holdouts, and quantify their impact on business and operational metrics.
  • Participate in code reviews, design discussions, and cross-team collaborations to contribute to a high-quality ML engineering culture.

Benefits

  • bonus
  • equity
  • benefits
  • Employee Travel Credits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service