Machine Learning Eval Engineer

Recruiting From ScratchSan Francisco, CA
Onsite

About The Position

As a Machine Learning Eval Engineer, you'll build the evaluation infrastructure that determines model quality, identifies failure modes, and drives improvements across production AI systems while working directly with machine learning, platform, and customer-facing teams. This is a rare opportunity to join one of the fastest-growing AI infrastructure companies where you'll directly influence how enterprise AI systems are measured, improved, and deployed at internet scale.

Requirements

  • 1–5 years of experience in Machine Learning Engineering, Software Engineering, or ML Infrastructure
  • Strong sweet spot around 2–4 years of experience
  • Experience building evaluation systems, ML tooling, or data infrastructure from zero-to-one
  • Experience working at high-bar technology companies, AI startups, quantitative firms, or leading research organizations
  • Experience working with production LLM applications
  • Experience building customer-facing ML tooling or internal AI platforms
  • Startup experience strongly preferred
  • Demonstrated ownership of high-impact technical initiatives
  • Strong Python engineering skills
  • Deep understanding of LLM evaluation methodologies including LLM-as-a-Judge
  • Strong prompt engineering experience
  • Strong understanding of precision, recall, statistical evaluation, and ML metrics
  • Experience building evaluation pipelines or benchmarking systems
  • Comfortable building lightweight web applications using Flask, TypeScript, or similar frameworks
  • Experience working with unstructured data including documents, PDFs, OCR, or document extraction
  • Strong debugging, experimentation, and software engineering fundamentals
  • Bachelor's degree in Computer Science, Mathematics, Physics, Machine Learning, or related technical field preferred
  • Strong academic background from a top engineering or quantitative program preferred
  • Formal machine learning education or research experience preferred

Nice To Haves

  • Familiarity with AWS S3, OLAP systems, Tinybird, or analytics infrastructure preferred
  • Experience working with Vision-Language Models or document AI preferred
  • AI infrastructure startups
  • LLM platform companies
  • Document AI companies
  • Machine learning platform teams
  • Quantitative trading firms
  • AI research organizations
  • Early-stage venture-backed startups
  • Evaluation infrastructure teams
  • Data infrastructure organizations
  • Engineers building production AI systems

Responsibilities

  • Design and build scalable evaluation systems for production LLM applications
  • Develop benchmarks, metrics, and automated evaluation pipelines measuring model quality
  • Build workflows that identify failure modes across large-scale unstructured datasets
  • Design statistical evaluation methodologies using precision, recall, and model quality metrics
  • Build internal tooling and lightweight applications for model visualization and evaluation analysis
  • Work hands-on with enterprise documents including PDFs, spreadsheets, OCR outputs, and unstructured data
  • Partner closely with ML engineers to prioritize model improvements using evaluation insights
  • Build customer-specific benchmarks demonstrating model performance across real-world workflows
  • Design evaluation infrastructure supporting production AI systems operating at massive scale
  • Collaborate with GTM, Product, and Engineering teams to communicate model performance
  • Prototype new evaluation techniques leveraging LLM-as-a-Judge methodologies
  • Own evaluation systems from initial design through production deployment

Benefits

  • Competitive Equity Package
  • Direct collaboration with ML and founding teams
  • Significant ownership over evaluation infrastructure
  • Opportunity to define model quality across enterprise AI systems
  • High-impact engineering role
  • Exposure to cutting-edge LLM and document AI technologies
  • Rapid career growth opportunities
  • Onsite collaboration with a world-class engineering team
  • Visa Sponsorship Available (Case-by-Case)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service