Forward Deployed Machine Learning Engineer

Advatix, Inc.•,
•$170,000 - $270,000•Remote

About The Position

Our Client is seeking a Forward Deployed Machine Learning Engineer (FDE MLE) with strong experience in machine learning evaluation, benchmark systems, backend infrastructure, and customer-facing engineering. This individual will be the first Machine Learning Engineer dedicated to the organization's Benchmarks and Evaluations vertical and will work closely with the General Manager, researchers, and early customers. The successful candidate will help establish the technical foundation for evaluating AI models across different domains and modalities. This is a highly hands-on role combining machine learning engineering, backend infrastructure, data pipelines, evaluation systems, and customer-facing technical delivery. The ideal candidate is comfortable operating in ambiguous, fast-moving environments and can independently take technical problems from feasibility through production delivery. Strong customer engagement skills and experience owning technical projects end-to-end are essential.

Requirements

  • Minimum 4+ years of professional engineering experience with hands-on machine learning model evaluation experience.
  • Minimum 3–8 years of relevant experience across machine learning engineering, evaluation systems, benchmark development, and end-to-end technical ownership.
  • Demonstrated experience deploying end-to-end ML evaluation or benchmark systems to production for foundation models.
  • Strong hands-on experience with ML evaluation frameworks and benchmark design, including approaches such as LLM-as-a-judge.
  • Prior ownership of backend and infrastructure systems, including: Data pipelines, Execution environments, Storage, Orchestration
  • Experience building data pipelines capable of supporting large-scale workloads.
  • Experience working with or deploying machine learning systems in production.
  • Customer-facing engineering experience, including direct interaction with enterprise customers.
  • Demonstrated ability to independently scope and execute ambiguous technical problems from feasibility through delivery.
  • Strong written communication skills for customer-facing and cross-functional technical work.
  • Strong bias toward action and ability to operate effectively in fast-moving, high-ambiguity environments.
  • Bachelor's degree or higher in Computer Science, Physics, or a related technical field.

Nice To Haves

  • Experience building benchmarks, evaluations, or human-data pipelines for large language models (LLMs) is strongly preferred.
  • Experience working directly with AI researchers or foundation model labs.
  • Experience developing evaluation systems for foundation models or LLM-based applications.
  • Published research, papers, or open-source contributions related to ML evaluations or benchmarks.
  • GitHub contributions involving ML evaluation, benchmark systems, LLM evaluation, or related infrastructure.
  • Experience with agentic AI evaluation environments involving tools, code execution, or multi-step workflows.
  • Experience developing human-data pipelines for machine learning evaluation.
  • Experience working across multiple AI modalities or evaluation domains.

Responsibilities

  • Partner directly with the General Manager, researchers, and early customers to define, design, and build AI benchmarks and evaluation systems.
  • Develop benchmarks and evaluation frameworks across multiple AI domains and modalities.
  • Build and own backend infrastructure supporting AI model evaluation.
  • Design and maintain data pipelines, execution environments, storage systems, and orchestration infrastructure.
  • Build sandboxed environments for agentic evaluations involving tools, code execution, and multi-step tasks.
  • Develop and deploy end-to-end evaluation systems for foundation models.
  • Own the engineering component of customer engagements from initial requirements through technical delivery.
  • Work directly with enterprise customers to understand technical requirements and develop practical evaluation solutions.
  • Identify repeatable evaluation patterns that can be transformed into scalable infrastructure and product capabilities.
  • Identify infrastructure gaps and opportunities that can inform future product development.
  • Scope ambiguous technical problems, assess feasibility, and independently drive solutions through implementation and delivery.
  • Collaborate with AI researchers and other technical stakeholders to translate research concepts into reliable production systems.
  • Develop scalable systems capable of supporting large-scale machine learning evaluation and benchmark workloads.
  • Communicate technical concepts, project status, tradeoffs, and solutions clearly to customers and internal stakeholders.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service