About The Position

Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group, you will lead the benchmarking and integration of state-of-the-art models for image and video understanding. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously test models in applied settings, uncover edge-case failure modes, and architect advanced agentic systems. If you are passionate about AI safety, robust evaluation, and building autonomous multi-modal workflows that bridge experimentation with production, we’d love to hear from you.

Requirements

  • MS and a minimum of 3 years relevant industry experience
  • 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation
  • Solid ML Foundation: Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance.
  • Computer Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs).
  • Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models.
  • Deep experience with custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.
  • Agentic Systems: Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.
  • Failure Analysis: Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments.
  • Engineering Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference, extracting embeddings, and building scalable evaluation pipelines.

Nice To Haves

  • Demonstrated ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers.
  • Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design
  • Knowledge of reinforcement learning, planning, or decision-making systems
  • Experience evaluating multi-modal or multi-agent systems
  • Prior work on AI reliability, safety, or benchmarking

Responsibilities

  • Lead the benchmarking and integration of state-of-the-art models for image and video understanding.
  • Apply deep CV and ML expertise to rigorously test models in applied settings.
  • Uncover edge-case failure modes.
  • Architect advanced agentic systems.
  • Build and evaluate LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.
  • Probe ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments.
  • Translate findings into actionable improvement recommendations.
  • Build scalable evaluation pipelines.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service