Applied Machine Learning Scientist

Vector InstituteToronto, ON
CA$125,800 - CA$157,300

About The Position

As an Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering, you will lead applied research on evaluation, observability, stress-testing, and systematic improvement of AI agents. The role focuses on assessing agent performance and safety across long-horizon, multi-step tasks, and on building methods and tools to help organizations understand whether those systems are working, why they fail, and how to make them measurably better. A core objective is developing adaptive evaluation approaches tailored to Canadian organizations, moving beyond static public benchmarks towards rigorous, organization-specific test environments of end-to-end agentic systems. Working alongside Vector researchers, research professionals, and external partners, the role balances high-quality applied research with the creation of practical technical systems that improve the reliability, safety, security, and effectiveness of deployed agents.

Requirements

  • PhD in computer science, computer engineering, machine learning, or a related discipline, or equivalent demonstrated research or engineering experience;
  • Research expertise in one or more of: evaluation of AI agents or language-model systems; automated red-teaming; program synthesis or automated software improvement; AI safety, security, or robustness; multi-agent systems;
  • Strong ability to design controlled experiments and reason about confounding variables, stochasticity, statistical power, evaluator reliability, and reproducibility;
  • Experience evaluating systems whose behaviour unfolds across multiple steps, tool interactions, or environmental state changes;
  • Strong knowledge of Python and experience building high-quality research software;
  • Experience working with modern language models and tool-using agent architectures;
  • Understanding of the distinction between model evaluation and evaluation of the broader model–harness–environment system;
  • Familiarity with open-source machine-learning and agent frameworks such as PyTorch, JAX, Google ADK, LangGraph, the OpenAI Agents SDK, or comparable systems;
  • Comfortable working at the boundary between open-ended research and production-quality engineering;
  • Able to communicate complex findings clearly to technical researchers, engineering leaders, domain specialists, and senior organizational stakeholders.

Responsibilities

  • Research and implement state-of-the-art methods for evaluating agents operating over long horizons, multiple tools, changing environments, and partially observable states;
  • Develop evaluations that assess complete agent trajectories, including planning quality, tool selection, intermediate decisions, state transitions, recovery behaviour, verification, termination decisions, resource consumption, and final outcomes;
  • Develop methods for creating organization-specific evaluations from production traces, human feedback, incidents, near misses, support interactions, domain-expert knowledge, and synthetic scenario generation;
  • Create techniques for converting discovered failures into durable regression evaluations that can be rerun across model, prompt, policy, tool, and harness changes;
  • Partner with Vector researchers, Applied ML Specialists, research professionals, and external collaborators – including Vector industry partners and members of the Canadian AI Safety Institute – to identify consequential agent use cases and create tools, reference agents, and evaluations required for trustworthy deployment;
  • Develop schemas and infrastructure for capturing structured traces of active agents;
  • Research representations of agent trajectories, such as event streams, causal graphs, tool-call graphs, state-transition graphs, and compact trajectory embeddings;
  • Develop approaches for identifying recurrent failure patterns and attributing outcomes to specific components or decisions within an agent system;
  • Build privacy-preserving and security-conscious methods for collecting and analyzing traces in sensitive organizational environments;
  • Research and build agent harnesses incorporating tools, memory, retrieval, sandboxes, permissions, validators, execution loops, recovery strategies, state management, and human approval mechanisms;
  • Develop automated or semi-automated methods for optimizing agent harnesses based on evaluation results and execution traces;
  • Develop safe mechanisms for agents to propose modifications to their own prompts, tools, policies, memory structures, workflow logic, or evaluation criteria while preserving auditability and human control;
  • Lead or contribute to peer-reviewed publications, technical reports, open-source software, benchmark releases, and reference implementations;
  • Contribute to training programs and technical workshops that help Vector partners and external stakeholders design, evaluate, debug, and govern agent systems;
  • Serve as a Vector expert on emerging methods in agent evaluation and harness engineering and connect external stakeholders with relevant members of the Vector research community; and,
  • Other related duties as assigned from time to time.

Benefits

  • vacation time
  • floater days
  • GRRSP
  • a Health Spending Account
  • a Summer Hours program
  • flexible work arrangements
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service