Staff AI Test and Evaluation Engineer

Anduril•Columbia, WA

About The Position

Discovery is Anduril's team for taking the newest problems across domains- space, missile systems, air, sensor capability, autonomy, and cyber- and proving out what is worth solving. We build the models, run the tests, and carry what works to the point where a program can pick it up. Discovery works alongside Air Defense, Space, Intelligence, Cyber, GNC, Hardware, and every other group at Anduril to incubate the solutions to the hardest problems. Discovery is building an expeditionary force: a team of engineers who want to solve difficult problems and who navigate unfamiliar territory as an operating standard. As a Discovery engineer, you pick up a concept nobody has proven and build the analysis or prototype that tests it to get to an answer. Whether you spend your time working on hard problems on our autonomy stack, acoustic sensor analysis, or battle space management, we tackle every challenge the same way: by questioning assumptions and building our way to the answer. We are focused on taking AI models from research into production—onto classified platforms and edge hardware where they have to perform reliably in the real world. As our model portfolio and classified work grow, rigorous, repeatable evaluation of how these models actually perform has become mission-critical.

Requirements

  • Bachelor's or Master's degree in Computer Science, Machine Learning, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field.
  • At least 12+ years of hands-on experience writing production-grade code, with strong Python skills for building evaluation pipelines, test harnesses, and tooling.
  • Demonstrated experience owning evaluation or automated testing for production ML or software systems—test scenarios, metrics, regression suites, and CI infrastructure.
  • Hands-on experience designing evaluation methodologies for AI or machine learning models—defining metrics, building benchmarks, and assessing model behavior against real-world scenarios.
  • Experience designing or working extensively with simulation environments to exercise model and system behavior.
  • Working knowledge of how to define, capture, and reason about performance metrics for AI models and traditional models, including systematic historic capture over time.
  • Ability to navigate and contribute to complex systems and established codebases.
  • Comfort operating between technical program management and software engineering—defining requirements, coordinating across teams, and documenting results.
  • Passion for building the evaluation infrastructure that proves AI works—directly influencing mission-critical outcomes.
  • Must be eligible for a US security clearance.

Nice To Haves

  • Experience as an ML Test & Evaluation Engineer, SDET for ML systems, ML/Evaluation Engineer, Simulation Engineer, or Technical Program Manager for AI/ML.
  • Experience designing test methodologies for agentic AI systems—tasking, decision-making, and scenario-based behavior validation.
  • Familiarity with validating and deploying AI models onto classified platforms, edge hardware, or resource-constrained environments.
  • Experience contributing to AI model cards, deployment readiness reviews, or similar model-governance and release-gating practices.
  • Familiarity with test and evaluation practices in aerospace or defense, including qualification testing, range operations, or operational assessment.
  • Experience with MLOps tooling, monitoring dashboards, and continuous validation pipelines for ML models in production.
  • Additional experience with Go, C++, or scripting for test automation and tooling.
  • Interest in growing from T&E ownership into deeper ownership of the AI evaluation and deployment stack over time.

Responsibilities

  • Develop comprehensive test scenarios, and simulation environments to assess agentic AI performance in classified simulations.
  • Lead the integration and validation of agentic AI systems onto classified platforms and edge hardware.
  • Build reusable evaluation pipelines, automated test harnesses, and monitoring dashboards for continuous validation.
  • Figure out what good metrics look like for AI models and traditional models—establishing systematic, historic capture of performance rather than one-shot, deployment-specific measurement.
  • Partner with cross-functional teams to define requirements, document test results, and contribute to AI model cards and deployment readiness reviews.
  • Architect, build, and maintain the evaluation and validation infrastructure for our agentic AI and ML models, from unit-level model checks through full-system, scenario-based assessment.
  • Analyze and resolve issues uncovered in evaluation and in deployment, ensuring reliability and operational success across every release.

Benefits

  • Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package.
  • Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service