Test & Evaluation Engineer

Primordial Labs•New Haven, CT
•$100,000 - $160,000•Remote

About The Position

Anura takes a spoken order and turns it into robot motion. Verifying that is not like verifying normal software. The input is a human voice, the system is probabilistic, and a failure looks like a vehicle doing something the operator didn't ask for. A green unit test suite tells you almost nothing about whether it works. We test against real airframes, simulation across ground, air, and multi-domain, and a harness that drives Anura and scores against test cases. This role takes that further: deeper coverage earlier in the cycle, harder evaluation sets, and the infrastructure to prove out a build before it ever reaches a range. You'll build and run the machinery that tells us the truth about our product. When we tell a program office the system performs, you own the evidence behind that claim.

Requirements

  • 3+ years in test engineering, test automation, systems test, ML evaluation, or a comparable role. Level and scope are flexible for the right candidate.
  • Strong proficiency in Python, comfort in Linux, and hands-on experience with CI.
  • Demonstrated depth in at least one of test automation, hardware and bench test, ML evaluation, or platform test.
  • Willingness to travel to test ranges and customer sites. This is not a desk-only role.
  • BS in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience.

Nice To Haves

  • Robotics, unmanned systems, or autonomy test experience
  • Defense test and evaluation experience, including test plans and government deliverable reporting
  • Experience testing systems with ML components
  • Part 107 or experience as a safety operator on small UAS
  • Ability to obtain a US Security Clearance

Responsibilities

  • Push Coverage Earlier. Build the sim, bench, and CI capability that qualifies a build before it reaches the field. Range time is the most expensive test environment we have, and it should be confirming what you already know, not discovering it.
  • Own the Rigs. Extend our test infrastructure and stand up software- and hardware-in-the-loop capability against representative hardware. This is the layer everything else runs on.
  • Score the System. Own evaluation sets and adversarial suites. Measure accuracy under degraded conditions and whether an operator's command actually produced the intended behavior end to end, not just whether a component returned something plausible.
  • Test on Real Platforms. Execute test cards on live hardware, including safety-relevant behavior. You'll be at the range, and you'll be the one who decides whether what you saw counts as a pass.
  • Close the Loop. Turn a day of range data into replayable regression cases and evaluation additions within days, not months. This is the difference between testing and learning.
  • Call It Ready. Produce the evidence that says a given build and configuration is flight-ready, and be willing to say "not yet" with data behind it.
  • You build the tools, run the tests, and tell us precisely what broke and under what conditions. The engineering teams own the fixes, including in the NLP and autonomy stack. You're the scorekeeper, not the cleanup crew.

Benefits

  • Significant equity stake
  • 100% Remote: This role is fully remote, open to residents across the US.
  • Comprehensive medical, dental, and vision plans
  • 401(k) package, complete with a generous 6% company match.
  • Generous Discretionary Time off (DTO) policy
  • Elite technology package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service