About The Position

NVIDIA is seeking a Senior Software Engineer to join the Open Harness Engineering team. This role focuses on building core libraries, evaluation systems, and reusable harness capabilities to enhance the safety, speed, and trustworthiness of AI agents. The engineer will work with open-source agent harnesses, sandboxed execution, model-provider infrastructure, and agent benchmarks to improve next-generation agents. Responsibilities include diagnosing failures, contributing to upstream projects, and demonstrating quality, reliability, cost, and latency improvements across models. This is an opportunity to build foundational technology for autonomous software systems.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Applied Math, or a related field, or equivalent experience.
  • 8+ years of hands-on software-engineering experience, with demonstrated technical ownership of production systems; and architecture design leadership.
  • Expert Python skills and working proficiency in Rust, Go, C++, or TypeScript, with the ability to debug systems across language and process boundaries.
  • Experience building or extending LLM agents, coding agents, tool-use loops, model-provider integrations, developer tools, or evaluation infrastructure.
  • Solid understanding of asynchronous execution, subprocesses, containers, sandboxing, callbacks, retries, networking, and distributed compute.
  • Experience crafting reproducible experiments, benchmark methodology, performance investigations, regression gates, or reliability analysis under nondeterministic conditions.
  • A record of shipping and maintaining open-source or developer-facing software with sound testing, API development, documentation, and code reviews.

Nice To Haves

  • Meaningful contributions to open-source coding agents, harnesses, agent runtimes, evaluation frameworks, developer tools, or observability projects.
  • Published research, technical writing, patents, or substantive open-source work in autonomous software engineering, agent evaluation, reliability, tool use, or inference efficiency.
  • Experience with SWE-bench, Terminal-Bench, AgentBench, or comparable public agent benchmarks.
  • Shipped improvements to context management, tool interfaces, retries, sandboxed repository execution, long-running agent loops, or evaluation environments.
  • Experience with trace-analysis and observability systems such as OpenTelemetry, OpenInference, structured event pipelines, exporters, or debugging tools.

Responsibilities

  • Define and evolve the technical approach for evaluating and improving open agent harnesses across frameworks, models, and benchmark environments.
  • Build and operate scalable, reproducible evaluation systems spanning CI, scheduled compute, sandboxed execution, artifact provenance, trace capture, and analysis.
  • Diagnose agent failures from traces, logs, tool results, and benchmark artifacts, separating harness, evaluator, environment, inference, and model causes.
  • Design controlled experiments and promotion criteria that distinguish genuine harness improvements from noise, benchmark artifacts, or model-specific gains.
  • Ship focused harness and runtime improvements, including upstream open-source contributions, detectors, regression tests, and maintainable documentation.
  • Partner with Relay, Hermes, model, infrastructure, and research teams to turn recurring optimization needs into reusable runtime, trace, and evaluation capabilities.

Benefits

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity
  • Benefits (details at www.nvidiabenefits.com/)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service