About The Position

Nuro is seeking a Senior/Staff Software Engineer to join their AI Agent Infrastructure team. This role focuses on building and maintaining the platform that enables AI agents to operate autonomously within Nuro's engineering organization, adhering to rigorous standards of proof and measurement. The team's mandate is to significantly amplify the output of every engineer and researcher by developing a trustworthy and highly capable autonomous system. This involves creating a robust closed-loop evaluation system, a reliable agent platform for orchestration and safety, and autoresearch infrastructure to automate the research loop. The engineer will work on agent-powered tooling across the engineering lifecycle, including code generation, review, debugging, and more. The role offers direct access to compute, systems automation, and leadership, with significant decision-making authority.

Requirements

  • 5+ years of software engineering experience (or 4+ with a Master's) in computer science, engineering, or equivalent practical experience. Staff-level candidates should demonstrate deeper scope and ownership.
  • Deep, current understanding of LLM research, including model training from scratch (data, tokenization, architecture, pretraining dynamics) and the post-training stack (SFT, preference optimization, RL).
  • Ability to reason about how training decisions impact model behavior and follow relevant literature.
  • Understanding of inference internals, including attention, KV-cache, batching, quantization, speculative decoding, prefix caching, and context handling, and their trade-offs.
  • Experience building and operating LLM-based agent systems in production (tool use, orchestration, sandboxing, retrieval, memory) and understanding their failure modes.
  • Strong backend and distributed systems background at scale, including cloud infrastructure, service design, storage, and queuing.
  • Strong programming skills in Python.
  • Strong opinions and experience with evaluation, including the ability to critically assess benchmarks.
  • Ability to work end-to-end without pre-decomposition of problems.
  • Hands-on post-training or fine-tuning experience (SFT, preference optimization, RL, distillation) with associated evaluation work.
  • Experience with ML training or research infrastructure (experiment orchestration, evaluation pipelines, hyperparameter search, data pipelines).
  • Experience running inference serving, cost, or capacity at meaningful scale.
  • Familiarity with agent architecture patterns (planning, reflection, long-horizon memory, multi-agent coordination).
  • Experience with open tool-integration protocols, plugin or skill frameworks, and model-routing or gateway layers.
  • Background in developer experience or platform engineering, observability, or security isolation.

Nice To Haves

  • Go, C++, or Rust experience in addition to Python.

Responsibilities

  • Build the closed-loop measurement layer to assess agent output acceptance, reversion, and override rates per workflow, and use this data to guide autonomy expansion.
  • Transition the autoresearch loop from assisted to unattended operation for a defined set of experiments, including the necessary evaluation and confidence machinery.
  • Design the isolation and permissioning model for agents to interact with production repositories and infrastructure, ensuring an auditable record of actions and justifications.
  • Develop and operate LLM-based agent systems in production, including tool use, orchestration, sandboxing, retrieval, and memory management.
  • Build and maintain backend and distributed systems for cloud infrastructure, service design, storage, and queuing.
  • Contribute to the automation of the research loop, including hypothesis formation, experiment execution, honest evaluation, and proposal generation.
  • Develop agent-powered tooling for the engineering lifecycle, such as code generation, review, debugging, test and CI failure attribution, knowledge retrieval, and triage.
  • Potentially work on post-training internal models if justified by workload.
  • Ensure the agent platform provides safe runtime operation through orchestration, sandboxing, tool/skill frameworks, memory, identity, permissioning, and gateway/observability layers.
  • Implement rigorous evaluation systems for AI work, focusing on measurable outcomes and statistical honesty.

Benefits

  • Annual performance bonus
  • Equity
  • Competitive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service