Senior/Staff Software Engineer, AI Agent Infrastructure

NuroMountain View, CA
$193,930 - $352,290

About The Position

Nuro is seeking a Senior/Staff Software Engineer to join their AI Agent Infrastructure team. This role focuses on building and operating the platform that enables AI agents to function autonomously within Nuro's engineering organization, adhering to rigorous standards of proof and measurement. The team's mandate is to significantly amplify the output of engineers and researchers by creating a trustworthy and autonomous system. Key areas of focus include closed-loop evaluation, agent platform development (orchestration, sandboxing, tool integration, memory, identity, observability), and autoresearch infrastructure for automating the research loop. The role involves hands-on work with production systems, direct access to compute and leadership, and making critical decisions about the platform's direction. The engineer will be central to building a closed-loop measurement layer, taking autoresearch loops from assisted to unattended, and designing isolation and permissioning models for agents operating on production systems. The position also involves building agent-powered tooling across the engineering lifecycle and potentially contributing to post-training internal models.

Requirements

  • 5+ years of software engineering experience (or 4+ with a Master's) in computer science, engineering, or equivalent practical experience. Staff-level candidates should bring correspondingly deeper scope and ownership.
  • Deep, current understanding of LLM research, including model training from scratch (data, tokenization, architecture, pretraining dynamics, post-training stack: SFT, preference optimization, RL) and reasoning about how training decisions affect model behavior.
  • Understanding of inference internals: attention and KV-cache behavior, batching and scheduling, quantization, speculative decoding, prefix caching, context handling, and their trade-offs.
  • Experience building and operating LLM-based agent systems in production (tool use, orchestration, sandboxing, retrieval, memory) and understanding their failure modes.
  • Strong backend and distributed systems background at scale (cloud infrastructure, service design, storage, queuing).
  • Strong programming skills in Python.
  • Opinionated about evaluation and able to critically assess benchmark validity.
  • Ability to work end-to-end without pre-decomposition of problems.
  • Hands-on post-training or fine-tuning experience (SFT, preference optimization, RL, distillation), including evaluation.
  • Experience with ML training or research infrastructure (experiment orchestration, evaluation pipelines, hyperparameter search, data pipelines).
  • Experience running inference serving, cost, or capacity at meaningful scale.
  • Familiarity with agent architecture patterns (planning, reflection, long-horizon memory, multi-agent coordination).
  • Experience with open tool-integration protocols, plugin or skill frameworks, and model-routing or gateway layers.
  • Background in developer experience or platform engineering, observability, or security isolation.

Nice To Haves

  • Go, C++, or Rust experience in addition to Python.

Responsibilities

  • Build the closed-loop measurement layer to track agent output acceptance, reversion, and override rates per workflow.
  • Use measurement data to guide the expansion or retraction of agent autonomy.
  • Transition autoresearch loops from assisted to unattended for specific experiment classes.
  • Develop the necessary evaluation and confidence machinery for unattended autoresearch.
  • Design the isolation and permissioning model for agents interacting with production repositories and infrastructure.
  • Ensure an auditable record of agent actions and their justifications.
  • Build and operate the agent platform, including orchestration, sandboxing, tool/skill frameworks, memory, identity, permissioning, and observability.
  • Develop agent-powered tooling for the engineering lifecycle (code generation, review, debugging, test/CI failure attribution, knowledge retrieval, triage).
  • Potentially contribute to post-training internal models where justified by internal workload.

Benefits

  • Annual performance bonus
  • Equity
  • Competitive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service