Software Engineering Manager - Build Agent

ServiceNowSanta Clara, CA

About The Position

The Build Agent team is ServiceNow's AI coding assistant, purpose-built for the platform's metadata-driven substrate, operating natively across ServiceNow's scoped applications, tables, and metadata types. This role leads the team responsible for the Build Agent evaluation framework, model support, and telemetry. The framework is critical to Build Agent success, having already driven measurable wins such as recovering Build Agent correctness, cutting latency and inference cost ahead of a major release, and catching high-severity defects that manual testing missed before they reached production. The team consists of 8 engineers and is currently a bottleneck on its own scale; this role exists to convert a high-performing initiative into a durable, scalable function. Key areas of focus include: Evaluation Infrastructure: Golden prompt sets, Pass@1 and functional scoring, failure categorization (plumbing vs. metadata-creation failures), and coverage across ServiceNow metadata types and UI workflows. Model Support & Benchmarking: Structured evaluation of candidate foundation models against the production default, with failure-consistency analysis to separate scaffold/tuning issues from genuine capability gaps. Telemetry: Token usage, cache efficiency, and inference cost tracked alongside correctness as first-class release signals. Cross-Team Arbitration: Prioritizing eval coverage, defining what "good" means across surfaces owned by different contributing teams, and turning eval signal into committed fixes by the owning teams rather than open-ended findings.

Requirements

  • Direct experience building or leading evaluation systems for LLM-based products — scoring methodology, benchmark design, and failure analysis, not just prompt engineering.
  • Experience driving cross-team accountability without direct authority — getting other teams to own and fix issues surfaced by your data.
  • A working understanding of coding agent architecture: tool use, agentic loops, context management, and the tradeoffs between generic and domain-specific scaffolding.
  • 6+ years of experience with technologies relevant to ServiceNow, including advanced coding skills and fluency in one or more of Java, C++, Ruby, Shell, or JavaScript.
  • Experience critically evaluating foundation models — distinguishing real capability differences from tuning or scaffold artifacts.
  • Ability to execute against ambiguous priorities, weighing context, risk, and desired outcomes — particularly where eval data is incomplete or contested.
  • People management experience is required.

Nice To Haves

  • experience managing at scale (8+ engineers) or managing managers is a plus.

Responsibilities

  • Own eval strategy and roadmap: prompt set design, scoring methodology, failure-mode taxonomy, and telemetry instrumentation.
  • Coordinate across the multiple teams contributing to Build Agent to arbitrate quality standards and drive eval findings to committed, owned fixes — not just reports.
  • Lead model benchmarking efforts, making data-backed recommendations on model support decisions (e.g., evaluating frontier coding models against the incumbent default).
  • Manage the daily activities of an 8-engineer team, including staffing, mentoring, and performance management, while building the scale (people, process, or automation) to remove the team as its own bottleneck.
  • Solve ambiguous, cross-cutting problems where eval signal, model capability, and platform architecture intersect — and where "correct" isn't obvious until you've defined the bar.
  • Represent eval results and model support recommendations to cross-functional and executive stakeholders, including in response to org-wide benchmark initiatives.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service