The Build Agent team is ServiceNow's AI coding assistant, purpose-built for the platform's metadata-driven substrate, operating natively across ServiceNow's scoped applications, tables, and metadata types. This role leads the team responsible for the Build Agent evaluation framework, model support, and telemetry. The framework is critical to Build Agent success, having already driven measurable wins such as recovering Build Agent correctness, cutting latency and inference cost ahead of a major release, and catching high-severity defects that manual testing missed before they reached production. The team consists of 8 engineers and is currently a bottleneck on its own scale; this role exists to convert a high-performing initiative into a durable, scalable function. Key areas of focus include: Evaluation Infrastructure: Golden prompt sets, Pass@1 and functional scoring, failure categorization (plumbing vs. metadata-creation failures), and coverage across ServiceNow metadata types and UI workflows. Model Support & Benchmarking: Structured evaluation of candidate foundation models against the production default, with failure-consistency analysis to separate scaffold/tuning issues from genuine capability gaps. Telemetry: Token usage, cache efficiency, and inference cost tracked alongside correctness as first-class release signals. Cross-Team Arbitration: Prioritizing eval coverage, defining what "good" means across surfaces owned by different contributing teams, and turning eval signal into committed fixes by the owning teams rather than open-ended findings.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager
Education Level
No Education Listed