Newton Research is a fast-growing software start-up founded by repeat entrepreneurs and well-funded by blue chip venture capital firms. We are building the next generation of the closed loop media lifecycle, developing AI agents that leverage the latest in LLMs and generative AI with specialized knowledge. Our products generate actionable business insights for our customers and partners, assisting in each step of the media planning, buying and measurement lifecycle. Newton ships on a sprint cadence through a develop, stage and customer-environment pipeline, and the product surface is wide: conversations, blueprints, connectors, scheduled tasks, permissions and sharing, SSO, and AI agents whose behavior is not fully deterministic. A missed regression lands in front of a media planner or a customer's security review. We run everything through an AI-first lens, because it is the only way quality scales. If a quality task is repeatable, an agent does it and you supervise; if it takes judgment, that is where you spend your time. The gap this hire fills: evals and skill-change testing. Code changes already have CI and review, including PRs written by Claude. What has no safety net is behavior change: an edit to a skill, a prompt, a tool definition or a model version can silently change what our agents do, and nothing tests that today. You own that layer: curated eval sets, scoring and regression tracking, so any change to how an agent behaves is measured before it ships. Measure the AI: agent output varies run to run, so “correct” is a range you define with evals, rubrics and scoring, not an exact-match assertion. Let agents run the checks: agents run suites, triage failures and draft repro-ready defects; you design their roles and guardrails and review what they produce. Make Newton verifiable by agents: you keep the product and pipeline observable and fixture-friendly so agents can verify it without a human in the loop. You are also a release-readiness partner (are we good?), alongside our existing QA lead, and that judgment stays human. But you are not hired to test Claude-driven PRs line by line.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed