We are building the evaluation backbone for safe, reliable, and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows. The role is ideal for engineers who can operate end to end: understand the workflow, design the benchmark, build the harness, implement validators and graders, run experiments, analyze quality and cost tradeoffs, connect results to production feedback, and help teams make evidence-based release decisions. Microsoft Security is moving toward agentic engineering systems for security triage, remediation, repo readiness, and scan-to-verified-closure workflows. Evals are the trust system for that shift. They help decide whether autonomy can safely expand, whether a release should stop, and which configuration achieves the required quality, safety, reliability, latency, and cost bar with the lowest practical human-review burden. As a Principal Security Research Manager on the AI Evaluation Systems team, you will build the common evaluation platform and methodology used by MSec agent programs. You will work across evaluation design, platform implementation, test infrastructure, telemetry, measurement, security workflow understanding, and production learning. Your work will make agentic systems measurable, reproducible, governable, and continuously improving.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal
Education Level
Ph.D. or professional degree