You’ll join the team that owns AIDR’s core reasoning engine — the multi-agent layer that turns telemetry from 70+ security vendors into investigated, correlated conclusions an analyst can act on. You’ll design the agent architectures, run the experiments that decide which ones ship, and take them into production environments where SOC teams depend on the output. Success is measured in precision, recall, latency, and cost, not demo quality. That standard shapes how we work: Accuracy and honesty. We’re clear about what we’ve measured versus what’s still a hypothesis. We expect the same of your results — including when one doesn’t hold up. Research into production. We move new research into the product quickly, then validate it against operational standards rather than demo conditions. Measure, then decide. We experiment in small increments and make calls from metrics, not intuition. In practice, you’ll spend as much time building the evaluation harnesses that tell us whether an agent is good as you will building the agent itself. If you want to work on LLM systems where “is it good?” has a number attached, this is that role.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed