Collinear AI builds environments, tasks, and evaluations that help frontier AI models improve at real work. We are growing our cybersecurity team and looking for someone who can turn practical security problems into environments where AI agents can investigate, act, and learn. We recently released CWE-bench, a defensive cybersecurity benchmark with 100 held-out audit-and-patch tasks across 54 weakness types. Agents must find and fix vulnerabilities in real codebases, with checks that confirm the vulnerability is resolved and existing functionality still works. You will help build what comes next: richer cybersecurity environments, realistic tasks, and reliable ways to measure whether agents succeed. You will own work from the initial security scenario through the runnable environment, task instructions, reference solution, and verifier. We are looking for hands-on security knowledge, strong programming skills, and curiosity about how AI agents fail. Deep cybersecurity expertise is enough to get started—you do not need prior AI or machine-learning experience. If you know how to investigate vulnerabilities, reason about security failures, and verify that a fix works, we want to hear from you. We will teach you our AI tooling and evaluation workflows.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed