Member of Technical Staff, Intern

RootlyToronto, ON
Hybrid

About The Position

Rootly AI Labs is a fellow-led research and open-source team exploring how AI can make complex systems more reliable and the people operating them more effective. We build benchmarks, prototypes, and practical tools that test frontier models against real reliability challenges, from incident response and on-call health to agent behaviour under pressure. We work in public, move quickly from ideas to experiments, and share what we learn with the broader engineering community. Our goal is to turn promising AI research into useful, trustworthy systems while keeping human judgment at the centre. The role of Member of Technical Staff, Intern will involve managing SRE Skills Bench, Rootly AI Labs' open-source benchmark for measuring how well AI models handle real-world reliability engineering tasks. This includes helping to move the project from research questions to public releases, designing realistic SRE scenarios, expanding the evaluation dataset, running model evaluations, improving the benchmark infrastructure, and sharing findings with the AI and reliability communities. The role combines applied AI research, technical project management, community building, and public communication. The intern will report directly to the Head of Rootly AI Labs and work closely with fellows, researchers, and engineers across the community.

Requirements

  • Hands-on experience or a strong interest in SRE, infrastructure, platform engineering, DevOps, or production operations.
  • Familiarity with LLM evaluations, benchmarks, agents, or applied AI research.
  • An experimental mindset, with an interest in designing reproducible tests, examining imperfect results, and learning quickly.
  • A self-driven approach, with the initiative to make sound decisions and move ambitious projects forward with minimal guidance.
  • The ability to explain technical work clearly and enthusiasm for sharing it through articles, social media, and talks at meetups or conferences.
  • Comfort working in a hybrid environment and collaborating asynchronously with a curious, distributed team.

Responsibilities

  • Participate in shaping the roadmap and research questions for SRE Skills Bench.
  • Help build an active community around the benchmark and encourage outside contributions.
  • Create realistic evaluation tasks based on incident response, infrastructure, debugging, observability, and production operations.
  • Help develop reliable and reproducible methods for evaluating frontier and open-source models.
  • Analyze results and turn them into reports, technical articles, and public datasets.
  • Collaborate with SREs, researchers, model providers, and open-source contributors.
  • Evaluate how new models, agents, and coding environments perform on operational work.
  • Help establish SRE Skills Bench as the industry standard for evaluating AI on reliability engineering.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service