Applied Computer Science Benchmark Specialist

Weekday AI
$66 - $84Remote

About The Position

This role is for one of our clients. We are seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. In this role, you will develop and validate rigorous multiple-choice questions across a broad range of computer science domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems. You will contribute through one of two primary task types: Question Authoring — Develop original, challenging multiple-choice questions within your area of computer science expertise, assess their difficulty, and submit them for review. Question Verification — Review existing questions for technical accuracy, clarity, completeness, and rigor. Make necessary edits, assess difficulty, and document the rationale behind your changes.

Requirements

  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related discipline.
  • A Master's degree may be considered for candidates with exceptional expertise in a specialized computer science domain.
  • Strong command of graduate-level computer science theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.
  • Excellent written English and the ability to communicate complex technical concepts clearly, accurately, and concisely.
  • Strong attention to detail and the ability to distinguish technically valid solutions from plausible but incorrect approaches.

Nice To Haves

  • Research publications, substantial industry experience at leading technology organizations, systems engineering experience, or competitive programming experience is a strong plus.

Responsibilities

  • Create original computer science questions that evaluate deep conceptual understanding, technical reasoning, and problem-solving rather than surface-level recall.
  • Ensure every question is unambiguous, self-contained, technically accurate, and sufficiently specified for a qualified expert to solve.
  • Classify questions by difficulty: Medium: Introductory undergraduate level, Hard: Advanced undergraduate level, Expert: Postgraduate level and above
  • Provide one correct answer alongside nine plausible but subtly incorrect alternatives designed to distinguish strong technical reasoning from superficial knowledge.
  • Develop clear, structured solution explanations that demonstrate the reasoning and technical principles required to reach the correct answer.
  • Provide 1–5 authoritative references per question, drawing from peer-reviewed research, academic publications, university resources, and other reputable technical sources.
  • For verification assignments, identify issues related to correctness, clarity, completeness, precision, or solvability and clearly explain the reasoning behind any recommended edits.
  • Apply consistent standards when evaluating questions and solutions to ensure benchmark quality and reproducibility.

Benefits

  • Opportunity to contribute to the development of high-quality benchmarks for evaluating advanced AI systems
  • Strong contributors may be considered for additional review, evaluation, or subject-matter expert opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service