About The Position

Cohere is seeking Generalist professionals with broad backgrounds for a part-time independent contractor position in Canada. This role is crucial for improving Large Language Models (LLMs) by labeling, ranking, auditing, and correcting model outputs. You will work with English-language data across text, image, and structured formats, applying strong analytical and judgment skills to evaluate, stress-test, and enhance model performance. This is a judgment-driven role focused on improving AI systems.

Requirements

  • 1+ years of experience in AI data annotation, LLM evaluation, content moderation, research, or a related analytical role.
  • Experience with quality assurance and/or preference ranking.
  • Experience applying detailed guidelines to complex and often ambiguous content.
  • Strong contextual and sociocultural judgment, sensitivity to nuance, tone, and register.
  • Ability to reason well in cases where there is no single correct answer.
  • Comfort with ambiguity and willingness to flag unclear edge cases.
  • Ability to work productively on novel, experimental tasks.
  • A sharp eye for inconsistencies, subtle errors, and model failure modes.
  • Excellent command of written English and strong reading comprehension.
  • Ability to clearly justify evaluations and write clear prompts and exemplars.
  • Strong attention to detail and commitment to accuracy.
  • Ability to maintain consistency across high-volume and monotonous tasks.
  • Comfort working with annotation platforms and structured formats such as JSON, CSV/TSV, Markdown, XML, and YAML.
  • Strong execution in a remote environment, including good time management, comfort using new tools, and ability to work independently.
  • Must be located within Canada.
  • Ability to commit to 16 hours per week.
  • Must have a laptop (BYOD).

Nice To Haves

  • Familiarity with how large language models behave, such as hallucination, sycophancy, and instruction-following gaps.
  • Fluency in another language.

Responsibilities

  • Evaluate and rank model outputs, assessing responses based on accuracy, helpfulness, tone, and safety, and providing clear justifications.
  • Stress-test models by probing for failure modes, unsafe behavior, and capability gaps, documenting reproducible cases.
  • Create datasets by authoring prompts, responses, and exemplars, and editing existing outputs.
  • Build and apply rubrics and taxonomies for grading criteria.
  • Annotate, audit, and correct inaccuracies across text, image, and structured data.
  • Participate in calibration exercises and inter-annotator agreement checks to maintain consistency.
  • Adapt to new and evolving task types, applying sound judgment to experimental work.
  • Report on model performance, providing feedback on where models succeed, fail, and degrade.

Benefits

  • Independent contractor agreement
  • Opportunity to work with cutting-edge AI models
  • Flexible scheduling (16 hours per week commitment)
  • Remote work environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service