About The Position

We are sharing a specialised part-time consulting opportunity for experienced AI safety, red-teaming, trust and safety, cybersecurity, investigative research, and scientific-risk professionals with expertise in adversarial evaluation of advanced AI systems. This role supports a frontier AI safety initiative focused on identifying vulnerabilities, unsafe behaviours, policy failures, and robustness gaps through structured adversarial testing. Selected professionals will design challenging prompts, evaluate model behaviour across high-risk and ambiguous scenarios, document findings, and contribute to safety benchmarks used to improve model alignment and reliability.

Requirements

  • At least 5 years of professional experience in AI safety, AI red teaming, trust and safety, cybersecurity, investigative journalism, life sciences, public policy, or a related field
  • Hands-on experience designing adversarial prompts or evaluating advanced AI systems
  • Strong analytical reasoning and the ability to identify subtle safety and policy failures
  • Experience working with complex, high-risk, or ambiguous subject matter
  • Excellent written communication and the ability to produce clear technical findings
  • Ability to work independently while applying structured evaluation standards
  • Professional residence in one of the eligible countries listed below
  • A bachelor's degree or higher in computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy, or a related discipline is required

Nice To Haves

  • Experience with AI red teaming, reinforcement learning from human feedback, supervised fine-tuning, AI alignment, or trust and safety
  • Familiarity with jailbreak testing, prompt engineering, model-behaviour analysis, or adversarial evaluation methodologies
  • Expertise in cybersecurity, biosecurity, political content, misinformation, fraud, or scientific safety
  • Experience developing safety benchmarks, evaluation rubrics, or structured testing frameworks
  • Knowledge of responsible disclosure, threat modelling, or vulnerability-severity assessment
  • Previous collaboration with AI researchers, policy teams, security engineers, or scientific specialists
  • Familiarity with frontier-model safety policies and model-alignment workflows
  • Graduate-level education in artificial intelligence, security, behavioural science, life sciences, or policy may be valuable
  • Equivalent specialist experience in adversarial testing, safety research, or high-risk investigations may also be considered
  • Research publications, safety evaluations, or relevant technical portfolios may strengthen an application

Responsibilities

  • Design sophisticated prompts that stress-test advanced AI systems
  • Develop realistic scenarios intended to expose model limitations, policy weaknesses, and inconsistent behaviour
  • Test direct, indirect, multi-turn, and context-dependent adversarial strategies
  • Create evaluations covering both clearly unsafe requests and complex grey-area situations
  • Identify jailbreak vulnerabilities, hallucinations, unsafe outputs, and instruction-following failures
  • Assess model robustness across misinformation, cybersecurity, biosecurity, fraud, political content, and scientific safety
  • Evaluate whether responses appropriately balance safety, usefulness, factual accuracy, and policy adherence
  • Distinguish isolated failures from broader or reproducible behavioural patterns
  • Document identified vulnerabilities through clear, structured, and reproducible reports
  • Explain testing methods, observed behaviours, severity, and potential impact
  • Classify failures according to established safety taxonomies and evaluation frameworks
  • Contribute findings to red-teaming reports, benchmark datasets, and model-improvement workflows
  • Collaborate with AI researchers, safety specialists, and other domain experts
  • Participate in calibration exercises to maintain consistent evaluation standards
  • Review adversarial tasks and findings developed by other contributors
  • Support improvements to testing methodologies, safety rubrics, and vulnerability classifications

Benefits

  • Flexible remote work
  • Competitive hourly compensation
  • Independent contractor role
  • Fully remote with flexible scheduling
  • Competitive rates between $65–$80 per hour depending on expertise and project scope
  • Weekly payments via Stripe or Wise
  • Projects may be extended, shortened, or adjusted depending on scope and performance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service