About The Position

We are sharing a specialised consulting opportunity for AI safety and red teaming professionals with native-level fluency in both English and Swedish and experience evaluating conversational AI systems through adversarial testing. This role supports an AI safety initiative focused on identifying vulnerabilities, failure modes, and systemic risks in conversational models and agents. Selected professionals will conduct structured red teaming, generate high-quality evaluation data, classify model failures, and document reproducible attack scenarios that can be used to improve model robustness and safety.

Requirements

  • Native-level fluency in both English and Swedish
  • Prior experience with AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing
  • Strong understanding of conversational AI systems and model failure modes
  • Experience developing structured adversarial tests and evaluation frameworks
  • Ability to identify subtle vulnerabilities and recurring behavioural patterns
  • Strong analytical reasoning and written communication skills
  • Ability to document findings clearly and reproducibly
  • Comfort working across changing projects, scenarios, and evaluation frameworks

Nice To Haves

  • Experience creating jailbreak or prompt-injection datasets
  • Familiarity with adversarial machine learning
  • Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts
  • Cybersecurity experience involving penetration testing, exploit development, or reverse engineering
  • Background analysing abuse, harassment, misinformation, or other socio-technical risks
  • Experience testing conversational AI systems
  • Strong creative writing, psychology, or behavioural-analysis skills applicable to adversarial testing
  • Previous experience producing structured human data for AI evaluation

Responsibilities

  • Conduct adversarial testing of conversational AI models and agents
  • Develop jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation strategies
  • Probe models for vulnerabilities that may be missed by automated evaluation systems
  • Test model behaviour across diverse conversational and adversarial scenarios
  • Apply systematic testing frameworks and established evaluation methodologies
  • Identify and classify model failures and safety vulnerabilities
  • Evaluate issues involving bias, misinformation, misuse, and potentially harmful model behaviours
  • Identify recurring or systemic patterns across model responses
  • Assess the severity, reproducibility, and practical significance of discovered vulnerabilities
  • Apply established taxonomies, benchmarks, and testing playbooks consistently
  • Produce high-quality human evaluation data from red teaming activities
  • Annotate model failures and categorise identified vulnerabilities
  • Create structured attack cases and supporting evaluation materials
  • Maintain consistency across repeated assessments and datasets
  • Provide clear rationale for classifications and safety judgments
  • Document adversarial scenarios in a clear and reproducible format
  • Produce reports, datasets, and structured findings that technical teams can act upon
  • Explain identified risks to both technical and non-technical stakeholders
  • Record testing methodology, model behaviour, and relevant failure patterns
  • Contribute to broader evaluation coverage across models and use cases

Benefits

  • Flexible remote consulting work
  • Competitive hourly compensation
  • Flexible scheduling
  • Weekly payments
  • Projects may be extended, shortened, or adjusted depending on scope and performance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service