Research Engineer (RE) Curator

DataForce by TransPerfectRemote, United States, AMERICAS
Remote

About The Position

DataForce by TransPerfect seeks a STEM Research Engineer to build next-generation AI evaluation benchmarks, perform red teaming, and advance LLM experimentation. DataForce by TransPerfect is seeking exceptional Research Engineers and AI Specialists for the role of Research Engineer (RE) Curator to drive the creation, execution, and refinement of next-generation AI evaluation benchmarks. In this high-priority role, you will leverage your deep analytical mindset, Python proficiency, and experimental research background to curate, evaluate, red team, and benchmark cutting-edge Large Language Models (LLMs) and generative AI systems, ensuring top-tier accuracy, safety, and reasoning capabilities.

Requirements

  • MSc or PhD in a STEM discipline (Computer Science, AI/ML, Statistics, Mathematics, Physics, Computational Sciences, or related field).
  • Proven background as a Research Engineer, Applied Scientist, ML Engineer, Research Scientist, or AI Researcher.
  • Strong hands-on Python programming skills, with daily fluency in Git, modern IDEs, and Jupyter/Colab environments.
  • Core experience in Machine Learning, LLMs, data analysis, and structured experimental research.
  • Strong research mindset paired with exceptional analytical and problem-solving capabilities.
  • Successful completion of a role-specific coding/research assessment and background check.

Nice To Haves

  • Direct experience in AI evaluation, benchmark curation/development, red teaming, or model testing frameworks.

Responsibilities

  • Design, curate, and implement complex AI evaluation benchmarks and dataset frameworks to test advanced LLM capabilities and limitations.
  • Conduct experimental research, red teaming, and systematic testing to identify model failure modes, edge cases, and reasoning gaps.
  • Utilize Python, Jupyter/Colab environments, and modern IDEs to write scalable code, process datasets, and automate evaluation pipelines.
  • Collaborate with AI research and engineering teams to translate complex subject-matter concepts into actionable model benchmarks.
  • Document research findings, track codebase and dataset changes via Git, and deliver precise sourcing and evaluation reports.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service