About The Position

Meta Superintelligence Labs is seeking a Content Engineer to define and implement quality standards for AI product capabilities. This role involves developing guidelines, evaluation frameworks, and measurement strategies to ensure high-quality AI products. The Content Engineer will lead human evaluation pipelines, collaborate with research scientists to build automated judges, and manage structured testing programs. A key aspect of this role is operating at speed to inform the current development cycle and identify emerging quality issues before they impact external users. The position also requires crafting system prompts and agentic behavior, and serving as a cross-functional leader to align various teams on quality standards. Leveraging AI-native tools to automate workflows and drive efficiency is crucial, as is a willingness to experiment and build solutions hands-on.

Requirements

  • Bachelor's degree or equivalent experience.
  • 5+ years of experience in digital content strategy, user experience, technical writing, journalism, production, or related fields.
  • Experience designing and running human evaluation pipelines at scale (annotator management, rubric design, golden set construction, calibration).
  • Experience translating ambiguous feedback into structured, objective, fixable categories.
  • Experience defining and operationalizing subjective quality dimensions into measurable benchmarks.
  • Experience running structured software testing/QA programs (designing test plans, triaging results, delivering actionable analysis quickly).
  • Experience making editorial and content quality decisions in a fast-paced environment.
  • Experience communicating complex technical concepts to cross-functional partners.
  • Experience leading through influence across teams without direct reporting lines.
  • Experience with AI-native tooling (LLM-based development tools, annotation platforms, prototyping environments).
  • Experience with LLM-as-judge development (building automated quality signals aligned with human judgment, validating alignment over time).
  • Experience working with product teams or programs from roadmapping through delivery.
  • Experience working in prompt engineering and agentic workflows.

Nice To Haves

  • A bias toward using AI-native tools to move faster.

Responsibilities

  • Define "great" for AI product capabilities by building guidelines, golden response sets, and frontier evaluations.
  • Develop failure mode taxonomies for engineering teams.
  • Own the human evaluation pipeline, including designing rubrics, guiding annotator teams, building calibration processes, and analyzing results.
  • Partner with Research Science to build and validate auto-judges aligned with human raters and define methodology for measuring alignment drift.
  • Construct and run large-scale evaluations to track quality metrics across product capabilities.
  • Lead structured dogfooding and testing programs, including designing test plans, running testing rounds, triaging results, and delivering prioritized summaries.
  • Operate at speed to unblock fast iteration and turn around quality assessments quickly.
  • Identify emerging quality issues before they reach external users.
  • Craft and tune system prompts and agentic behavior to support product vision and model outcomes.
  • Serve as a key member of cross-functional teams, aligning engineering, research science, product, policy, and design on quality standards and priorities.
  • Lead through influence and collaboration across teams.
  • Leverage AI-native tools to replace manual workflows with scalable, repeatable processes.
  • Continuously evaluate and adopt emerging tools to operate at speed.
Β© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service