R&D Spring Co-op

Johnson & Johnson Innovative MedicineHopewell Township, NJ
$24 - $53Onsite

About The Position

At J&J we are developing Generative AI systems to support scientific work across discovery and translational research, including reports, hypotheses, summaries, and analyses. Many qualities that matter in these outputs — such as scientific plausibility, reasoning quality, insightfulness, novelty, and usefulness for future research — are partly subjective. Expert reviewers may reasonably disagree, and there may be no single ground-truth answer. This internship will explore how AI judges can be calibrated to different scientific users or reviewer groups so they better reflect expert judgment, uncertainty, and disagreement while remaining grounded in evidence and scientific standards. The role is intended for a current PhD student with research experience in subjective alignment, human preference modeling, disagreement modeling, LLM-as-judge methods, or related evaluation methods who wants to apply that expertise to a real pharmaceutical R&D problem.

Requirements

  • Currently enrolled in a PhD program in NLP, machine learning, computer science, biomedical informatics, data science, human-centered AI, or a related field.
  • Research experience, publications, or active PhD work in one or more of the following areas: subjective alignment, human preference modeling, disagreement modeling, pluralistic alignment, LLM-as-judge methods, model evaluation, or evaluation under uncertainty.
  • Proficiency in Python and standard NLP, machine learning, or data science tooling.
  • Ability to translate qualitative, disagreement-heavy feedback into testable evaluation criteria, datasets, or analysis plans.
  • Ability to analyze model outputs, compare reviewer judgments, and identify patterns of agreement or disagreement.
  • Clear written and verbal communication skills, including the ability to explain research framing to scientific collaborators outside NLP or machine learning.
  • Interest in applying AI evaluation research to scientific or pharmaceutical R&D problems.

Nice To Haves

  • Familiarity with LLM-as-judge approaches, calibration methods, preference modeling, reward modeling, or representation-learning methods for qualitative attributes.
  • Exposure to retrieval-augmented generation, grounding, citation evaluation, or context-aware evaluation frameworks.
  • Interest or prior exposure to biomedical, scientific, clinical, regulatory, or other expert domains where evaluation requires judgment.
  • Experience designing expert review workflows, annotation guides, adjudication processes, or inter-rater agreement analyses.
  • Experience working with qualitative feedback, rubric design, benchmark construction, or human evaluation studies.

Responsibilities

  • Investigate how methods for modeling subjective or pluralistic human judgment can be adapted to scientific evaluation contexts.
  • Explore approaches for building AI judges that can be calibrated to different scientific users, reviewer groups, or evaluation styles.
  • Help define scientific quality beyond correctness, including how an AI judge might assess plausibility, reasoning quality, novelty, usefulness, and research value.
  • Compare approaches for representing evaluator uncertainty, disagreement, or multiple valid interpretations.
  • Help define a practical use case grounded in accessible data and available expert feedback.
  • Build small benchmark datasets or annotation samples that capture expert disagreement on scientific report or hypothesis quality.
  • Support the design of evaluation protocols that preserve evaluator uncertainty rather than collapsing feedback into a single score.
  • Analyze where AI judge outputs align or diverge from expert reviewers.
  • Work with scientific and data science colleagues to gather qualitative feedback and translate it into testable evaluation criteria.
  • Participate in reviews with the Evaluation & Standards team to check assumptions against how scientific experts reason about output quality.
  • Communicate research tradeoffs, limitations, and findings to both technical and non-technical collaborators.
  • Document methods, findings, and open questions so the team can build on the work after the internship.
  • Contribute to reusable evaluation assets, such as a calibrated AI judge approach, a disagreement-modeling method, an annotation guide, or a readiness rubric.
  • Prepare a final readout summarizing the research approach, findings, limitations, and recommended next steps.

Benefits

  • Company sponsored employee medical benefits
  • Sick time benefits (up to 40 hours per calendar year; for employees who reside in the State of Washington, up to 56 hours per calendar year)
  • Consolidated retirement plan (pension)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service