Applied Data Scientist, Health AI Evaluation & Datasets

Innodata Inc.
$210,000 - $240,000

About The Position

Innodata is a global data engineering company focused on enabling the responsible advancement of artificial intelligence. We provide data, evaluation frameworks, and human expertise for building trustworthy AI systems at scale. Our mission is to support Generative AI/AI builders and adopters with transferable solutions, platforms, and services, building on our 36+ year legacy of delivering high-quality data and outstanding outcomes. Healthcare is a critical domain for generative AI, requiring clinical accuracy, patient safety, regulatory compliance, health equity, auditability, and workflow fit. Innodata collaborates with foundation model labs, medical AI startups, payers, providers, pharma, and digital health companies to build LLMs, multimodal systems, and AI agents for healthcare and life sciences. As an Applied Data Scientist, Health AI Evaluation & Datasets, you will be responsible for the design, measurement quality, and clinical validity of datasets used for training, fine-tuning, and evaluating health-domain models. This role requires a blend of clinical or biomedical fluency and data science rigor, enabling you to interpret clinical guidelines, payer policies, medical literature, and patient communication workflows, translate them into measurable datasets and evaluation plans, and effectively communicate your methodology to clinical, data science, and ML stakeholders. You will work collaboratively within a dedicated pod comprising a Technical Solutions Architect, Applied Research Scientist, AI/ML Research Engineer, and Language Data Scientists, ensuring that data, rubrics, review workflows, and measurement evidence are clinically realistic, statistically defensible, compliant, and valuable for evaluation and post-training.

Requirements

  • 5+ years of data science experience, with at least 2+ years involving healthcare, clinical, biomedical, payer, provider, pharma, life sciences, or comparable regulated health data.
  • Working knowledge of healthcare data and standards such as EHR structure, clinical documentation conventions, ICD-10, CPT, SNOMED CT, LOINC, RxNorm, and familiarity with FHIR, HL7, or equivalent interoperability concepts.
  • Hands-on experience designing ML datasets, including writing annotation guidelines, sizing cohorts, setting quality thresholds, designing QA checks, and delivering data for training or evaluation.
  • Familiarity with LLM-based health AI workflows, including prompt design, rubric-based evaluation, retrieval-augmented generation, LLM-as-judge methods, model comparison, and understanding the limitations of automated evaluation in clinical settings.
  • Strong Python and SQL skills; comfort with pandas, scikit-learn, statsmodels or equivalent tools; and working familiarity with modern LLM tooling such as Hugging Face, evaluation frameworks, prompt development tools, or model APIs.
  • Statistical literacy encompassing sampling design, bias and fairness analysis, inter-annotator agreement metrics (Cohen or Fleiss kappa, Krippendorff alpha), confidence intervals, significance testing, and error analysis.
  • Solid grasp of healthcare privacy, compliance, and governance, including HIPAA, de-identification standards (Safe Harbor and Expert Determination), practicalities of working with PHI safely, auditability, access control, and documentation suitable for regulated AI programs.
  • Ability to work credibly with clinicians, biomedical SMEs, research scientists, engineers, technical solutions teams, annotators, and customer stakeholders.
  • A bias toward clinical realism, prioritizing datasets that reflect real-world scenarios over those that appear impressive but are impractical.
  • Degree in a relevant field such as biostatistics, epidemiology, computational biology, health informatics, computer science with a health focus, statistics, a clinical degree with quantitative training, or equivalent demonstrated experience.

Nice To Haves

  • Clinical credentials (MD, RN, PharmD, MPH, PhD, or health informatics backgrounds) are encouraged, though not required, as candidates must be able to work credibly with clinicians and health AI customers.

Responsibilities

  • Translate customer goals into dataset specifications, taxonomies, rubrics, sampling plans, and acceptance criteria for applications like improving differential diagnosis, evaluating clinical note summarizers, testing RAG-based medical literature assistants, and creating preference data for patient-facing chatbots.
  • Focus on multimodal health AI by designing training and evaluation datasets across diverse data types including clinical text, medical images, waveforms, structured EHR data, claims, trial data, medical literature, patient communications, payer policies, drug information, and other clinical artifacts, for use cases such as clinical reasoning, medical QA, note summarization, medical coding, patient communication, utilization management, and literature synthesis.
  • Design evaluations for retrieval-augmented and source-grounded health AI systems, assessing aspects like evidence citation, faithfulness, contraindication handling, guideline adherence, source freshness, and failure modes arising from incomplete, conflicting, or stale context.
  • Define sampling strategies, label schemas, inter-annotator agreement targets, adjudication workflows, and SME review patterns in collaboration with Language Data Scientists, clinicians, biomedical experts, and quality teams.
  • Develop statistical and ML checks to ensure the trustworthiness of healthcare datasets, including stratified sampling across specialties and patient subgroups, bias and representation analysis, leakage detection, distribution shift checks, uncertainty estimates, reliability metrics, and subgroup performance analysis.
  • Partner with Applied Research Scientists and AI/ML Research Engineers to integrate datasets into evaluation and post-training pipelines, utilizing rubric-grounded LLM-as-judge prompts, regression suites, model comparison workflows, experiment tracking, and model-improvement feedback loops.
  • Evaluate health AI behavior beyond surface accuracy, assessing calibration, hallucination on safety-critical content, refusal appropriateness, robustness under ambiguity, equity across patient subgroups, and safe handoff in agentic or workflow-integrated systems, with a focus on clinical workflow fit, necessary evidence for trust, uncertainty surfacing, and risk differences across use cases.
  • Own data quality throughout the process, from source intake to delivery, covering de-identified clinical text, medical literature, synthetic cases, structured records, client policies, and knowledge bases, with attention to PHI/PII handling, provenance, audit trails, versioning, and compliance documentation.
  • Stay current with the health AI landscape, including regulatory developments (e.g., FDA guidance, EU AI Act), benchmark releases (e.g., MedQA, MedMCQA, HealthBench), and emerging clinical evaluation methodologies.
  • Support customer discovery and proposal work by scoping dataset programs, estimating annotation and SME review effort, identifying regulatory or data-access constraints, and explaining methodology choices to client leadership.
  • Contribute to Innodata's internal intellectual property by developing reusable health-domain taxonomies, evaluation rubrics, golden datasets, clinical review playbooks, dataset quality checks, and methodology templates.

Benefits

  • The expected salary range for this position is $210,000 – $240,000 USD per year, based on experience, skills, and qualifications.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service