Principal Data Scientist

WileyHoboken, NJ
$140,000 - $200,733

About The Position

We are seeking a Principal Data Scientist to own domain-specific content modeling work end-to-end, from the eval set through the pipeline stage that ships it. You will join a small, senior team where data scientists own their models in production. This is a hands-on role for someone who wants to see their models through to real users in a rapidly evolving market. The role involves building production NLP pipelines that process millions of journal articles, extracting entities, classifications, claim tuples, and summaries optimized for use by downstream agentic applications.

Requirements

  • Deep Python experience in production, at scale, with a clear understanding of concurrency tradeoffs (asyncio vs. threads vs. queues).
  • Strong NLP background covering modern (LLMs, transformers, embeddings, retrieval) and classical (NER, classification, sequence labeling) approaches.
  • Experience building and learning from NLP evaluations.
  • A habit of comparing approaches and choosing the right one for the task, backed by evaluation and cost estimates.
  • A track record of shipping systems that deliver value to real users.

Nice To Haves

  • Experience working with scientific or scholarly text.
  • Familiarity with AWS (S3, Batch, Lambda, SageMaker) and Parquet or Iceberg data lake patterns.
  • Experience running LLMs under real cost and latency budgets in production.
  • Some exposure to agentic AI applications: tool use, multi-step reasoning, guardrails, and evaluation of trajectories.

Responsibilities

  • Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientific full-text at scale.
  • Compare NLP approaches to extraction and enrichment against LLM-based approaches, and pick the right tool for each task, defending choices with evaluation, cost, and operational tradeoffs.
  • Own evaluation: Build golden sets, choose metrics, and make tradeoffs between speed, quality, and cost.
  • Write production-quality Python, managing concurrency and cost for high-volume LLM workloads, and structuring code for engineers and other data scientists.
  • Collaborate with data engineers to orchestrate work in data pipeline and data build tools like Airflow and Dagster, designing reliable and evaluable pipeline stages.
  • Contribute to agentic AI application work, shaping how agents ground and defend their answers.
  • Work directly with editors, product managers, and engineers, bringing a modeling perspective to product decisions and translating stakeholder feedback into modeling work.

Benefits

  • Continual learning
  • Internal mobility
  • Meeting-free Friday afternoons
  • Robust employee programming
  • Competitive compensation
  • Comprehensive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service