About The Position

Ancestry is seeking an exceptional and highly motivated Data Science Co-Op to join our Content AI team, a dynamic group at the forefront of Document Understanding. You’ll play a vital role in developing innovative AI models that extract and organize text and image information from billions of historical and genealogical records enabling customers to discover, share, and connect with their family history. As a Co-Op on the Content AI team, you will build, train and fine-tune models that process historical documents to detect meaningful, personalized insights within historical documents that connect people to their ancestors. You will also work closely with engineering teams to train, optimize, and deploy models that promote product development, customer success, and content creation across our Family History business.

Requirements

  • Currently pursuing an advanced degree (Master's or PhD preferred) in Computer Science, Data Science, Statistics, Mathematics, Linguistics, Engineering or related quantitative field with a strong data focus.
  • Specialization in generative AI & LLMs, embeddings, LoRA, QLoRA, vector databases, transformer models, Natural Language Processing (NLP), with software development expertise including data structures, distributed model training, and inference optimizations.
  • Exhibit strong proficiency in Python and relevant tools and libraries, including those for transformer models, multi-modal models, and general NLP (e.g., Hugging Face Transformers, agentic frameworks and workflows, LangChain, LangGraph, NLTK).

Nice To Haves

  • Familiarity with cloud platforms and related AI/ML services such as Google Gemini API, Vertex AI, AWS EC2, S3, SageMaker, Model Registry, and Bedrock is a plus.

Responsibilities

  • Implement and experiment with cutting-edge transformer and generative AI solutions for key Document Understanding tasks, including OCR, handwriting recognition, transcription, Named Entity Recognition (NER), Relation Extraction (RE), Coreference Resolution, Summarization, and Knowledge Graphs working with diverse genealogical and historical collections spanning newspapers, city directories, family history books, and vital records (birth, marriage, death).
  • Evaluate the performance of multi-modal models in zero-shot and few-shot learning scenarios for comprehensive document understanding.
  • Partner closely with ML Ops and Data Science Engineers to seamlessly deploy datasets, truth sets, models, and pipelines for training and inference in cloud environments.
  • Clearly and confidently present your findings, deliverables, and proposed solutions to technical and non-technical audiences, including teams, stakeholders, and executives.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service