Senior Data Scientist

INFO ORIGIN INCWoodlawn, MD
$80 - $90Onsite

About The Position

We are seeking an experienced Senior Data Scientist to join our team in Woodlawn, MD. The ideal candidate will have strong hands-on experience in Natural Language Processing (NLP), text processing, information extraction, Python, SQL, and data matching/entity resolution. This role will focus on developing advanced data processing solutions, building scalable data pipelines, improving data quality, and implementing techniques to identify and match records across large and complex datasets.

Requirements

  • Bachelor's or Master's degree in Computer Science, Statistics, Applied Mathematics, Information Science, or a related field.
  • 10+ years of IT experience.
  • Strong hands-on experience with Python and SQL.
  • Strong practical experience with NLP, text processing, and information extraction.
  • Experience with Named Entity Recognition (NER).
  • Experience with entity resolution, record linkage, data matching, or deduplication.
  • Knowledge of blocking/indexing and string distance metrics.
  • Experience with TF-IDF and cosine similarity.
  • Experience with Regex for text processing and data cleansing.
  • Knowledge of address standardization and phonetic encoding.
  • Strong analytical, problem-solving, written, and verbal communication skills.

Nice To Haves

  • Experience with spaCy
  • Experience with Scikit-learn
  • Experience with Splink
  • Experience with Dedupe
  • Experience with FastLink
  • Experience with recordlinkage
  • Experience with PostgreSQL
  • Experience with DB2
  • Experience with Oracle
  • Experience with SQL Server
  • Experience with Hadoop
  • Experience with Jenkins
  • Experience with CI/CD
  • Experience with Apache Airflow or other pipeline automation tools
  • Experience working on federal, state, or local government projects.
  • Experience working with large-scale, legacy, or distributed data systems.
  • Experience migrating and integrating data from multiple enterprise sources.
  • Experience designing automated data-quality and validation pipelines.
  • Ability to explain complex data-matching algorithms and business rules to technical and non-technical stakeholders.
  • Ability to independently own projects from data discovery through development, deployment, and post-production support.

Responsibilities

  • Design, develop, and maintain advanced data processing and entity resolution pipelines using Python and SQL.
  • Process, clean, transform, and manage large-scale datasets from multiple sources.
  • Apply NLP, text processing, and information extraction techniques to structured and unstructured data.
  • Implement data matching techniques such as Named Entity Recognition (NER), blocking and indexing, string similarity/distance metrics, TF-IDF, cosine similarity, phonetic matching, and address standardization.
  • Develop data cleansing and validation solutions using Python and Regex.
  • Write and optimize complex SQL queries and database operations for performance and scalability.
  • Perform data validation, testing, deployment, and production monitoring.
  • Participate in code reviews and follow version control, data security, reproducibility, and software development best practices.
  • Collaborate with technical teams and business stakeholders to understand requirements and translate complex data-processing logic into practical solutions.
  • Troubleshoot data quality and pipeline issues and implement reliable, scalable solutions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service