Senior Data Scientist

Omm IT SolutionsWoodlawn, MD
Onsite

About The Position

We are seeking a Senior Data Scientist to develop analytics solutions, focusing on data hygiene, management, and performance optimization. The role involves designing, implementing, and maintaining advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms. You will be responsible for cleaning, transforming, and managing large-scale datasets from diverse sources, ensuring data integrity, reliability, and security. Additionally, you will optimize complex SQL queries and database operations, participate in code reviews, enforce version control, and uphold best practices for code quality, reproducibility, and data privacy. The position also requires supporting data validation, testing, deployment, and post-implementation monitoring in a fast-paced environment.

Requirements

  • Master's degree and 10+ years of experience, Bachelor's degree and 12+ years of experience, or 18+ years of experience in lieu of a degree.
  • Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with experience in NLP, Text Processing, Information Extraction, Python, SQL, Regex and specialized libraries/frameworks.
  • Overall 10+ years’ experience in IT industry.
  • Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction including Named Entity Recognition and Address Standardization.
  • Practical knowledge of data matching strategies, including Blocking and Indexing, String Distance Metrics, TF-IDF/Cosine Similarity, and Phonetic Encoding.
  • Strong Python development skills for building analytics solutions and manipulating data.
  • Advanced SQL proficiency for complex data querying, optimization, and database operations.
  • Practical experience using Regex for advanced text processing, data cleansing, and pattern matching.
  • Familiarity with specialized libraries and frameworks including linkage libraries such as Splink / FastLink, Dedupe, or recordlinkage, and core Python data science libraries, specifically spaCy for NLP tasks and Scikit-Learn for general machine learning and clustering.
  • Familiarity with code reviews, version control, and maintaining data security and reproducibility standards.
  • Excellent communication skills.

Nice To Haves

  • Prior experience delivering IT or data initiatives within federal, state, or local government environments.
  • Proven ability to operate independently, take ownership of data pipeline architectures, and drive projects from data discovery through to post-implementation.
  • Experience retrieving, migrating, and manipulating data from legacy and distributed systems, including PostgreSQL, DB2, Oracle, SQL Server, and Hadoop, as well as unstructured flat files.
  • Experience utilizing Jenkins to automate continuous integration, testing, and deployment (CI/CD) for data validation pipelines.
  • Experience with pipeline automation tools to schedule and monitor complex data cleansing jobs.
  • Strong ability to translate complex algorithmic decisions (such as probabilistic match thresholds) into clear business logic for executive leadership and non-technical stakeholders.
  • Excellent problem-solving skills and proven verbal/written communication skills when collaborating across cross-functional teams.

Responsibilities

  • Develop Analytics Solutions: Design, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms.
  • Data Hygiene & Management: Clean, transform, and manage large-scale datasets from diverse, complex sources, ensuring absolute data integrity, reliability, and security.
  • Performance Optimization: Optimize complex SQL queries and database operations to ensure efficient data access, processing, and scalability.
  • Engineering Standards: Actively participate in code reviews, enforce version control, and uphold best practices for code quality, reproducibility, and data privacy.
  • End-to-End Delivery: Support data validation, testing, deployment, and post-implementation monitoring in a fast-paced environment.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service