Senior Data Quality & Active Learning AI Engineer (Plymouth)

PhilipsPlymouth, MN
$156,240 - $187,488Onsite

About The Position

Build compliant, audit-ready medical imaging datasets and active learning systems that accelerate the development of artificial intelligence (AI) and non-AI software for Philips products. Own the design and operation of compliant data-quality and active-learning pipelines for multimodal clinical data, primarily medical images and metadata, on cloud infrastructure. Establish data contracts, de-identification workflows, validation and release gates, dataset versioning, readiness metrics, and traceable documentation so data are audit-ready and fit for model and software development. Join the growing Data and AI Platforms team within the Software and AI organization in research and development, working closely with data and AI engineers, medical informaticians, architects, data stewards, clinical partners, privacy experts, and external annotation teams. Partner with algorithm, AI, and software developers who use these data to create AI-enabled and non-AI software that becomes part of Philips products. Build active-learning loops that prioritize labeling, measure label quality and coverage, and accelerate safe dataset improvement; collaborate on preprocessing, augmentation, feature extraction, embeddings, and handoff to the data platform. You will help shape modern data and AI engineering practices while growing your expertise in regulated medical-device development, cloud platforms, synthetic and simulated data, and AI-assisted and agentic engineering workflows.

Requirements

  • 5+ years of relevant individual-contributor experience in data quality, data engineering, or machine learning data pipelines, ideally in healthcare and with medical imaging data.
  • Bachelor's degree in engineering, computer science, biomedical or health informatics, data science, mathematics, statistics, or a related field; a master's degree with 3+ years of relevant experience or a doctorate with relevant experience is also acceptable.
  • Python and Structured Query Language (SQL) for data extraction, transformation, validation, and automation.
  • Experience building scalable, event-driven and batch-oriented data pipelines in Amazon Web Services (AWS) using distributed processing frameworks such as Apache Spark or Ray.
  • Ability to design workflows that trigger when new imaging or metadata becomes available and apply data-quality checks, lineage, dataset versioning, and release gates.
  • Hands-on experience with Digital Imaging and Communications in Medicine (DICOM) and Picture Archiving and Communication System (PACS) concepts, 2D/3D imaging volumes, de-identification, and protected health information (PHI) handling.
  • Understanding of the Health Insurance Portability and Accountability Act (HIPAA), General Data Protection Regulation (GDPR), privacy by design, and audit-ready documentation under medical-device design controls.
  • Proficiency with AI-assisted engineering tools such as OpenAI Codex and Claude Code, plus agentic AI workflows.
  • Clear, collaborative communicator who surfaces risks, resolves blockers, shares progress and knowledge gaps, and works effectively with globally distributed technical, clinical, governance, and software stakeholders.

Nice To Haves

  • Health Level Seven (HL7) and Fast Healthcare Interoperability Resources (FHIR) familiarity.
  • Active-learning expertise, including uncertainty, diversity, and error-based sampling, label-quality measurement, coverage analysis, and bias/drift assessment.
  • Proficiency with PyTorch or TensorFlow.
  • Graphics processing unit (GPU)-accelerated preprocessing.
  • Experience creating, validating, or assessing synthetic or simulated data.

Responsibilities

  • Own the design and operation of compliant data-quality and active-learning pipelines for multimodal clinical data, primarily medical images and metadata, on cloud infrastructure.
  • Establish data contracts, de-identification workflows, validation and release gates, dataset versioning, readiness metrics, and traceable documentation so data are audit-ready and fit for model and software development.
  • Build active-learning loops that prioritize labeling, measure label quality and coverage, and accelerate safe dataset improvement.
  • Collaborate on preprocessing, augmentation, feature extraction, embeddings, and handoff to the data platform.
  • Help shape modern data and AI engineering practices while growing expertise in regulated medical-device development, cloud platforms, synthetic and simulated data, and AI-assisted and agentic engineering workflows.

Benefits

  • Generous PTO
  • 401k (up to 7% match)
  • HSA (with company contribution)
  • Stock purchase plan
  • Education reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service