AI Data Engineer

Bright Vision TechnologiesTroy, MI
$80,000 - $100,000Remote

About The Position

We are seeking an AI Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
  • Strong proficiency in Python and at least one JVM or systems language.
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
  • Hands-on experience operating petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, and storage formats.
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows.
  • Familiarity with high-throughput data loading for accelerator-based training.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Nice To Haves

  • Experience with multimodal datasets at large scale.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.
  • Experience supporting frontier model training pipelines.

Responsibilities

  • Build and operate large-scale data systems for AI training and evaluation pipelines.
  • Focus on data ingestion, transformation, quality assurance, lineage, and high-throughput delivery.
  • Manage petabyte-scale data systems.
  • Ensure data infrastructure choices support model quality and training efficiency.
  • Implement dataset versioning, lineage, and reproducibility for ML workflows.
  • Support high-throughput data loading for accelerator-based training.
  • Apply strong software engineering practices including testing, CI/CD, and code review.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service