Data Engineer

Innodata Inc.
$100,000 - $120,000

About The Position

Innodata is a global data engineering company focused on the intersection of data and Artificial Intelligence (AI). Our mission is to enable the responsible advancement of AI by providing the necessary data, evaluation frameworks, and human expertise to build trustworthy AI systems at scale. We offer a variety of transferable solutions, platforms, and services for Generative AI / AI builders and adopters, leveraging our 36+ year legacy of delivering high-quality data and exceptional customer outcomes. This role is designed to develop and implement enterprise data warehouses, data lakes, and pipelines that support data-driven decision-making within data center supply chain and real estate operations. The Data Engineer will be responsible for creating scalable, secure, and optimized ETL infrastructure on GCP/AWS, facilitating advanced AI/ML use cases such as RAG, copilots, and agentic AI for predictive analytics and workflow automation.

Requirements

  • Advanced proficiency in SQL (complex queries, optimization) and Python (data engineering, scripting, APIs).
  • Experience building ETL/ELT pipelines operating on structured and unstructured data sources.
  • Knowledge of enterprise data warehouse and data lake architectures.
  • Exposure to data pipelines for AI/ML (vector DB ingestion, embeddings, RAG pipelines, copilots, agents).
  • Strong hands-on expertise with GCP services: BigQuery, Dataflow, Pub/Sub, Cloud Storage, Looker/BI (or similar, preferred).

Nice To Haves

  • Familiarity with supply chain or data center operations data is a strong plus.
  • Experience with ML Engineering, data visualization tools (Looker, Tableau, Power BI) and MLOps practices.

Responsibilities

  • Design and implement data-driven solutions on GCP including BigQuery, Cloud Storage, Dataflow, Pub/Sub, and Looker/BI.
  • Build ETL scripts using SQL and Python to extract, clean, and transform structured and unstructured data from ERP, procurement, logistics, and facility management systems.
  • Develop and optimize data pipelines for ingestion, transformation, and loading into enterprise data lakes and warehouses.
  • Build and extend end-to-end data and BI solutions, spanning extraction, storage, transformation, and visualization layers.
  • Partner with supply chain, real estate, and AI/ML teams to provide pipelines for AI solutions (e.g., RAG ingestion, Copilot integration, multi-agent workflows).
  • Ensure data governance, lineage, and compliance across supply chain datasets.
  • Continuously optimize query performance, ETL processes, and pipeline reliability.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service