About The Position

We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, Apache Spark, AWS, distributed data processing, and scalable data architecture to build and operate infrastructure supporting AI-driven products and research initiatives. The role will focus on designing and scaling distributed data pipelines, managing large datasets across cloud environments, and building reliable systems for analytics, experimentation, and model development.

Requirements

  • Strong professional experience in data engineering or distributed data systems
  • Advanced proficiency in Python and SQL
  • Hands-on experience with Apache Spark or comparable distributed-processing frameworks
  • Strong experience with AWS data services and cloud-native architecture
  • Experience with SQL and NoSQL databases
  • Demonstrated experience processing large-scale datasets
  • Strong understanding of partitioning, performance optimisation, and scalable architecture
  • Familiarity with orchestration, automation, monitoring, and data-quality workflows

Nice To Haves

  • Exposure to AI/ML or research environments is advantageous
  • Familiarity with LLM training, evaluation, or experimentation datasets is beneficial
  • Experience with data-visualisation tools such as Matplotlib, Seaborn, or Plotly is a plus

Responsibilities

  • Design, build, and maintain large-scale pipelines for structured and unstructured data
  • Develop distributed processing workflows using Apache Spark or comparable frameworks
  • Optimise transformations, partitioning strategies, and computational workloads
  • Identify and resolve performance bottlenecks across high-volume data systems
  • Support downstream analytics, experimentation, and model-development requirements
  • Design scalable AWS-based data architectures across SQL and NoSQL systems
  • Build reliable ingestion, transformation, storage, and distribution workflows
  • Write efficient Python and SQL for production data processing
  • Evaluate storage and database technologies against workload requirements
  • Improve scalability, maintainability, accessibility, and operational efficiency
  • Implement monitoring, validation, and automation across data workflows
  • Identify failures, anomalies, and data-quality issues
  • Maintain integrity and reliability throughout pipelines and storage layers
  • Collaborate with AI researchers, data scientists, and engineering teams
  • Support data infrastructure for AI/ML training, evaluation, and experimentation

Benefits

  • Full-time engagement
  • Fully remote
  • Base compensation: $140,000–$180,000/year
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service