About The Position

At Apple, great ideas have a way of becoming phenomenal products, services, and customer experiences very quickly. Our team is building a massive, real-time platform that transforms continuous streams of multimodal data (including structured, image, and log data) into an intelligent, searchable foundation. We are seeking a Principal Data Engineer to lead and drive not only our team's data processing systems, but also to partner at a larger scale, coordinating and synching strategically with other business groups and organizations within Apple. We are seeking a Principal Data Engineering Lead with deep expertise in ETL/ELT, data architecture, and applied ML pipelines to drive the design, build, and operations of this infrastructure. As a key member of our team, you will be responsible for driving critical decisions and operations across the entire system while aligning strategically across Apple.

Requirements

  • Masters Degree
  • 12+ years of experience in data engineering, including building and maintaining large-scale ETL/ELT data pipelines
  • Proficiency in data modeling, especially dimensional modeling, and designing schemas optimized for analytics and reporting
  • Experience with leveraging databases including SQL/NoSQL Databases (including Postgres / Cassandra / Redis)
  • Strong experience with distributed data processing frameworks including Apache Spark
  • Strong experience with Parallel processing frameworks: BigTable/Hadoop
  • Strong software engineering fundamentals and proven experience with Scala, Java
  • Hands-on experience with Apache Kafka, Iceberg, and Flink.
  • Experience with workflow orchestration tools including Apache Airflow and Beam
  • Experience with AWS: e.g., S3, EMR, Lambda, Glue, Redshift, BigQuery, Kinesis, or similar services
  • Experience with Analytics frameworks including Trino (Presto, BigQuery, Snowflake)
  • Hands-on experience with big data lake architectures
  • Experience with containerization and orchestration (Docker, Kubernetes/EKS) and CI/CD tooling including Jenkins
  • Experience in Python and PySpark
  • Familiarity with graph databases such as TigerGraph
  • Experience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation
  • Hands-on experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar).
  • Experience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline
  • Knowledge of data governance principles, data security best practices, and data privacy regulations
  • Proven experience delivering a consumer-oriented solution by participating at every stage of the development life-cycle.
  • Excellent communication skills and a collaborative mindset with past experience presenting and partnering with VP and C level decision makers.

Nice To Haves

  • Experience with data versioning tools and frameworks (e.g., DVC, Delta Lake)
  • Experience storing/serving embeddings (e.g., pgvector, Milvus, FAISS)

Responsibilities

  • Drive the design, build, and operations of data processing systems.
  • Partner at a larger scale, coordinating and synching strategically with other business groups and organizations within Apple.
  • Drive critical decisions and operations across the entire system while aligning strategically across Apple.
  • Build and maintain large-scale ETL/ELT data pipelines.
  • Design schemas optimized for analytics and reporting.
  • Build pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation.
  • Deploy, serve, and optimize LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar).
  • Tune batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline.
  • Deliver a consumer-oriented solution by participating at every stage of the development life-cycle.
  • Present and partner with VP and C level decision makers.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service