Principal Data Engineer

SS&C TechnologiesWaltham, MA
$160,000 - $170,000Hybrid

About The Position

We are seeking a highly skilled and hands-on Principal Data Engineer to lead the design, development, and deployment of scalable AI/ML data platforms, distributed data processing systems, and cloud-native data services. This role requires deep expertise in Python and Java-based backend engineering, microservices architecture, machine learning pipelines, distributed systems, cloud-native platforms, Kubernetes, and AWS technologies. The ideal candidate will have extensive experience building enterprise-scale data platforms, developing production-ready ML pipelines, implementing scalable microservices, and driving engineering best practices. This role will collaborate closely with Product, Architecture, Data Science, Application Development, Analytics, SRE, and DevOps teams to deliver highly scalable, reliable, and intelligent data solutions.

Requirements

  • Strong expertise in Python, Java, Spring Boot, REST API development, and Microservices Architecture.
  • Experience building production-grade AI/ML platforms, ML pipelines, and distributed data processing applications.
  • Strong understanding of ML SDLC, MLOps, model deployment, and productionizing Python/Java applications.
  • Hands-on experience with Apache Kafka, Kafka Connect, Kafka Streams, Apache Flink, Spark, Airflow, or similar distributed data processing technologies.
  • Extensive experience designing and developing cloud-native applications on AWS.
  • Solid expertise in Kubernetes, Docker, Terraform/CloudFormation, ECS, and EKS environments.
  • Experience with ML frameworks such as PyTorch, TensorFlow, Keras, or scikit-learn.
  • Strong knowledge of Oracle, PostgreSQL, Amazon Redshift, Amazon Aurora, MongoDB, and vector databases such as Milvus, Pinecone, or Chroma.
  • Experience with data modeling, feature engineering, database optimization, query tuning, indexing, partitioning, and performance improvement strategies.
  • Proven experience developing and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Maven/Gradle, SonarQube, and Infrastructure as Code.
  • Strong expertise in monitoring, logging, and observability tools including CloudWatch, Prometheus, Grafana, ELK Stack, Splunk, and OpenTelemetry.
  • Strong Linux proficiency and software engineering best practices.

Nice To Haves

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, Data Science, or a related field.
  • 8+ years of software engineering, data engineering, or AI platform engineering experience.
  • Experience building scalable ML pipelines, feature engineering workflows, and enterprise AI platforms.
  • Experience with Generative AI, Retrieval-Augmented Generation (RAG), LLM deployment, embedding pipelines, and semantic search technologies.
  • Experience with ML orchestration frameworks such as Kubeflow, MLflow, Airflow, or similar platforms.
  • Strong experience designing scalable, fault-tolerant, and highly available distributed systems.
  • Experience leading enterprise-scale platform initiatives and mentoring engineering teams.
  • Excellent problem-solving, communication, architectural, and technical leadership skills.

Responsibilities

  • Lead the design and development of scalable AI/ML data platforms, distributed data processing systems, and cloud-native applications.
  • Design and implement end-to-end ML pipelines including data ingestion, feature engineering, model training, validation, deployment, monitoring, and automated retraining.
  • Build scalable batch and streaming data pipelines using technologies such as Apache Kafka, Apache Flink, Spark, or similar distributed processing frameworks.
  • Develop scalable microservices, REST APIs, reusable platform services, and enterprise data processing components.
  • Design and implement Retrieval-Augmented Generation (RAG) pipelines, embedding generation services, semantic search capabilities, and vector database integrations.
  • Drive platform modernization, technical design reviews, engineering standards, and adoption of innovative technologies to improve scalability, reliability, performance, and operational efficiency.
  • Design and maintain cloud-native infrastructure, CI/CD pipelines, deployment automation, containerized applications, and ML deployment workflows using Kubernetes, Docker, Terraform/CloudFormation, ECS/EKS, and AWS services including EC2, S3, Lambda, Redshift, Aurora, RDS, Glue, CloudWatch, and MSK.
  • Design optimized relational, NoSQL, and vector data models using PostgreSQL, MongoDB, Redshift, Aurora, Milvus, Pinecone, Chroma, or similar technologies, including performance tuning, indexing, partitioning, and query optimization.
  • Troubleshoot complex production issues, perform root-cause analysis, and collaborate with SRE and DevOps teams to improve platform stability, scalability, deployment automation, and ML operational workflows.
  • Provide technical leadership, mentorship, and guidance to engineering teams while driving best practices, governance, architecture, and continuous improvement initiatives.

Benefits

  • medical, dental, and vision coverage
  • a 401(k) plan with company match
  • paid time off, holidays, and parental leave
  • professional development reimbursement opportunity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service