Senior Data Engineer

Red HatRaleigh, NC
Remote

About The Position

This is a telecommuting role to be performed anywhere in the U.S. The Senior Data Engineer will architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation. This includes developing macro-driven data quality tests and validation frameworks. The role involves orchestrating pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics. The engineer will administer the enterprise Databricks platform, configure IAM roles for secure Amazon S3 bucket access, manage application credentials and secrets, and establish workspace governance policies and cluster configurations. Additionally, the role requires designing and deploying intelligent retrieval architecture and AI-driven workflows using vector-based search methods, building marketing retrieval and decision-automation applications, and operationalizing MLOps methodologies using MLflow. The engineer will implement end-to-end machine learning models, deliver stakeholder-facing analytical outputs, manage CI/CD pipelines, build and manage container images, lead the deployment and maintenance of containerized data science models on Red Hat OpenShift, and lead application security initiatives. This includes completing security compliance assessments, performing static application security testing, executing vulnerability scanning, completing Privacy Impact Assessments, and conducting threat modeling. Collaboration with enterprise information security teams for vulnerability remediation, compliance audits, and centralized logging/monitoring through Splunk is also a key responsibility.

Requirements

  • Master's degree (U.S. or foreign equivalent) in Computer Science or related field and three (3) years of experience in the job offered or related role OR Bachelor's degree (U.S. or foreign equivalent) in Computer Science or related field and five (5) years of experience in the job offered or related role.
  • Three (3) years of experience with architecting and implementing high-volume data pipelines between cloud data warehouse (Snowflake) and lakehouse (Databricks) platforms using PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality test frameworks and validation logic.
  • Three (3) years of experience orchestrating and scheduling data pipeline workflows using Apache Airflow, including configuring DAG-based dependency management, automated failure recovery, and pipeline monitoring for enterprise analytics workloads.
  • Three (3) years of experience administering enterprise Databricks environments, including configuring IAM roles for secure cloud object storage (Amazon S3) access, managing application secrets through platform vault systems and OpenShift secrets, and establishing workspace governance and cluster policies for cross-functional teams.
  • Three (3) years of experience implementing end-to-end machine learning models by: 1) building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise text analysis; 2) developing predictive models using gradient boosting frameworks (XGBoost) and Scikit-learn; 3) constructing deep learning architectures using Keras; and 4) designing time-series forecasting models for event-based prediction.
  • Three (3) years of experience delivering full-scale information retrieval systems for enterprise data by researching, evaluating, and implementing Transformer architectures and Transfer Learning methodologies using deep learning frameworks for semantic search, text classification, and vector-based clustering.
  • Three (3) years of experience operationalizing MLOps methodologies using MLflow for experiment tracking and model registry management, and implementing automated post-production model monitoring to track performance degradation and optimize predictive accuracy.
  • Three (3) years of experience managing CI/CD pipelines using Git and Tekton, building and publishing container images using buildah and skopeo to internal container registries, and deploying containerized applications on Red Hat OpenShift (Kubernetes) with network route management, TLS termination, and high-availability configurations.
  • Three (3) years of experience leading enterprise security compliance assessments, including performing static application security testing (SAST) using SonarQube, executing vulnerability scans using Qualys, completing Privacy Impact Assessments (PIA), and conducting STRIDE-based threat modeling.

Responsibilities

  • Architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality tests and validation frameworks.
  • Orchestrate pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics.
  • Administer the enterprise Databricks platform by configuring IAM roles for secure Amazon S3 bucket access, managing application credentials and secrets through Databricks' built-in vault system and OpenShift secrets, and establishing workspace governance policies and cluster configurations for cross-functional data science and engineering teams.
  • Design and deploy intelligent retrieval architecture and AI-driven workflows using vector-based search methods and enterprise data platforms, building marketing retrieval and decision-automation applications that integrate multiple data sources and APIs.
  • Operationalize MLOps methodologies using MLflow for experiment tracking and model registry management, and Lakehouse monitoring for automated post-production model performance tracking to optimize predictive accuracy and increase marketing return on investment.
  • Implement end-to-end machine learning models and deliver stakeholder-facing analytical outputs by building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise systems analysis, developing predictive models using XGBoost and Scikit-learn, constructing deep learning architectures using Keras, and designing time-series forecasting models for event-based user adoption prediction.
  • Manage CI/CD pipelines using Git and Tekton to ensure reliable, repeatable code delivery for production applications.
  • Build and manage container images using buildah and skopeo, pushing to internal container registries for deployment.
  • Lead the deployment and maintenance of containerized data science models and enterprise applications on Red Hat OpenShift (Kubernetes), managing network routes, TLS termination, and container orchestration for high-availability services.
  • Lead application security initiatives by completing comprehensive enterprise security compliance assessments encompassing 20+ security controls across the full technology stack, aligned with industry frameworks such as NIST and CIS Controls.
  • Perform static application security testing (SAST) using SonarQube, execute vulnerability scanning using Qualys and pip-audit, complete Privacy Impact Assessments (PIA), and conduct STRIDE-based threat modeling.
  • Collaborate with enterprise information security teams to remediate identified vulnerabilities, navigate compliance audits, and maintain centralized logging and monitoring through Splunk.

Benefits

  • bonus
  • commission
  • equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service