Cactus Wellhead - Data Engineer & Architect

Cactus WellheadPiney Point Village, TX
Hybrid

About The Position

The Senior Data Engineer / Data Architect is responsible for designing, building, and governing the enterprise data platform on Azure Databricks. This role owns both hands-on data engineering delivery and data architecture design, ensuring that enterprise data is reliable, scalable, secure, and ready to support analytics, reporting, and AI initiatives. The position plays a critical role in transforming data from ERP, CRM, and other enterprise systems into trusted, reusable data products that enable business insights and future AI capabilities. The role operates with high ownership across the full lifecycle, from ingestion and transformation to data modeling, governance, and platform optimization, aligned with a modern Lakehouse architecture.

Requirements

  • 6–10+ years of experience in data engineering, data architecture, or enterprise data platforms
  • Hands-on experience building and supporting data pipelines in cloud environments
  • Experience working with enterprise systems (ERP, CRM, or similar)
  • Strong expertise in Azure Databricks / Apache Spark
  • Strong expertise in Python / PySpark
  • Strong expertise in SQL (advanced)
  • Experience with Azure Data Lake Storage (ADLS)
  • Experience with CI/CD with Github
  • Experience with API and data integration patterns
  • Strong understanding of data modeling (dimensional and normalized)
  • Strong understanding of data governance and security (RBAC, lineage)

Nice To Haves

  • Experience with Delta Lake and Lakehouse architecture
  • Exposure to ML/AI pipelines or data science workflows
  • Experience in manufacturing, oil & gas, or ERP-heavy environments
  • Familiarity with modern programming languages and frameworks (C#, .NET, SQL, JavaScript)
  • Experience with DevOps, CI/CD, and source control for ERP development
  • Knowledge of cloud-based ERP solutions and migration strategies
  • Strong background in data integration, ETL, and reporting tools (Power BI, SSRS, Tableau)

Responsibilities

  • Design and build end-to-end data pipelines (batch and near real-time) using Databricks and Spark
  • Implement and maintain Bronze / Silver / Gold data architecture layers
  • Ingest data from ERP, SaaS applications, APIs, and legacy systems into Lakehouse
  • Optimize pipelines for performance, scalability, and cost efficiency
  • Implement CI/CD pipelines using GitHub
  • Define and maintain the enterprise data architecture and standards
  • Design logical and physical data models across core business domains (finance, operations, field service)
  • Establish consistent definitions for critical entities (customer, job, revenue, etc.)
  • Ensure data structures support both analytics and AI/ML workloads
  • Implement and manage Unity Catalog–based governance (RBAC, lineage, auditability)
  • Define and enforce data standards for naming, access control, and data quality
  • Ensure compliance with security policies (e.g., sensitive HR and financial data protection)
  • Establish monitoring, validation, and reconciliation processes to ensure data accuracy and reliability
  • Define scalable patterns for data ingestion, transformation pipelines, and data access (BI tools, APIs, AI models)
  • Integrate Databricks with Azure services (ADLS, APIs, enterprise systems)
  • Ensure platform reliability, observability, and operational excellence
  • Prepare and structure data for AI/ML use cases and automation initiatives
  • Partner with analytics and AI teams to deliver datasets supporting predictive models and AI-driven workflows
  • Enable reusable datasets and feature-ready data pipelines
  • Partner with business stakeholders, application teams, and analytics teams to translate requirements into scalable solutions
  • Provide technical leadership for data engineering best practices and architecture decisions
  • Mentor junior engineers or contractors as the platform scales
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service