Cactus Wellhead - Sr. Data Platform Engineering

Cactus WellheadHouston, TX
Hybrid

About The Position

The Data Platform Lead is responsible for implementing, configuring, and governing the technical foundation of the company’s enterprise data Lakehouse on Azure Databricks. This role combines hands-on data platform ownership with data engineering leadership, ensuring that data from ERP systems, product SaaS PostgreSQL databases, APIs, and other enterprise sources is ingested, modeled, governed, secured, and prepared for analytics, reporting, automation, and AI use cases. This position serves as the internal technical owner for Databricks platform standards, Bronze/Silver/Gold data architecture, Unity Catalog governance, pipeline design patterns, source-to-target mapping standards, data quality implementation, vendor technical review, and production readiness. The role is hands-on, with responsibility for platform configuration, environment setup, access controls, catalog/schema structure, compute standards, and operational readiness, while also providing technical directions to data engineers, contractors, vendors, and future internal team members through strong standards, practical architecture, and disciplined delivery.

Requirements

  • 7+ years of experience in data engineering, data architecture, cloud data platforms, or enterprise data integration.
  • Hands-on experience designing and supporting production data pipelines in cloud environments.
  • Experience with modern Lakehouse architecture, preferably Azure Databricks and Delta Lake.
  • Experience working with enterprise source systems such as ERP, CRM, SaaS applications, PostgreSQL, SQL Server, Oracle, or similar relational databases.
  • Experience reviewing vendors or contractor technical deliverables and enforcing engineering standards.
  • Strong expertise with Azure Databricks, Apache Spark, Python/PySpark, advanced SQL, Delta Lake, and Lakehouse design patterns.
  • Working knowledge of Unity Catalog, RBAC, lineage, data classification, metadata, and access governance.
  • Experience with Azure Data Lake Storage, GitHub or Azure DevOps, CI/CD, secrets management, and cloud integration patterns.
  • Strong understanding of data modeling, dimensional modeling, normalized models, medallion architecture, data quality, and reconciliation.
  • Ability to design ingestion patterns for ERP data, application databases, APIs, files, and incremental source changes.
  • Ability to operate as both a hands-on technical lead and a manager of delivery standards.
  • Strong judgment to challenge designs that are not secure, scalable, documented, or production ready.
  • Strong communication skills with technical teams, business stakeholders, vendors, and leadership.
  • Ability to translate business data needs into scalable platform and engineering solutions.
  • Strong ownership mindset, documentation discipline, and ability to work in a growing data organization with evolving standards.

Nice To Haves

  • Experience in manufacturing, oil & gas, field services, industrial operations, or ERP-heavy environments.
  • Exposure to Power BI, Tableau, semantic models, reporting migration, or analytics product delivery
  • Exposure to ML/AI pipelines, feature engineering, GenAI use cases, automation, or AI-ready data product development.
  • Experience with infrastructure-as-code, automated testing, data observability, or enterprise data catalog tools.
  • Databricks, Azure Data Engineer, Azure Solutions Architect, or related cloud/data certifications.

Responsibilities

  • Own the technical architecture for the Azure Databricks Lakehouse, including workspace structure, catalogs, schemas, compute patterns, storage strategy, and environment separation.
  • Define and maintain bronze, silver, and gold layer standards, including naming conventions, table ownership, audit columns, refresh patterns, and production readiness criteria.
  • Implement and govern Unity Catalog standards for access control, lineage, data classification, catalog/schema organization, and least-privilege access.
  • Partner with cybersecurity, infrastructure, and DevOps teams to align Databricks with enterprise identity, networking, secrets management, monitoring, and compliance expectations.
  • Establish cost controls, cluster policies, job standards, and usage monitoring to ensure the platform is reliable, scalable, and cost-effective.
  • Design, build, and oversee production-grade data pipelines using Databricks, Spark, Python/PySpark, SQL, Delta Lake, and approved orchestration patterns.
  • Lead ingestion from ERP systems, product PostgreSQL databases, SaaS platforms, APIs, files, and other enterprise data sources into the Lakehouse.
  • Define engineering patterns for full loads, incremental loads, CDC where applicable, reprocessing, error handling, logging, reconciliation, and pipeline recovery.
  • Ensure every production pipeline includes source-to-target mapping, ownership, data quality rules, monitoring, alerting, and operational handover documentation.
  • Review vendor and contractor deliverables for technical quality, maintainability, security, performance, and production readiness.
  • Implement practical data quality controls for completeness, uniqueness, validity, freshness, referential integrity, and reconciliation to source systems.
  • Support data governance by ensuring datasets have clear owners, stewards, classifications, lineage, refresh frequency, and required documentation before go-live.
  • Work with business, ERP, and product teams to understand source system meaning, schema changes, business logic, and downstream impact.
  • Enable certified silver and gold datasets that can support analytics, executive reporting, operational dashboards, automation, and AI/ML use cases.
  • Own technical incident response for pipeline failures, data refresh issues, root cause analysis, and corrective actions.
  • Define enterprise data architecture standards, reference architectures, and long-term roadmap.
  • Establish data domain ownership and enterprise data governance models.
  • Lead architecture decisions for analytics, AI/ML, master data, and enterprise reporting platforms.
  • Define standards for semantic models, reusable data products, and self-service analytics.
  • Act as the internal technical authority for Databricks platform and data engineering decisions.
  • Provide direction to vendors, contractors, and future internal data engineers to ensure delivery follows company standards.
  • Partner with Cactus IT Leadership team on execution, prioritization, architecture decisions, production risk, and vendor acceptance.
  • Collaborate with analytics, business applications, ERP, product engineering, cybersecurity, infrastructure, and business stakeholders.
  • Promote engineering discipline, documentation quality, reusable patterns, and operational excellence across the data function.

Benefits

  • None explicitly mentioned
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service