Senior Data QA-Onshore

V4C.ai
Onsite

About The Position

v4c.ai was founded with a clear goal: to make data, AI, and machine learning accessible and impactful for every organization. As a Databricks partner, we deliver end-to-end solutions that transform complex data challenges into strategic outcomes. We are seeking a Senior Data QA Automation Engineer to lead the quality strategy, design, and implementation of automated testing frameworks for our big data platforms. In this senior role, you will own the end-to-end data validation strategy within our Databricks Lakehouse architecture, ensuring high-quality, reliable, and compliant data across Delta Lakes, ETL pipelines, and enterprise data models. You will work closely with Data Engineering leadership to establish rigorous quality gates and mentor mid-to-junior engineers on data testing best practices.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related quantitative field.
  • 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET), with at least 2+ years of dedicated experience architecting test automation in Databricks.
  • Mastery of Python and PySpark (DataFrames and SQL APIs) for processing and profiling large datasets.
  • Deep expertise in writing advanced SQL queries, optimization techniques, and understanding Spark query execution plans.
  • Hands-on mastery of big-data validation libraries (e.g., Great Expectations, pytest, Delta Live Tables expectations).
  • Strong operational knowledge of Databricks deployment on a major cloud provider (AWS, Azure, or GCP).

Nice To Haves

  • Databricks Certified Data Engineer Professional or Databricks Certified Machine Learning Professional.
  • Experience validating real-time event-streaming architectures (Kafka, Event Hubs, Kinesis).
  • Solid understanding of DataOps culture, testing infrastructure as code, and data observability principles.

Responsibilities

  • Architect, build, and scale automated test frameworks from scratch natively within Databricks using PySpark, Python, and SQL.
  • Design robust automated assertions for Delta Lake tables, including checking data drift, schema evolution, and historical data validation via time-travel functions.
  • Code complex automated scenarios to validate large-scale batch and real-time streaming data pipelines (Structured Streaming), ensuring source-to-target integrity.
  • Programmatically verify data lineage, audit logs, and access controls implemented via Databricks Unity Catalog.
  • Lead the integration of automated data quality tests into enterprise CI/CD pipelines (e.g., Azure DevOps, GitHub Actions), leveraging Databricks Workflows, APIs, or Airflow.
  • Act as the subject matter expert for data quality; mentor junior team members, establish QA standards, and advocate for data quality principles across engineering teams.
  • Design and execute automated performance and scalability tests on Spark jobs, large clusters, and complex query optimizations.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service