Databricks Engineer

RevStar
Remote

About The Position

Reports To: Data & AI Practice Lead Location: Remote (US-Based / Eastern or Central Time Zone Preferred) Employment Type: Contract Ready to build greenfield Lakehouse solutions at the bleeding edge of AI and big data? RevStar is an innovation shop and official Databricks Partner launching a dedicated, cloud-agnostic Data, ML, and AI practice. We are seeking a high-caliber Databricks Engineer to build, optimize, and deploy high-performance data pipelines and MLOps frameworks for enterprise clients across AWS, Azure, and GCP. In this role, you will work directly with data architects, scientists, and client leaders to turn complex data into scalable, production-ready AI models. Above all, the ideal candidate embodies RevStar’s core values: Self-Mastery : We hold a high bar for how we think, communicate, and improve. Ownership : We own outcomes, not just effort. Shared Destiny : We rise or fall together. Your Impact Pillars Your technical contributions are organized into four strategic pillars: 1. Lakehouse Architecture & Pipeline Engineering Design, build, and optimize scalable ETL/ELT pipelines using Apache Spark and Delta Lake across multi-cloud environments (AWS S3, Azure Data Lake, GCS). Implement robust Lakehouse architectures that seamlessly process both structured and unstructured data at enterprise scale. Automate data ingestion and storage workflows to support downstream analytics and real-time operational reporting. 2. Performance Optimization & Automation Fine-tune Spark jobs for low latency, high throughput, and maximum cloud cost-efficiency. Implement CI/CD pipelines and Infrastructure-as-Code (Terraform, Databricks CLI) for automated deployments. Build automated monitoring, alerting, and data quality validation frameworks to guarantee pipeline reliability. 3. MLOps & AI Integration Partner with ML engineers and data scientists to build production-grade feature engineering pipelines. Support model training, tracking, versioning, and deployment inside Databricks using MLflow. Operationalize AI/ML models into secure, scalable production environments for client applications. 4. Data Governance & Client Excellence Enforce enterprise data security, access controls, and compliance standards (GDPR, HIPAA, SOC 2). Establish best practices for data lineage, metadata management, and operational documentation. Collaborate with client-facing stakeholders to align technical implementations with critical business outcomes.

Requirements

  • 3+ years of hands-on experience in data engineering, with a focus on big data processing and cloud-native architectures.
  • 2+ years of hands-on experience with Databricks, including Apache Spark, Delta Lake, and MLflow.
  • Databricks Certified Data Engineer Associate (or higher)
  • Proficiency in Python, SQL, and Spark-based frameworks.
  • Experience in developing and optimizing large-scale ETL/ELT pipelines.
  • Strong understanding of Lakehouse architecture and cloud-agnostic data solutions.
  • Familiarity with CI/CD pipelines and Infrastructure-as-Code (IaC) for Databricks (e.g., Terraform, Databricks CLI).
  • Knowledge of data governance, security, and compliance best practices.
  • Experience working in Agile development environments, following DevOps/MLOps best practices.

Nice To Haves

  • Additional Databricks Certifications (e.g., Databricks Certified Machine Learning Associate).
  • Experience with real-time streaming solutions (e.g., Kafka, Kinesis, Event Hub).
  • Familiarity with cloud storage and orchestration tools (e.g., Apache Airflow, Prefect).
  • Background in AI/ML integration within Databricks, assisting in feature engineering and model deployment.
  • Experience working in client-facing roles or consulting environments.

Responsibilities

  • Design, build, and optimize scalable ETL/ELT pipelines using Apache Spark and Delta Lake across multi-cloud environments (AWS S3, Azure Data Lake, GCS).
  • Implement robust Lakehouse architectures that seamlessly process both structured and unstructured data at enterprise scale.
  • Automate data ingestion and storage workflows to support downstream analytics and real-time operational reporting.
  • Fine-tune Spark jobs for low latency, high throughput, and maximum cloud cost-efficiency.
  • Implement CI/CD pipelines and Infrastructure-as-Code (Terraform, Databricks CLI) for automated deployments.
  • Build automated monitoring, alerting, and data quality validation frameworks to guarantee pipeline reliability.
  • Partner with ML engineers and data scientists to build production-grade feature engineering pipelines.
  • Support model training, tracking, versioning, and deployment inside Databricks using MLflow.
  • Operationalize AI/ML models into secure, scalable production environments for client applications.
  • Enforce enterprise data security, access controls, and compliance standards (GDPR, HIPAA, SOC 2).
  • Establish best practices for data lineage, metadata management, and operational documentation.
  • Collaborate with client-facing stakeholders to align technical implementations with critical business outcomes.

Benefits

  • Paid Time Off
  • Remote-First Working Environment
  • Comprehensive Health Coverage – Medical, Dental, Vision
  • 401(k) Retirement Plan
  • Annual Learning & Development Stipend
  • Peer Mentorship & Coaching
  • Professional Growth Opportunities
  • Company Outings & Volunteer Opportunities
  • Collaborative, Innovative Culture
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service