Data Engineer RD-017

Irth Solutions
CA$75,000Remote

About The Position

We are looking for a Data Engineer to design, build, and maintain data ingestion and processing pipelines in Databricks, transforming high-volume external data sources into clean, structured, reliable inputs for downstream analysis and intelligence. This role will primarily support our Stakeholder Engagement offering, working closely with our Data Scientist and application teams.

Requirements

  • 3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines.
  • Familiarity with data modeling, data quality, and schema evolution.
  • Solid understanding of data pipeline reliability practices: monitoring, alerting, and handling failures gracefully in a continuously running system.
  • Hands-on experience with Databricks (or an equivalent Spark-based environment): schema design, Delta Lake, performance tuning, and pipeline orchestration.
  • Experience with at least one major cloud (Azure preferred; AWS/GCP also beneficial).
  • Experience integrating with external APIs at scale: authentication, pagination, rate limiting, retries, error handling.
  • Strong proficiency in Python and SQL
  • Comfortable working with unstructured/semi-structured text data at scale.

Nice To Haves

  • LLM prompting experience and/or basic understanding of AI/NLP concepts
  • Exposure to medallion architecture or lakehouse best practices.
  • Experience with orchestration frameworks (ADF, Workflows, Airflow, DBX, etc.).
  • Experience with CI/CD tools and version control (Git, GitHub Actions or equivalent).
  • Basic understanding of security practices: RBAC, encryption, credential management.
  • Databricks certification (Data Engineer Associate or equivalent).

Responsibilities

  • Design, build, and maintain ingestion pipelines from high-volume external APIs, capable of running continuously and reliably at scale.
  • Implement ingestion and transformation workflows using Databricks (Spark/PySpark, SQL, Delta Live Tables), applying medallion architecture patterns (Bronze → Silver → Gold) to move from raw ingested content to clean, structured, analysis-ready data.
  • Build the infrastructure for deduplication and relevance filtering of incoming content, implementing filtering logic and quality criteria defined in collaboration with the Data Scientist.
  • Implement schema evolution handling and data validation rules as data sources and formats change over time.
  • Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
  • Design and evolve data schemas that balance query performance, cost, and maintainability as data volume grows.
  • Maintain clear metadata and documentation of table structures to support easy consumption by the Data Science and application teams.
  • Ensure pipeline reliability and observability: error handling, retries, monitoring, and alerting for a continuously running system.
  • Adapt pipelines to evolving external API contracts, rate limits, authentication changes, and new data sources.
  • Troubleshoot pipeline failures, perform recovery, and tune performance as needed.
  • Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar orchestration tools.
  • Contribute to CI/CD pipelines for code deployment, versioning, and environment management.
  • Work closely with the Data Scientist to expose clean, well-structured data feeding LLM/NLP pipelines and downstream models.
  • Participate in technical decisions around data architecture and propose structuring solutions as the team's needs evolve.
  • Document pipelines, data dictionaries, job schedules, and transformation logic.
  • Support the onboarding of new data sources and pipelines as the product expands to additional solution areas.

Benefits

  • Competitive Salary
  • Medical, Dental, and Vision Insurance
  • 401(k) Plan with Company Match
  • Generous Paid Time Off (PTO)
  • Company-Paid Holidays
  • Flexible Work Options
  • On-Call Compensation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service