Data Engineer

Sand Tech Holdings LimitedBaltimore, MD
Onsite

About The Position

Sand Technologies is seeking a Data Engineer to design, build, and maintain scalable data architecture that underpins their decision-support applications. These applications range from traditional Analytics (data warehouse) to Machine Learning, Digital Twins, and serving LLMs and Agentic workflows. The role involves working closely with cross-functional teams, contributing to the strategic direction of data initiatives, and operating with a strong code-first, “data as a product” mindset where testing, reliability, observability, and performance are critical. The company partners with governments, cities, and enterprises to improve essential systems in healthcare, water, energy, telecommunications, and infrastructure, using data and AI to make these industries work better. Sand Technologies has a global presence with colleagues across Africa, Europe, the UK, and the US.

Requirements

  • 3+ years designing and operating large-scale semi-distributed data platforms (hybrid centralised and distributed) in cloud or hybrid environments.
  • Proven experience architecting modern data systems (lakehouse, data mesh, or equivalent) supporting both analytical (descriptive and predictive) and operational workloads.
  • Deep hands-on expertise with distributed processing frameworks (e.g., Spark) and streaming/event systems (e.g., Kafka or similar).
  • Strong experience building secure, governed data environments with robust access controls, encryption, lineage, and audit capabilities.
  • Experience designing secure data platforms in regulated or government environments, with strong understanding of compliance, auditability, and data protection standards.
  • Experience integrating heterogeneous data sources, including legacy systems, APIs, telemetry/IoT systems, and relational databases.
  • Demonstrated ability to design highly available, observable, production-grade data systems.
  • Experience enabling machine learning and advanced analytics through robust data infrastructure and feature pipelines.
  • Strong proficiency in Python, SQL, and ideally DBT with a track record of writing clean, production-quality code.
  • Experience deploying and operating solutions in AWS, Azure, or GCP, including CI/CD and infrastructure-as-code is beneficial.
  • Ability to operate effectively in complex, multi-stakeholder environments.
  • Strong systems-thinking mindset with a focus on scalability, modularity, and long-term platform evolution.
  • Experience designing data platforms in U.S. public sector or highly regulated environments, with working knowledge of applicable federal and state data privacy and security requirements (e.g., HIPAA, CJIS, FERPA, state-level privacy acts), and the ability to embed compliance, auditability, and data governance principles into architectural design.
  • Ability to travel to client sites in Baltimore 4 days a week minimum.
  • Drive and ethic to succeed in working in small teams physically but in larger efforts virtually.
  • Self-drive to communicate constantly using web collaboration and video conferencing is essential.

Nice To Haves

  • Experience deploying and operating solutions in AWS, Azure, or GCP, including CI/CD and infrastructure-as-code.

Responsibilities

  • Architect and build a secure, scalable urban data platform integrating multi-agency and infrastructure datasets at scale.
  • Design resilient cloud-native architectures supporting batch, streaming, and near-real-time operational workloads.
  • Lead development of high-performance ingestion and transformation pipelines across legacy systems, APIs, IoT/telemetry, and structured data sources.
  • Implement distributed and event-driven processing systems (e.g., Spark, Kafka or equivalent) for large-scale analytical and operational use cases.
  • Establish platform reliability standards, including observability, automated data quality validation, lineage, monitoring, and defined SLAs/SLOs.
  • Design and enforce strong data governance and access control frameworks, including RBAC, encryption, auditability, and secure data handling practices.
  • Build modern lakehouse or equivalent architectures that enable advanced analytics, GIS, and production-grade machine learning.
  • Partner closely with data scientists, ML engineers, and senior stakeholders to operationalize AI and analytics at scale.
  • Optimize platform performance, scalability, and cost efficiency as adoption grows.
  • Contribute to long-term architectural direction and mentor engineering team members.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service