Staff Data Engineer

Cantor FitzgeraldDallas, TX
Hybrid

About The Position

This role involves owning and driving the technical architecture for complex, cross-team data initiatives. The engineer will design, build, and maintain scalable data pipelines and platforms in a cloud-native environment. A key focus will be on architecting and leading enterprise Master Data Management (MDM) solutions and integrating agentic AI and LLM-driven workflows. The position also includes building data foundations for machine learning and AI, designing backend data services and APIs, setting engineering standards, establishing data governance, and mentoring other engineers. The role requires partnering with leadership to translate business strategy into data roadmaps and identifying/resolving systemic issues. Maintaining awareness of emerging technologies and industry trends is also crucial.

Requirements

  • Bachelor's degree in Computer Science, Engineering, MIS, or related field preferred.
  • 12+ years of experience in data engineering or software engineering, with demonstrated experience architecting data platforms and pipelines at scale.
  • Expert-level SQL and strong proficiency in Python (Scala or Java a plus) for large-scale data processing and transformation.
  • Deep experience with cloud data platforms (e.g., Databricks, Snowflake, Synapse, BigQuery, Redshift) and cloud-native architecture patterns.
  • Deep understanding of distributed systems, data modeling (dimensional, data vault, lakehouse), and ETL/ELT architecture.
  • Hands-on experience designing and implementing Master Data Management (MDM) solutions, including entity resolution, match/merge, golden records, and reference/hierarchy management (e.g., Informatica, Reltio, Profisee, or similar).
  • Hands-on experience building or integrating agentic AI systems, LLM-powered applications, RAG pipelines, or AI agent orchestration frameworks (e.g., LangChain, AutoGen, Semantic Kernel, MCP).
  • Experience building backend data services and APIs (REST/GraphQL), with comfort working across the full stack.
  • Strong background with both relational (SQL) and NoSQL data stores, plus data lake/lakehouse formats (Delta, Iceberg, Parquet).
  • Deep understanding of CI/CD pipelines, infrastructure as code, and DevOps/DataOps practices.
  • Proven track record of leading large-scale technical initiatives across multiple teams.
  • Demonstrated ability to mentor engineers and influence technical direction without direct reporting authority.

Nice To Haves

  • Experience with data governance, lineage, and cataloging tools (e.g., Unity Catalog, Microsoft Purview, Collibra, Alation).
  • Experience designing multi-agent systems, tool-calling architectures, or retrieval-augmented generation (RAG) pipelines.
  • Experience with event-driven architectures and streaming/real-time data processing (e.g., Kafka, Event Hubs, Kinesis, Flink, Spark Structured Streaming).
  • Experience building the data layer for ML/AI, including feature stores, vector databases, embeddings, and ML/LLMOps.
  • Familiarity with containerization and orchestration (Docker, Kubernetes) and workflow orchestration (Airflow, Dagster, dbt).
  • Prior experience in commercial real estate, fintech, or operations/transaction systems.
  • Track record of speaking, writing, or open-source contributions that demonstrate technical thought leadership, especially in applied AI or data.

Responsibilities

  • Own and drive the technical architecture for complex, cross-team data initiatives spanning ingestion, transformation, storage, and serving layers.
  • Design, build, and maintain scalable, high-performance data pipelines and distributed data platforms in a cloud-native environment (Azure, AWS, or GCP).
  • Architect and lead enterprise Master Data Management (MDM), including golden records, entity resolution, data domains, reference and hierarchy management, and stewardship, to create trusted, authoritative data across the business.
  • Architect and integrate agentic AI and LLM-driven workflows (autonomous agents, RAG pipelines, AI copilots) into data platforms and pipelines to drive efficiency and new capabilities.
  • Build and support the data foundations for machine learning and AI, including feature stores, vector stores, embeddings, and ML/LLMOps pipelines.
  • Design and deliver backend data services and APIs (REST/GraphQL), and contribute across the stack to expose curated datasets to applications, analytics, and BI consumers.
  • Set engineering standards and best practices for data quality, modeling, testing, observability, and deployment across the organization, including responsible use of AI-assisted development tools.
  • Establish data governance, lineage, cataloging, and quality frameworks across the data estate.
  • Lead technical design reviews and provide architectural guidance to multiple engineering and data teams.
  • Partner with product, analytics, and engineering leadership to translate business strategy into scalable data roadmaps, including AI-driven capabilities.
  • Identify and resolve systemic performance, reliability, and scalability issues across the data stack.
  • Mentor and coach senior and mid-level engineers, raising the technical bar across the organization on data engineering, MDM, and AI practices.
  • Drive adoption of modern frameworks, tools, and engineering practices, including agentic AI and LLM tooling, to improve delivery velocity and platform resilience.
  • Maintain awareness of emerging technologies and industry trends, particularly in agentic AI, master data management, and modern data platforms, and assess their applicability to the business.

Benefits

  • Competitive compensation
  • Growth opportunities
  • Access to world-class engineering, data, and AI resources
  • Collaborative Culture
  • Access to world-class learning resources and mentorship
  • Flexible working hours
  • Hybrid options
  • Comprehensive health, dental and vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service