Data Analytics Lead Engineer

CitiIrving, TX
$125,760 - $188,640Hybrid

About The Position

Citi is looking for a Data Analytics Lead Engineer to design, build, and operate scalable data pipelines and cloud-based data architectures within its Lending business, spanning Mortgage and Personal Loans. This is a hands-on data engineering role where the individual will develop and maintain production-grade data systems — working across big data platforms, data lakes, and cloud infrastructure — that directly power lending analytics at scale. The role also requires an understanding of AI and ML integration as an additional capability applied within a strong data engineering foundation.

Requirements

  • 6+ years of hands-on experience building and managing data pipelines, data warehouses, and data lake solutions using technologies such as Hadoop, Apache Spark, PySpark, Databricks, Delta Lake, Hive, Impala, and Iceberg.
  • Practical experience with cloud data platforms including Snowflake, Cloudera, used to build and automate ETL and data ingestion workflows.
  • Fluency in one or more scripting languages — Python, Scala, or Shell Scripting — applied actively to data engineering, pipeline development, and automation tasks.
  • Strong ability to design and query relational and non-relational data stores, with a clear understanding of schema design trade-offs and data modelling principles.
  • Hands-on experience with workflow scheduling tools such as Autosys or Apache Airflow to manage and orchestrate data pipeline execution.
  • Confident use of DevOps practices including version control, build tools, unit testing, monitoring, and change management to support reliable and repeatable delivery.
  • Experience with data visualization platforms such as Tableau, Cognos to support data presentation and reporting needs.
  • A Bachelor's degree or equivalent university qualification; a Master's degree is preferred.

Nice To Haves

  • Exposure to cloud-based AI and ML services such as Amazon SageMaker, Azure Machine Learning, or Google AI Platform, used to integrate predictive models within data pipelines.
  • Familiarity with NoSQL database technologies such as HBase, MongoDB, Couchbase, Cassandra, or Neo4j.
  • Databricks certification or cloud platform certification in AWS, Azure, or GCP.
  • A proactive approach to troubleshooting — able to independently investigate root causes and resolve pipeline or data issues with thoroughness and pace.

Responsibilities

  • Build, deploy, and manage end-to-end data pipelines that ingest, transform, and deliver large-scale lending datasets across Mortgage and Personal Loans with high reliability and performance.
  • Design and implement scalable data architectures on cloud platforms, selecting the right tools and approaches across data lakes, data warehouses, and streaming environments.
  • Architect and implement data schemas — choosing from relational, dimensional, normalized, or partitioned models — to meet performance, scalability, and business requirements.
  • Write and optimize complex SQL queries against large-scale datasets, applying sound decisions around distributed and parallel processing to improve pipeline efficiency.
  • Monitor, diagnose, and resolve operational and data quality issues across pipelines to ensure accuracy, completeness, and timely delivery of data.
  • Apply generative AI tools to accelerate core engineering tasks such as code generation, query optimization, and data summarization where appropriate.
  • Contribute to data engineering standards and collaborate with Business Analysts, Data Engineers, and Data Governance teams to translate business requirements into robust technical solutions.

Benefits

  • medical, dental & vision coverage
  • 401(k)
  • life, accident, and disability insurance
  • wellness programs
  • paid time off packages, including planned time off (vacation), unplanned time off (sick leave), and paid holidays.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service