Big Data Engineer

NexivaRockville, MD
Hybrid

About The Position

We are seeking a Big Data Engineer with expertise in building enterprise-scale data solutions. The role involves designing, maintaining, and optimizing big data pipelines, implementing automated testing and data quality validation, and enabling analytics and data science teams. This position requires strong problem-solving skills to diagnose and resolve issues in production environments, optimize workloads, and automate tasks. The ideal candidate will have experience with Agile methodologies, CI/CD practices, and leveraging AI-assisted tools to enhance development productivity.

Requirements

  • 5+ years building enterprise-scale data solutions using Spark, Hadoop, Hive, and Scala
  • Strong scripting skills (Python or Perl)
  • Expert-level complex SQL (window functions, multi-joins)
  • AWS cloud experience required (S3, EMR, Glue, Athena)
  • Experience with Agile delivery, CI/CD pipelines, automated testing, and GitHub workflows
  • Delivered end-to-end pipelines using Spark and Hadoop ecosystem tools
  • Optimized SQL and pipeline performance with measurable improvements
  • Deployed and supported AWS data workloads (EMR, Glue, Athena, S3)
  • Implemented CI/CD and automated testing for data pipelines
  • Used AI coding assistants and GitHub workflows in team-based development

Nice To Haves

  • Financial services or regulated industry experience preferred
  • Prompt engineering and safe use of AI coding assistants
  • Collaboration and communication in fast-paced, cross-functional environments

Responsibilities

  • Design and maintain scalable, reliable big data pipelines
  • Optimize Spark/Hadoop workloads for performance, scalability, and cost efficiency
  • Implement automated testing and data quality validation
  • Enable analytics and data science teams with high-quality, accessible datasets
  • Leverage AI-assisted tools (Copilot, ChatGPT, Q Developer) to improve development productivity
  • Diagnose and resolve Spark performance bottlenecks and data pipeline failures
  • Optimize complex SQL transformations and large-scale joins
  • Troubleshoot data quality, latency, and reliability issues in production
  • Improve AWS workload efficiency through tuning and resource optimization
  • Automate repetitive engineering tasks using AI-assisted development tools
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service