Senior Data Engineer

SteerBridgeVienna, VA
$135,000 - $160,000Hybrid

About The Position

SteerBridge is seeking a highly skilled and motivated Senior Data Engineer to join our Modern Disability Claims AI/ML program. This role involves aligning data solutions with business requirements by planning and managing data infrastructure and strategy. The team is focused on using AI/ML to increase claims processing throughput and reduce adjudication wait times for veterans. The Senior Data Engineer will perform Data Engineering tasks within existing systems of record with multiple databases, enhancing and optimizing data entry, management, and extraction. Key activities include data quality checks, analysis, data presentation, and process documentation. The ideal candidate is a quick learner, curious, innovative, results-oriented, and possesses strong interpersonal skills.

Requirements

  • Must be a U.S. Citizen.
  • Bachelor's Degree or Above in Systems Engineering, Computer Science, or related field.
  • Must hold, or be able to obtain, a Public Trust clearance (an active Secret or Top Secret clearance also satisfies this requirement).
  • Minimum 6+ years of experience.
  • Experience in data pipelines, utilizing advanced analytics tools and platforms and Python.
  • Experience in scripting, tooling, and automating large-scale computing environments.
  • Extensive experience with major tools such as Python, Pandas, PySpark, NumPy, SciPy, SQL, and Git.
  • Minor experience with TensorFlow, PyTorch, and Scikit-learn.
  • Advanced data modeling (conceptual, logical, and physical) with emphasis on scalability and maintainability.
  • Strong understanding of database paradigms (relational, NoSQL, graph, time-series, and document-based).
  • Expertise with modern data warehousing platforms (Redshift, Snowflake, BigQuery).
  • Deep understanding of dimensional modeling (star/snowflake schemas) and data vault techniques.
  • Experience designing for both OLTP and OLAP workloads.
  • Proficiency with schema evolution, metadata-driven pipelines, and data versioning strategies.
  • Implementing data retention, archival, and lifecycle policies.
  • Hands-on experience with distributed processing tools (Apache Kafka, Airflow, Spark, Flink, NiFi).
  • Skilled in building and orchestrating batch and real-time pipelines on cloud platforms (AWS Glue, GCP Dataflow, Azure Data Factory).
  • Deep understanding of incremental processing, idempotency, schema evolution, and backfill logic.
  • Proficient in pipeline automation, observability, and monitoring (metrics, logging, alerting).
  • Strong Python development for ETL — modular, testable, reusable, and performance-optimized.
  • Knowledge of workflow dependency management, retries, and failure recovery strategies.
  • Deep expertise in AWS, GCP, or Azure data ecosystems.
  • Experience building and managing cloud-native data solutions (Data Lakes, Data Warehouses, Data Mesh).
  • Strong understanding of cloud storage (S3, Blob), managed databases (RDS, DynamoDB), and compute (EMR, Dataproc, ECS).
  • Cost governance and performance optimization for large-scale data workloads.
  • Knowledge of serverless data patterns (AWS Lambda + Athena, GCF + BigQuery).
  • Experience with hybrid/multi-cloud architecture and inter-cloud data movement.
  • Hands-on experience with distributed computing frameworks (Hadoop, Spark, Hive, Presto).
  • Proficiency with data lake and lakehouse architectures (Delta Lake, Apache Iceberg, Apache Hudi).
  • Understanding of partitioning, data compaction, schema evolution, and ACID compliance.
  • Strong knowledge of query optimization on massive datasets (Athena, Trino, Presto).
  • Performance tuning in petabyte-scale distributed systems.
  • Advanced SQL/NoSQL query tuning, indexing, sharding, and partitioning strategies.
  • Proficient with replication, backups, and disaster recovery across distributed systems.
  • Skilled in analyzing query execution plans and applying cost-based optimization.
  • Experience optimizing data-intensive application code and database interfaces.
  • Familiarity with temporal tables, data versioning, and caching strategies.
  • Implementing data privacy and compliance frameworks (GDPR, CCPA).
  • Experience with data cataloging, lineage, and metadata management (DataHub, Collibra, Alation).
  • Role-based access control and sensitive data protection across multi-tenant systems.
  • Integration of data quality validation and data contract testing within CI/CD pipelines.
  • Automation of governance and security policies using cloud-native tools.
  • Strong proficiency in Python and SQL for data processing, automation, and API integration.
  • Expertise in object-oriented programming (OOP) and design patterns in Python.
  • Deep understanding of algorithmic complexity (Big O) and code performance optimization.
  • Familiarity with parallel and distributed computing frameworks (Spark, Dask, Ray).
  • Skilled with version control (Git) and CI/CD tools (GitLab, Jenkins, CircleCI).
  • Proficient in software engineering best practices: testing (pytest/unittest), documentation, type hinting, and linting.
  • Ability to debug, profile, and optimize large-scale data workflows.
  • Collaboration with data scientists on feature engineering, data preparation, and model deployment.
  • Knowledge of ML orchestration and experiment tracking (MLflow, Kubeflow).
  • Familiarity with feature stores and data lineage for ML.
  • Integration of batch and streaming data pipelines for real-time inference.
  • Hands-on experience with analytics and visualization tools (Tableau, Power BI).
  • Mentoring and guiding junior engineers in data design, coding standards, and performance optimization.
  • Leading cross-functional projects with data scientists, analysts, and business partners.
  • Promoting best practices for data engineering and governance within the organization.
  • Effective stakeholder communication, documentation, and Agile project management.
  • Ability to conduct technical reviews and enforce design and scalability standards.

Nice To Haves

  • DevOps/DataOps: Infrastructure as code, Docker/Kubernetes, automated deployment of data infrastructure.
  • Testing & CI/CD: Git-based workflows, automated integration testing, and continuous delivery for data pipelines.
  • Performance & Cost Optimization: Tuning query execution, pipeline efficiency, and resource utilization.
  • Automation: Building self-healing data pipelines with retry logic, monitoring, and alerting.
  • Documentation: Strong communication of technical architecture using tools like Lucidchart, PlantUML, or Draw.io.

Responsibilities

  • Perform Data Engineering tasks within existing systems of record with multiple databases.
  • Enhance and optimize data entry, management, and extraction within databases for usability within a proprietary system.
  • Conduct data quality checks and analysis.
  • Present data and document data processes.
  • Develop and maintain data pipelines, utilizing advanced analytics tools and platforms.
  • Script, tool, and automate large-scale computing environments.
  • Perform advanced data modeling (conceptual, logical, and physical) with emphasis on scalability and maintainability.
  • Build and orchestrate batch and real-time pipelines on cloud platforms.
  • Implement data retention, archival, and lifecycle policies.
  • Build and manage cloud-native data solutions (Data Lakes, Data Warehouses, Data Mesh).
  • Optimize query execution, pipeline efficiency, and resource utilization.
  • Implement data privacy and compliance frameworks.
  • Collaborate with data scientists on feature engineering, data preparation, and model deployment.
  • Mentor and guide junior engineers in data design, coding standards, and performance optimization.
  • Lead cross-functional projects with data scientists, analysts, and business partners.
  • Promote best practices for data engineering and governance within the organization.
  • Communicate technical architecture and lead technical reviews.

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Life Insurance
  • 401(k) Retirement Plan with matching
  • Paid Time Off
  • Paid Federal Holidays
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service