Database Engineer

Bigbear.aiMcLean, VA
Remote

About The Position

BigBear.ai is hiring Data Engineers to build and maintain the source adapters and normalization logic that translate raw data from disparate systems into a common risk-signal schema. This position focuses on reliable ingestion and transformation—turning heterogeneous legacy inputs (APIs, feeds, databases, files, and event streams) into consistent, high-quality signals that downstream scoring and adjudication workflows can trust. This position is remote but may require travel in the DMV area.

Requirements

  • Must maintain an active Top Secret security clearance
  • Bachelor's Degree and 8 to 10 years of experience; Master's Degree and 6 to 8 years of experience
  • 3–5 years of experience in data engineering, including building production-grade ingestion and transformation pipelines.
  • Strong experience with API integrations and ETL/ELT development in complex environments.
  • Experience integrating heterogeneous and/or legacy systems with inconsistent schemas and data quality.
  • Experience with REST/API frameworks and building maintainable, well-tested integration services.

Nice To Haves

  • Engineering discipline: writes maintainable, testable code and builds robust pipelines that handle edge cases.
  • Curiosity and persistence: digs into messy source data and drives it to consistent outcomes.
  • Collaboration: works effectively across data architecture, scoring/analytics, and application teams.
  • Operational mindset: builds pipelines that are observable, debuggable, and supportable in production.
  • Solid SQL skills and working familiarity with NoSQL data stores.
  • Proficiency in Python or Java for building data services and transformation logic.
  • Hands-on experience producing/consuming events in Kafka (producers/consumers) or an equivalent event streaming platform

Responsibilities

  • Build source adapters/connectors to ingest data from APIs, legacy systems, databases, and event streams
  • Develop normalization and mapping logic to translate source-specific fields into the common risk-signal schema (including validation, enrichment, and standardization)
  • Implement ETL/ELT pipelines with strong engineering rigor: testing, observability, error handling, retries, and backfills
  • Produce and consume streaming events (e.g., Kafka topics) to support near-real-time signal delivery and downstream processing
  • Partner with data architecture and domain SMEs to define and maintain data contracts, mappings, and lineage from source to normalized signal
  • Ensure data quality and consistency (deduplication patterns, schema evolution handling, and reconciliation against source systems)
  • Optimize pipeline performance and reliability (throughput, latency, and scalable processing patterns)
  • Create and maintain technical documentation for adapters, transformations, and operational runbooks
  • Some travel may be required within the DMV area
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service