Data Engineer opportunity in New York

Capgemini•New York, NY
•$86,129 - $127,189•Hybrid

About The Position

Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world. This is a hybrid role located in New York.

Requirements

  • 5+ years of data engineering experience, including production Databricks work: Delta Lake, Unity Catalog, Lakeflow/DLT pipelines, structured streaming, and SQL warehouses.
  • Working GCP fluency - BigQuery, Cloud Storage, Cloud Run, IAM.
  • Expert-level SQL and strong Python for transformation logic and data-quality automation.
  • Hands-on experience with at least one clean room technology, with a real understanding of aggregate versus row-level constraints.
  • Deep understanding of identity resolution: deterministic matching, probabilistic graph construction, household-level aggregation, and device graph assembly.
  • Strong PII governance knowledge in financial services: data residency, consent frameworks, GLBA, Fair Lending, and UDAAP.
  • CI/CD and infrastructure-as-code experience for data platforms.
  • Ability to operate in client-facing consulting delivery contexts - translating business requirements into technical specifications.

Responsibilities

  • Build and operate ingestion into a Delta raw landing zone on GCS using Lakeflow Connect managed connectors, Zerobus, structured streaming, and batch ETL/ELT.
  • Ingest transaction signals, behavioural events, and third-party identity graphs - LiveRamp RampID, UID2, GCLID chains, and household device graphs.
  • Own schema enforcement and evolution at the landing boundary, so upstream change surfaces as a controlled event rather than a downstream failure.
  • Build the SLM/LLM document-processing tier that converts unstructured blobs - Adobe AEM exports, asset libraries, PDFs, brand guidelines, compliance rule sets, creative - into governed Silver tables.
  • Implement parsing and OCR, entity extraction, data hygiene and deduplication, normalisation and mapping, schema inference and enforcement, and PII detection and masking as pipeline stages with measurable quality gates.
  • Treat model-based extraction as a transformation step under test, with validation and deterministic fallbacks - not as a black box.
  • Implement identity resolution at scale: deterministic matching, probabilistic graph construction, and household- and device-level cluster assembly across 1B+ data points.
  • Own feature engineering and the data-quality expectations framework - reconciliation, lineage, deduplication - reducing manual reconciliation cycles from weeks to hours.
  • Build and operate the serving layer behind the analytical model agents: MTA (7- and 28-day), MMM, incrementality, LTV:CAC, cohorts, reach/frequency, cobrand overlap, channel viability, card activation, product holdings, scenario reallocation, and channel universe.
  • Maintain clean room connectors and privacy-preserving data exchange pipelines across GMP, AMC, XMi, LiveRamp DCR or equivalent, at both aggregate and row level.
  • Serve the Gold layer the compliance agent tier reads, including the product-terms and compliance-rule data required for factual claim verification.
  • Build activation integrations to owned channels and paid media, supporting real-time audience push and match-rate monitoring.
  • Own Unity Catalog governance: access control, lineage, and tenant isolation across client organisations.
  • Implement PII governance controls at the pipeline layer - redacted ID egress, consent signal propagation, and guardrail validation aligned to GLBA, Fair Lending, UDAAP, and TCPA/CAN-SPAM.
  • Own CI/CD, infrastructure-as-code, cost management, and pipeline observability.
  • Meet the serving contract the agentic layer depends on: a fixed set of signal shapes, queried live on every conversational turn, with predictable latency and per-client isolation.

Benefits

  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave
  • Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
  • Other benefits as provided by local policy and eligibility
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service