Graph Data Engineer

Redhorse CorporationArlington, VA
$140,000 - $170,000

About The Position

We are seeking an analytical, forward-thinking Graph Data Engineer to design, build, scale, and maintain the Enterprise Semantic Map — our ontology-grounded metadata graph. In this role, you will move the enterprise beyond traditional, static cataloging by leading an automation-first approach. You will architect and deliver programmatic data and API integrations, design and configure graph database structures, and build the agentic workflows that discover and catalog disparate data sources across the enterprise. Partnering with graph, data, and engineering teams, you will align these assets to enterprise semantic and provenance layers so data is discoverable, understandable, trusted, and dynamically composable for human analysts, applications, and downstream AI agents. Success in this role requires strong hands-on engineering skills and a systems-thinking mindset: the ability to reason about how data pipelines and tool integrations affect the broader enterprise architecture, search and discovery, and downstream agentic research workflows — and to make and defend design decisions that others will build on.

Requirements

  • Bachelor’s Degree with 5+ of relevant professional experience or equivalent.
  • Active TS SCI Clearance.
  • Foundational proficiency across the following areas, demonstrated in any comparable technology: Programming and scripting for automation (e.g., Python, Java, or a comparable general-purpose language)
  • Relational database querying (e.g., SQL)
  • Structured and semi-structured data formats (e.g., JSON, XML, YAML)
  • Graph query languages for retrieval, validation, and manipulation (e.g., Cypher for property graphs, SPARQL for RDF/triple stores)
  • Knowledge graph concepts, including nodes, edges, relationships, and metadata schemas
  • Hands-on experience building and operating production data pipelines or ETL (Extract, Transform, Load) processes, including error handling, monitoring, and scheduling.
  • Practical experience with at least one enterprise graph database platform, including schema design and query performance considerations.
  • Experience integrating heterogeneous systems through APIs across legacy, cloud, and distributed environments.
  • Ability to reason about how individual pipelines and modeling choices propagate through a broader enterprise ecosystem, and to weigh trade-offs explicitly.
  • Precision in aligning metadata terms, formatting data endpoints, and maintaining technical schemas.
  • Ability to explain semantic and architectural decisions to both engineering peers and non-technical mission stakeholders, and to document them durably.

Nice To Haves

  • Working experience with formal ontology or semantic web standards (e.g., RDF, OWL, SHACL) and with established government- or defense-related semantic models.
  • Experience with LLM orchestration, retrieval-augmented generation, or agentic workflows, particularly where a graph provides grounding.
  • Applied experience with open lineage specifications or metadata management frameworks.
  • Experience with metadata catalog environments and data stewardship systems.
  • Experience with pipeline scheduling and orchestration tooling.
  • Familiarity with cloud data platforms, containerized deployment, and CI/CD practices.
  • Prior experience supporting defense, intelligence community, or other regulated enterprise data environments.

Responsibilities

  • Design, build, and deploy automated pipelines that programmatically discover enterprise data assets and interface with existing data catalogs.
  • Scan, catalog, and ingest technical metadata — including schemas, tables, columns, and API endpoints — from legacy, cloud, and distributed environments to establish baseline assets for alignment to the Enterprise Core Ontology.
  • Establish the automated pipelines and orchestrated workflows that ingest metadata at scale, replacing manual, field-by-field mapping.
  • Own the practices that keep the ontology current as a dynamic, living “semantic control plane” rather than a static document.
  • Define how the technical origin of ingested data is captured at the point of ingestion, creating the foundation for automated provenance chains that track where data originated and how it changes over time.
  • Lead the alignment of discovered data elements from local systems to the shared Enterprise Core Ontology and specialized Domain Ontologies, with particular attention to compatibility with established institutional frameworks (e.g., DIA’s DIKEM).
  • Preserve local naming conventions while establishing standardized, shared meaning, and resolve modeling conflicts as they arise.
  • Design and maintain data lineage chains within the Provenance Layer, applying industry lineage standards to document where data originates, how it is transformed, and who governs it.
  • Write, optimize, and review graph queries supporting metadata retrieval, logical validation, and graph manipulation.
  • Establish reusable query patterns and validation checks the wider team can build on.
  • Assess how newly integrated data sources and automated pipelines affect the broader Enterprise Semantic Map, selected use cases, downstream consumers, and enterprise search and discovery — and adjust the design accordingly.
  • Connect data assets to relevant mission metadata so technical capabilities can be clearly linked to the mission workflows they support.
  • Ensure enterprise assets are associated with appropriate governance metadata, including ownership, classifications, handling rules, and access constraints.
  • Translate complex data policies into machine-readable semantic structures.
  • Maintain and optimize the Enterprise Semantic Map within enterprise graph database platforms so human analysts, applications, and autonomous AI agents can efficiently search, navigate, and discover resources.
  • Tune schema and query performance as the graph grows.
  • Partner with AI engineers so planning, research, and tool agents can dynamically query the graph, and help define the grounded, trustworthy reasoning and retrieval strategies those agents depend on.
  • Guide junior engineers on graph modeling, query construction, and pipeline development, and review their work.
  • Document schema decisions, modeling rationale, and runbooks so the design is reproducible, and represent technical positions clearly to architects, program leadership, and government stakeholders.

Benefits

  • Comprehensive benefits programs
  • Performance-based or other incentive compensation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service