Senior Software Engineer, Robot Data Infrastructure

GRAMSan Francisco, CA
$190,000 - $240,000Onsite

About The Position

GRAM is a self-replication company creating machine labor for the physical economy. Our first research frontier is self-preservation: the base case of physical self-replication. We are building a new class of machines called insectoids that can survive, coordinate, and recover without humans. We believe scalable machine labor requires more than single-agent task generality or machines shaped in our image. You will build the infrastructure that turns physical operation into reproducible training and evaluation datasets. Your scope begins at the capture contract and spans multimodal ingestion, temporal alignment, provenance, quality controls, storage, dataset construction, replay, and reliable access for training and evaluation. This is a senior software engineering role responsible for the systems after capture and the contracts that keep recorded experience compatible with downstream use. Success means a model behavior can be traced through its dataset, run, software, calibration, commands, interventions, outcomes, and hardware state—and that dataset revisions remain reproducible rather than becoming ungoverned data volume.

Requirements

  • Bachelor's degree in computer science, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
  • Strong Python programming ability and working proficiency in C++, Rust, Java, or Go.
  • Experience operating object-storage and streaming or batch-processing systems against defined throughput, freshness, data-integrity, or reliability targets, including schema evolution, orchestration, and safe backfills.
  • Direct experience preserving provenance and temporal relationships across video, time-series, sensor, event, or other multimodal data.
  • Demonstrated ownership of a production pipeline incident that caused silent data loss, corrupt data, or an unavailable downstream dataset, including detection, root cause, recovery or backfill, and a test or monitor that prevented recurrence.

Nice To Haves

  • Robotics, autonomous vehicles, fleet telemetry, teleoperation, or another embodied-data environment.
  • Training-data systems, multimodal alignment, active-learning queues, dataset observability, or reproducible replay.
  • Edge capture in bandwidth-constrained or intermittently connected environments.

Responsibilities

  • Build reliable ingestion and processing systems for synchronized video, sensor, command, capture-device state, intervention, outcome, and machine-health data.
  • Define versioned schemas and lineage connecting every run to its robot configuration, calibration, software, model, operator protocol, and experiment.
  • Develop automated checks for time drift, missing streams, corruption, calibration faults, schema breaks, weak coverage, and silent data loss.
  • Build systems for indexing, filtering, sampling, curation, dataset versioning, replay, and delivery into training and evaluation workflows.
  • Instrument the edge-to-training path with service-level metrics, traceability, failure isolation, and safe reprocessing.
  • Design storage and compute paths that support large multimodal records without losing reproducibility or making iteration dependent on manual recovery.
  • Use training and evaluation evidence to revise capture contracts, quality thresholds, dataset composition, and retention policy.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service