AI Engineer – Robotics Data Preprocessing

Persona AI IncHouston, TX
Hybrid

About The Position

Persona AI is building humanoid robots for demanding environments in heavy industry, performing dangerous and physically demanding work. We are backed by leading investors and engaged with global industrial leaders. We are seeking a highly skilled AI Engineer to architect systems that turn raw, unstructured multimodal data into high-fidelity training assets for our robots. This role is critical as models are only as good as the data they learn from, and in humanoid robotics, creating this data is the bottleneck. You will architect and scale the infrastructure that turns raw, messy reality into training-grade data, extracting, augmenting, and aligning human dexterous manipulation data from massive multi-sensor and egocentric video datasets. You'll build advanced pre-processing algorithms to infer what sensors can't directly see, such as quantifying grasp dynamics, estimating contact forces, reconstructing occluded hand poses, and lifting 3D geometry from 2D frames. You will also maximize the value of expensive teleoperation data through spatial, temporal, and cross-modal augmentation, directly impacting how fast our models learn.

Requirements

  • M.S., or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.
  • Deep expertise in Python and extensive experience with PyTorch, specifically in handling custom dataloaders for multimodal datasets.
  • Experience analyzing and processing complex time-series data from force-torque (F/T) sensors, load cells, or tactile arrays, ensuring pristine alignment with visual frames.
  • Mastery of video processing pipelines and libraries (OpenCV, FFmpeg, Decord) and managing the I/O bottlenecks of terabyte-scale video datasets.
  • Solid working knowledge of 3D geometry and robotics data: coordinate frames and transforms, rotation representations, camera intrinsics/extrinsics, forward/inverse kinematics, URDF.
  • Proven ability to implement programmatic and generative data augmentation techniques for computer vision and time-series data.

Nice To Haves

  • Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).
  • Familiarity with the modern perception toolbox as a user: segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), 6-DoF object pose tracking, point tracking.
  • Familiarity with distributed data processing systems (Ray, Apache Spark) for cluster computing.
  • Background in generating or utilizing synthetic robotic data via simulation (Omniverse, MuJoCo).
  • Experience integrating spatial awareness or tactile data representations (e.g., Fourier encoding) into visual pipelines.

Responsibilities

  • Design cross-modal validation systems to verify agreement between video, proprioception, force/haptic signals, and language annotations.
  • Orchestrate hand-tracking, segmentation, depth estimation, 3D reconstruction, and pose-tracking modules; retarget human demonstrations into robot trajectories; and run simulation-in-the-loop validation.
  • Implement robust data augmentation strategies (spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection).
  • Develop unified state–action representations across differing embodiments, coordinate frames, rotation conventions, gripper/hand parameterizations, and sampling rates.
  • Build tooling for researchers to query, visualize, and audit datasets, and translate model-failure analyses into curation rules and re-collection requests.
  • Architect end-to-end ingestion pipelines to produce indexed, queryable, training-ready datasets from raw recordings, including temporal segmentation, metadata extraction, embedding-based retrieval, and language annotation workflows.

Benefits

  • Competitive compensation
  • Performance-based bonus
  • 99% employer covered medical benefits
  • Early-stage equity
  • Competitive PTO
  • Company-wide paid winter break between December 24th and January 2nd
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service