About The Position

Stord operates the largest independent e-commerce fulfillment network in the US, with over 20 fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are building a new business line that turns this operational infrastructure into some of the most valuable training data assets in physical AI. We are looking for an experienced computer vision engineer and technologist to build and scale this business from the ground up.

Requirements

  • Experience standing up and scaling an egocentric perception stack, having built and run a similar product end to end at a robotics or AI data company, driving the full lifecycle: hardware setup, embedded perception, data pipelines, ensuring quality, and delivering it to production teams who depend on it.
  • 8+ years building and shipping production computer-vision/perception systems (or an MS/PhD in CV, ML, or robotics plus 6+ years hands-on), including systems that ran on messy real-world data, not just benchmarks.
  • Deep expertise in computer vision and tooling — track record of leveraging existing tooling and designing, training, and debugging CNNs and vision transformers from scratch.
  • Strong command of geometric computer vision: camera calibration, depth estimation, and 2D/3D pose estimation.
  • End-to-end ownership of a major perception problem: from data and model design through evaluation, optimization, and deployment, with measurable accuracy and reliability outcomes.
  • Track record of setting technical direction for a team or large workstream and raising the bar for other engineers.
  • Proven ability to take ambiguous, 0→1 problems with no established playbook and drive them to a working system with limited resources.
  • Experience with large unstructured datasets (video/multimodal) and the eval discipline to instrument accuracy rather than eyeball it.
  • Expert Python and strong software-engineering fundamentals; C++ where performance demands it.

Responsibilities

  • Define and deliver the product, owning the data product across quality tiers — from RGB egocentric video through depth-enhanced and full multimodal capture with hand pose and annotations. Decide what gets built, in what order, based on what buyers will actually pay for, and hold the line on quality.
  • Run the capture and delivery program, standing up the warehouse capture operation: camera and rig hardware selection, enrollment, edge processing, and the processing pipelines that package datasets for delivery. Coordinate across warehouse operations, engineering, and customers to ship datasets on spec and on schedule.
  • Build the perception stack, including detection, tracking, and segmentation, plus depth/3D reconstruction and 6DoF, multi-view 3D hand/body pose estimation from egocentric and fixed-camera capture.
  • Stand up VLM-assisted and automated labeling with human-in-the-loop QA to drive down cost per annotated hour, and integrate the annotation tooling.
  • Own the hardware<>vision intersection, including camera calibration, epipolar/multi-view geometry, and frame-accurate time-sync across multi-camera and egocentric rigs; derive 3D pose by triangulation where no direct sensor exists.
  • Train and ship models by designing, fine-tuning, and optimizing CV/multimodal models on large unstructured video datasets, and get them reproducible and production-ready, not stuck in a notebook.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service