About The Position

Stord is building a new business line that transforms its real-world operational infrastructure into high-value training data for the next generation of physical AI. We are looking for an experienced Computer Vision Engineer and technologist to help build and scale this business from the ground up. You will work at the intersection of computer vision, robotics, data, and warehouse operations, turning real-world environments and workflows into high-quality datasets that enable smarter, more capable AI systems. This is an opportunity to build a new physical AI data business from the ground up, with access to an operating environment that would be difficult to replicate anywhere else. You will have a structural data advantage, direct access to real-world warehouse environments, workflows, and human activity at significant scale. You will be building data products for companies developing the next generation of robotics and physical AI in a rapidly growing market. This role offers true 0→1 ownership, allowing you to shape the product, technology, team, and operating model from the beginning. You will work closely with the CTO & Co-Founder to define the technical and commercial direction of the business.

Requirements

  • 8+ years building and shipping production computer vision/perception systems, or an MS/PhD in computer vision, ML, or robotics with 6+ years of hands-on industry experience.
  • Experience building and scaling an egocentric perception or video data stack end to end, ideally within robotics, physical AI, or an AI data company.
  • Deep expertise in computer vision, including detection, tracking, segmentation, depth, 2D/3D pose estimation, and vision transformers.
  • Strong command of geometric computer vision, including camera calibration, multi-view geometry, synchronization, and 3D reconstruction.
  • Proven ownership of complex perception problems from data and model design through evaluation, optimization, and deployment, with measurable improvements in accuracy and reliability.
  • Experience working with large, unstructured video and multimodal datasets, with a disciplined approach to evaluation and quality measurement.
  • A track record of setting technical direction and raising the bar for other engineers.
  • Proven ability to take ambiguous 0→1 problems from concept to working system with limited resources and no established playbook.
  • Expert-level Python and strong software engineering fundamentals; C++ experience where performance demands it.
  • Comfortable moving between hardware, data, models, infrastructure, and operations to make the system work.

Responsibilities

  • Own the early egocentric video and perception stack—from data collection and camera rigs through vision models, processing pipelines, and dataset delivery.
  • Define and build the data product, owning the product across quality tiers—from RGB egocentric video to depth-enhanced and multimodal capture with hand pose, body pose, and annotations.
  • Prioritize what gets built based on customer demand and hold a high bar for data quality.
  • Stand up the capture operation, owning camera and rig selection, hardware setup, enrollment, edge processing, data ingestion, and the pipelines that turn raw capture into production-ready datasets.
  • Partner closely with warehouse operations, engineering, and customers to deliver on spec and on schedule.
  • Build the perception stack, developing detection, tracking, segmentation, depth/3D reconstruction, 6DoF, and multi-view 3D hand/body pose estimation across egocentric and fixed-camera systems.
  • Automate labeling at scale by building VLM-assisted and automated labeling workflows with human-in-the-loop QA, reducing the cost of annotation while maintaining rigorous quality standards.
  • Own the hardware–vision intersection, driving camera calibration, epipolar and multi-view geometry, frame-accurate synchronization, and 3D pose triangulation across multi-camera and egocentric rigs.
  • Train and ship models by designing, fine-tuning, optimizing, and deploying computer vision and multimodal models against large, unstructured video datasets.
  • Build reproducible systems that perform reliably in production.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service