AI Research Engineer: Vision AI / VLM / Physical AI-1

Centific
$180,000 - $240,000Hybrid

About The Position

Centific is seeking passionate AI Research Engineers to join its cutting-edge labs, focusing on Vision AI, Multimodal Large Models (VLMs), and Physical AI. The role involves translating cutting-edge research into production systems that perceive, reason, and act in the real world. You will own high-leverage experiments from paper to prototype to deployable module within Centific's platform. This includes advancing visual perception through model building and fine-tuning for various tasks, training and evaluating VLMs for grounding and reasoning, prototyping perception-in-the-loop policies for physical AI, curating datasets, and developing robust evaluation protocols. You will also package research into reliable services on a modern stack and orchestrate multi-agent pipelines for complex tasks. Example problems include long-horizon video understanding, 3D scene grounding, privacy-preserving on-device perception, and robust multi-modal evaluation. The role offers the opportunity to work with state-of-the-art tools and frameworks in areas like 3D reconstruction, embodied AI, and robotics simulation.

Requirements

  • Masters/Ph.D in CS/EE/Robotics (or related), actively publishing in CV/ML/Robotics (e.g., CVPR/ICCV/ECCV, NeurIPS/ICML/ICLR, CoRL/RSS).
  • Strong PyTorch (or JAX) and Python; comfort with CUDA profiling and mixed precision training.
  • Demonstrated research in computer vision and at least one of: VLMs (e.g., LLaVA style, video-language models), embodied/physical AI, 3D perception.
  • Proven ability to move from paper → code → ablation → result with rigorous experiment tracking.

Nice To Haves

  • Experience with video models (e.g., TimeSFormer/MViT/VideoMAE), diffusion or 3D GS/NeRF pipelines, or SLAM/scene reconstruction.
  • Prior work on multimodal grounding (referring expressions, spatial language, affordances) or temporal reasoning.
  • Familiarity with ROS2, DeepStream/TAO, or edge inference optimizations (TensorRT, ONNX).
  • Scalable training: Ray, distributed data loaders, sharded checkpoints.
  • Strong software craft: testing, linting, profiling, containers, and reproducibility.
  • Public code artifacts (GitHub) and first-author publications or strong open-source impact.

Responsibilities

  • Build and fine-tune models for detection, tracking, segmentation (2D/3D), pose & activity recognition, and scene understanding (incl. 360° and multi-view).
  • Train/evaluate vision–language models (VLMs) for grounding, dense captioning, temporal QA, and tool use; design retrieval-augmented and agentic loops for perception-action tasks.
  • Prototype perception-in-the-loop policies that close the gap from pixels to actions (simulation + real data). Integrate with planners and task graphs for manipulation, navigation, or safety workflows.
  • Curate datasets, author high-signal evaluation protocols/KPIs, and run ablations that make results irreproducible impossible.
  • Package research into reliable services on a modern stack (Kubernetes, Docker, Ray, FastAPI), with profiling, telemetry, and CI for reproducible science.
  • Orchestrate multi-agent pipelines (e.g., LangGraph-style graphs) that combine perception, reasoning, simulation, and code generation to self-check and self-correct.

Benefits

  • Salary: $180,000-$240,000 + OTB
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service