Staff AI Engineer, Applied AI — Smart Vision

Arlo TechnologiesMilpitas, CA

About The Position

About Arlo: At Arlo, we're passionate about creating innovative and reliable solutions that help people protect what matters most to them. Our team is dedicated to delivering products that exceed our customers' expectations, while always pushing the boundaries of what's possible in the world of protection technology. We believe that everyone deserves to feel safe and secure, whether they're at home or away, and we're committed to providing our customers with the peace of mind they need to live their lives without worry. Arlo’s deep expertise in AI- and CV-powered analytics, cloud services, user experience, product design, and innovative wireless and RF connectivity enables the delivery of a seamless, smart security experience for Arlo users that is easy to set up and interact with every day. Smart Vision is the AI team behind Arlo's intelligence layer: object and person detection, animal/vehicle/package recognition, custom-trained detections, video captioning and scene description, and natural-language search over a user's video library. Our models run across the edge, the cloud, and third-party foundation models, and they process events from millions of cameras every day. About the role As a Staff AI Engineer for Applied AI, you'll be a technical owner of the models behind Arlo's smart features — from computer vision detectors running on-camera to vision-language models that describe what happened, to the retrieval and agent layers that let customers ask questions about their video. You'll pick the right approach for each problem (train, fine-tune, prompt, or retrieve), prove it with solid evals, and take it all the way to production at consumer scale. This is a hands-on applied role: you ship models, not papers.

Requirements

  • BS in Computer Science or a related technical field with 8+ years of experience; MS/PhD in ML, CV, or a related field preferred (or equivalent practical experience).
  • 8+ years building production ML/AI systems, with a track record of owning models end to end — problem framing, data, training, evaluation, deployment, iteration.
  • Strong computer vision depth: detection, classification, segmentation, tracking, video understanding; you've trained and shipped CV models in a real product.
  • 3+ years working with LLMs or VLMs in production — multimodal modeling, prompt and context design, fine-tuning, and evaluation.
  • Strong Python + PyTorch; solid distributed systems and cloud fundamentals (AWS, Docker).
  • Rigorous about evaluation and data quality — you build the benchmark before you build the model.
  • Experience with embeddings and vector search at scale; multimodal or video retrieval strongly preferred.
  • Experience shipping LLM agents with tool use, and pragmatic judgment about when not to use an agent.
  • Working knowledge of inference optimization (serving stacks, quantization, batching, GPU performance) and cost/latency tradeoffs at scale.
  • Demonstrated technical leadership without formal authority: influencing roadmaps, mentoring engineers, driving cross-team decisions.
  • Bias to ship, comfort with ambiguity, strong written communication.

Nice To Haves

  • Edge/on-device inference (quantization-aware training, pruning, NPU/DSP toolchains) and streaming inference.
  • Video processing at scale (decoding, frame sampling, FFmpeg-class tooling).
  • Audio or sensor-fusion models complementing video; privacy-preserving or on-device personalization.
  • Open-source contributions to CV, serving, retrieval, or agent frameworks; consumer IoT or camera/security domain experience.

Responsibilities

  • Build, train, and fine-tune computer vision models for detection, classification, tracking, re-identification, and video understanding, and improve them against real-world customer footage — night, weather, motion blur, odd camera angles, edge compute limits.
  • Own our video-understanding pipeline built on vision-language models: frame selection and temporal context, prompt and output-schema design, grounding and hallucination control, multi-event reasoning, and quality tuning for captioning and scene description.
  • Adapt models to our domain: SFT, LoRA/QLoRA, preference tuning, distillation into small deployable models, and knowing when a 200M-parameter specialist beats a frontier model.
  • Own the data and evaluation loop — dataset curation, labeling strategy, hard-negative and failure mining, active learning, benchmark suites, and offline/online metrics that reliably predict customer-perceived quality.
  • Own the embedding and retrieval stack behind video search: multimodal/video embeddings, vector index design and tuning, hybrid search and re-ranking, and natural-language queries over a user's library.
  • Build agentic experiences on top of the vision stack: tool/function calling, multi-step reasoning over event history, RAG and memory, guardrails, and tracing/observability for agent runs.
  • Keep production inference fast and economical — serving stack choice and tuning (vLLM/TensorRT-LLM/Triton), quantization, batching, GPU utilization, and routing between hosted foundation models and self-hosted open models against clear cost and latency targets.
  • Partner with product, data, backend, and firmware/edge teams to turn ambiguous product ideas into shipped AI features, with safe rollout (canary, A/B, feature flags) and real production telemetry.
  • Raise the engineering bar — architecture and design reviews, MLOps practices, documentation, and mentoring senior and mid-level engineers.

Benefits

  • We provide reasonable accommodations to applicants and employees with disabilities, who are pregnant or have a related medical condition, or who have sincerely held religious beliefs, observances, and practices.
  • Pursuant to applicable state and municipal Fair Chance Laws and Ordinances, the Company will consider for employment qualified applicants with arrest and conviction records.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service