Fleet Response Engineer

Rhoda AIMountain View, CA

About The Position

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

Requirements

  • Strong systematic debugging of complex systems — the ability to reason from symptoms to root cause and reverse-engineer unexpected behavior.
  • Comfortable in Linux and the command line, and reading logs, telemetry, and system data.
  • Coding proficiency in Python and/or C/C++ (Bash a plus).
  • Clear written and verbal communication, and calm judgment under pressure.
  • Willingness to take part in an on-call rotation, including some nights and weekends as the fleet grows.

Nice To Haves

  • 3+ years working with robotic, autonomous, automotive, aerospace, or industrial-automation systems in an engineering, reliability, or support capacity.
  • Familiarity with ROS/ROS2 and robot subsystems (sensors, actuators, perception, networking).
  • Prior on-call, site-reliability, or release-engineering experience.
  • Experience building tools that help others debug, and a track record of solving unusual bugs.
  • Comfort with hardware and hardware–software interfaces; willingness to travel occasionally to deployment sites.

Responsibilities

  • Triage and incident response: Monitor incoming faults, alarms, and alerts from robots deployed at customer sites and in internal testing; triage and prioritize by severity and operational impact.
  • Act as the first engineering responder and point of contact for real-time escalations from field technicians and operations.
  • Debug complex issues across subsystems through log analysis, telemetry review, and reproduction testing.
  • Drive incidents to fast resolution or mitigation to restore operation and maximize fleet uptime.
  • Escalate to and coordinate with domain engineering teams (AI, robot software, cloud, hardware) to drive resolution.
  • Own the on-call rotation and paging, and keep the escalation process clear and current.
  • Communicate status to operations, engineering, and customer-facing teams throughout an incident.
  • Own each incident through to confirmed recovery and a clean handoff.
  • Perform root-cause analysis — identify contributing factors, themes, and corrective actions — and track follow-ups to closure.
  • Document investigations and fixes in a shared knowledge base and troubleshooting guide so future issues resolve faster.
  • Feed field insights back to engineering to improve product reliability and issue detection.
  • Track reliability trends and flag recurring or systemic failures.
  • Build (or spec, with the software team) the monitoring, alerting, and diagnostic tooling that catches issues fleet-wide.
  • Establish diagnostic procedures that let technicians and operators self-serve common issues.
  • Continuously reduce manual, repetitive response work through automation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service