Senior Fleet Software Engineer
Software Engineering
Mountain View, CA, USA
At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.
The Role
You'll build the software and infrastructure that lets a small team operate and monitor a growing fleet of deployed robots. The observability, alerting, and automation you build are what the team watches during a shift and what on-call responds to when something goes wrong. You'll contribute to the fleet's operational software end to end, from the telemetry we collect on every robot to the dashboards, pipelines, and tooling that act on it, and work closely with the operations and response teams so the fleet gets easier to run as it scales.
What You'll Do
Observability and telemetry
Build and own fleet observability: the metrics, logs, traces, and dashboards that give the team full visibility into live robots
Design telemetry and data pipelines that reliably move robot data to the cloud for monitoring, debugging, and model training
Build fleet-health dashboards and reports that make performance and regressions easy to spot
Alerting and automation
Build alerting and on-call tooling that catches issues fast and routes them to the right responder with the right context
Automate diagnostics and incident capture so responders can start debugging instead of gathering data
Automate manual, error-prone work and reduce operational toil
Deployment and release infrastructure
Contribute to CI/CD pipelines that deliver code reliably from development to the fleet
Work with software teams to automate over-the-air (OTA) software and firmware updates across the fleet, with staged rollout, monitoring, and rollback
Build provisioning and configuration tooling to keep the fleet consistent and reproducible
Required Qualifications
3+ years of software engineering, with real ownership of internal tooling, infrastructure, or reliability systems
Strong proficiency in Python plus at least one systems language (Go, C++, or Rust)
Experience building and operating CI/CD pipelines and deployment automation
Experience with a major cloud (AWS or GCP), containers, and orchestration (Docker, Kubernetes)
Solid distributed-systems fundamentals across data transport, monitoring, and fault tolerance
Ability to debug across the full stack, from a device on the network to a service in the cloud
Preferred Qualifications
Background in robotics, autonomous vehicles, or other latency- or safety-critical domains
Experience with observability stacks (e.g., Prometheus/Grafana, OpenTelemetry, Foxglove)
Experience with OTA or fleet deployment and safe-rollout patterns such as staged rollout and auto-rollback
Familiarity with ROS/ROS2, edge or embedded Linux, or low-latency data transport for real-time systems
Experience building tooling for an operations, on-call, or field team