Model Releases

Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations

arXiv:2607.15708v2 Announce Type: replace Abstract: Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors

DGX agentpaper
model-releasesarxiv-cs-ro

arXiv:2607.15708v2 Announce Type: replace Abstract: Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors that are ill-suited for GPS-denied environments. We present a learned, vision-only estimator that maps each robot's monocular image, together with messages exchanged over a communication graph, directly to its 6-DoF relative pose. Its key ingredient is the implicit virtual leader (IVL): a non-physical reference frame at the team centroid, implicitly learned inside a Transformer-based graph neural network, so that estimation has no privileged node and needs no absolute localization. The estimator additionally reports well-calibrated aleatoric (heteroscedastic GNLL) uncertainty alongside epistemic (MC~Dropout) uncertainty, compared systematically across simulation and real-world test sets. Trained only in simulation, the estimator generalizes to unseen scenes, to larger unseen team sizes, and to an external real-world benchmark. It exhibits no single point of failure: removing any one robot costs at most 1.24imes the median removal, and removing 71% of the communication links costs 1.77imes in position error without retraining. Trained on real-robot data from a single platform, it transfers without modification to a heterogeneous team, estimating relative pose to 0.22,m and 1.6^irc on physical robots, where it drives closed-loop formation control.

Source: arXiv cs.RO | 2026-08-11

Loading related sources…