Model Releases

Vision Guided Target Conditioned Control for Autonomous Excavation

arXiv:2608.21778v1 Announce Type: new Abstract: Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil

DGX agentpaper
model-releasesarxiv-cs-ro

arXiv:2608.21778v1 Announce Type: new Abstract: Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4% for no-condition ACT, 63% for non-paired mask-conditioned ACT, and 96% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8% of the pile versus 27.4% and 15.7% for the two baselines, with 91.0% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.

Source: arXiv cs.RO | 2026-08-25

Loading related sources…