Research
A Taxonomy of Construction Task Activities for Robot Workers
arXiv:2608.25395v1 Announce Type: new Abstract: Recent vision-language-action models offer a path toward robots with broader repertoires than conventional task-specific systems. Construction deploymen
arXiv:2608.25395v1 Announce Type: new Abstract: Recent vision-language-action models offer a path toward robots with broader repertoires than conventional task-specific systems. Construction deployment, however, requires a precise inventory of worker activities and the capabilities needed to execute them. We present TARCAT, an occupation-grounded taxonomy derived from 91 O*NET tasks across seven high-employment construction occupations and 30 instructional videos of physical work. TARCAT defines 41 action primitives in 12 groups and three classes and provides a mechanism for composing parameterized primitive sequences into reusable skills. This human-interpretable structure can organize demonstrations, specify robot requirements, and support coding agents that retrieve and extend skill libraries. We also demonstrate selected primitives on a DOBOT CR3 arm with a CRAFT hand. TARCAT thereby provides a common vocabulary for analyzing human work and developing general-purpose construction robots. Annotations are available at https://github.com/AICPS/TARCAT-Taxonomy.
Related
- RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
- Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
Source: arXiv cs.RO | 2026-08-27