A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
arXiv:2607.24148v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deploy