From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
arXiv:2608.08904v1 Announce Type: cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action