Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
DGX agentarXiv:2605.18177v1 Announce Type: new Abstract: Query-based Vision Transformer segmentation models typically reconstruct dense spatial feature maps to predict masks, inheriting design patterns from co