Model Releases
Revisiting Shape and Texture Reliance with Category-Separability-Calibrated Suppression
arXiv:2607.16298v2 Announce Type: replace Abstract: Feature-suppression evaluations infer model reliance on shape or texture from the accuracy loss caused by attenuating each type of information. Such
arXiv:2607.16298v2 Announce Type: replace Abstract: Feature-suppression evaluations infer model reliance on shape or texture from the accuracy loss caused by attenuating each type of information. Such losses, however, conflate feature reliance with the amount of category-relevant information removed by the corresponding transformation. Because shape and texture are suppressed using different operators, their effects are not directly comparable. We introduce the Semantic Degradation Index (SDI), which quantifies the suppression-induced reduction in category separability relative to clean images in a fixed clean-reference discriminative space constructed from handcrafted features. On an ImageNet16-like benchmark, we use SDI to compare Gaussian blur for texture suppression with grid distortion for shape suppression over their overlapping degradation range. At comparable SDI values, all five evaluated ImageNet-trained convolutional neural networks (CNNs) retain less accuracy under Gaussian blur than under grid distortion. This results supports stronger texture than shape reliance under the evaluated operators, contrasting with the shape-dominant conclusion obtained from unmatched suppression conditions. The evaluated Vision Transformers (ViTs) also generally retain more accuracy than CNNs under both operators. To determine whether this advantage extends beyond classification, We evaluate fixed brain-encoding models using clean and suppressed images from the Natural Scenes Dataset. Under both operators, ViT features show smaller suppression-induced decreases in noise-ceiling-normalized explained variance than CNN features. These findings establish category separability as an important reference for interpreting suppression-based feature reliance and show that the CNN-ViT robustness difference extends to model representations predictive of human visual cortical responses.
Related
- TwistNet-2D: Learning Second-Order Channel Interactions via Spiral Twisting for Texture Recognition
- Shape: A Self-Supervised 3D Geometry Foundation Model for Industrial CAD Analysis
- REVNET: Rotation-Equivariant Point Cloud Completion via Vector Neuron Anchor Transformer
Source: arXiv cs.CV | 2026-08-17