Research
Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces
arXiv:2608.25876v1 Announce Type: cross Abstract: Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as 'more elegant' or 'more minimalis
arXiv:2608.25876v1 Announce Type: cross Abstract: Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.
Related
- Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
- Consistency Regularised Gradient Flows for Inverse Problems
- Linguistic Context Recodes Visual Representations in Vision-Language Models
Source: arXiv cs.CV | 2026-08-27