Research
Scalable unsupervised feature selection via weight stability
arXiv:2506.06114v4 Announce Type: replace Abstract: Unsupervised feature selection is critical for improving clustering performance in high-dimensional data, where irrelevant features can obscure mean
arXiv:2506.06114v4 Announce Type: replace Abstract: Unsupervised feature selection is critical for improving clustering performance in high-dimensional data, where irrelevant features can obscure meaningful structure. In this work, we introduce the Minkowski weighted k-means++, a novel initialisation strategy for the Minkowski Weighted k-means. Our initialisation selects centroids probabilistically using feature relevance estimates derived from the data itself. Building on this, we propose two new feature selection algorithms, FS-MWK++, which aggregates feature weights across a range of Minkowski exponents to identify stable and informative features, and SFS-MWK++, a scalable variant based on subsampling. We support our approach with a theoretical analysis, demonstrating that, under explicit assumptions on noise features and cluster structure, relevant features are assigned consistently higher weights than noise features across a range of Minkowski exponents. Our software can be found at https://github.com/xzhang4-ops1/FSMWK.
Related
- Robust gene prioritization for Dietary Restriction via Fast-mRMR Feature Selection techniques
- TreeGrad-Ranker: Feature Ranking via O(L)-Time Gradients for Decision Trees
- The Wasserstein transform
- Does Dimensionality Reduction via Random Projections Preserve Landscape Features?
- Distributionally Robust K-Means Clustering
Source: arXiv cs.LG | 2026-04-16