Research
Mutual information and sensitivity analysis for feature selection in customer targeting: a comparative study
arXiv:2608.20447v1 Announce Type: new Abstract: Feature selection is a highly relevant task in a data-driven knowledge discovery project. Several techniques have been developed aiming at finding the f
arXiv:2608.20447v1 Announce Type: new Abstract: Feature selection is a highly relevant task in a data-driven knowledge discovery project. Several techniques have been developed aiming at finding the features that influence most an outcome to predict, including mutual information and, in recent years, the data-based sensitivity analysis. The present research focus on analyzing the advantages and disadvantages of each of these two techniques, by applying both to a bank telemarketing case. Thereafter, a logistic regression model is built on the tuned set of features identified by each of the two techniques as the most influencing set of features on the success of a telemarketing contact, in a total of 13 features for mutual information and 9 features for the data-based sensitivity analysis. The latter performs better for lower values of false positives while the former is slightly better for a higher false positive ratio. Thus, mutual information becomes a better choice if bank managers intend to reduce slightly the cost of contacts without risking losing a high number of successes. Such results show that mutual information, although not recent, is still a valid method for feature selection. On the other side, the data-based sensitivity analysis selection achieved good prediction results with less features.
Related
- An Empirical Study of Feature Selection Granularity
- MIST: Mutual Information Estimation Via Supervised Training
- Fourier Preconditioning for Neural Feature Learning
- FSEVAL: Feature Selection Evaluation Toolbox and Dashboard
- ML-PWS: Estimating the Mutual Information Between Experimental Time Series Using Neural Networks
Source: arXiv cs.LG | 2026-08-24