Volume 10,Issue 7
Feature selection is essential for dimensionality reduction on big data, but it faces considerable challenges when applied to high-dimensional and sparse datasets. To address these challenges, this paper proposes Unconstrained Latent Factorization-based Improved Relief-F (ULF-IR), a novel feature selection method tailored for such complex scenarios. The method integrates two main components: (1) a double factorization (DF)-based unconstrained latent factor model is employed to accurately reconstruct missing data without relying on pre-imputation or strict non-negativity constraints; (2) an improved Relief-F (IRelief-F) algorithm assigns reliable importance weights to features, effectively differentiating among highly similar features even in the presence of noise introduced during imputation. Comprehensive experiments on three real-world datasets show that ULF-IR consistently surpasses state-of-the-art methods in both classification accuracy and robustness, demonstrating its effectiveness as a dependable solution for feature selection on high-dimensional, incomplete data.