ESSEC METALAB

RESEARCH

IMPUTATION STRATEGIES FOR CLUSTERING MIXED-TYPE DATA WITH MISSING VALUES

[ARTICLE] This paper addresses the challenge of clustering incomplete mixed-type data by analyzing two procedures based on the k-prototypes algorithm. It adapts the k-POD algorithm for handling missing values and introduces a cluster aggregation strategy after multiple imputation.

by Adalbert WILHELM (ESSEC Business School),  Rabea ASCHENBRUCK, Gero SZEPANNEK

Incomplete data sets with different data types are difficult to handle, but regularly to be found in practical clustering tasks. Therefore in this paper, two procedures for clustering mixed-type data with missing values are derived and analyzed in a simulation study with respect to the factors of partition, prototypes, imputed values, and cluster assignment. Both approaches are based on the k-prototypes algorithm (an extension of k-means), which is one of the most common clustering methods for mixed-type data (i.e., numerical and categorical variables). For k-means clustering of incomplete data, the k-POD algorithm recently has been proposed, which imputes the missings with values of the associated cluster center. We derive an adaptation of the latter and additionally present a cluster aggregation strategy after multiple imputation. It turns out that even a simplified and time-saving variant of the presented method can compete with multiple imputation and subsequent pooling.

[Please read the research paper here]

Research list
arrow-right
Résumé de la politique de confidentialité

Ce site utilise des cookies afin que nous puissions vous fournir la meilleure expérience utilisateur possible. Les informations sur les cookies sont stockées dans votre navigateur et remplissent des fonctions telles que vous reconnaître lorsque vous revenez sur notre site Web et aider notre équipe à comprendre les sections du site que vous trouvez les plus intéressantes et utiles.