A complexity measure for binary classification problems based on lost points

Lancho, Carmen; Martín de Diego, Isaac; Cuesta, Marina; Aceña, Víctor; M. Moguerza, Javier

doi:10.1007/978-3-030-91608-4_14

Lancho, Carmen; Martín de Diego, Isaac; Cuesta, Marina; Aceña, Víctor; M. Moguerza, Javier

URI: https://hdl.handle.net/10115/39274

DOI: 10.1007/978-3-030-91608-4_14

Fecha: 2021

Resumen

Complexity measures are focused on exploring and capturing the complexity of a data set. In this paper, the Lost points (LP) complexity measure is proposed. It is obtained by applying k-means in a recursive and hierarchical way and it provides both the data set and the instance perspective. On the instance level, the LP measure gives a probability value for each point informing about the dominance of its class in its neighborhood. On the data set level, it estimates the proportion of lost points, referring to those points that are expected to be misclassified since they lie in areas where its class is not dominant. The proposed measure shows easily interpretable results competitive with measures from state-of-art. In addition, it provides probabilistic information useful to highlight the boundary decision on classification problems.

Mostrar el registro completo del ítem

Colecciones