Knowledge discovery for higher education student retention based on data mining: Machine learning algorithms and case study in Chile

Palacios Rojas, Carlos; Reyes-Suárez, José A.; Bearzotti, Lorena A.; Leiva, Víctor; Marchant-Fuentes, Carolina

Mostrar el registro sencillo de la publicación

dc.contributor.author	Palacios Rojas, Carlos
dc.contributor.author	Reyes-Suárez, José A.
dc.contributor.author	Bearzotti, Lorena A.
dc.contributor.author	Leiva, Víctor
dc.contributor.author	Marchant-Fuentes, Carolina
dc.date.accessioned	2021-12-21T13:11:59Z
dc.date.available	2021-12-21T13:11:59Z
dc.date.issued	2021
dc.identifier.uri	http://repositorio.ucm.cl/handle/ucm/3649
dc.description.abstract	Data mining is employed to extract useful information and to detect patterns from often large data sets, closely related to knowledge discovery in databases and data science. In this investigation, we formulate models based on machine learning algorithms to extract relevant information predicting student retention at various levels, using higher education data and specifying the relevant variables involved in the modeling. Then, we utilize this information to help the process of knowledge discovery. We predict student retention at each of three levels during their first, second, and third years of study, obtaining models with an accuracy that exceeds 80% in all scenarios. These models allow us to adequately predict the level when dropout occurs. Among the machine learning algorithms used in this work are: decision trees, k-nearest neighbors, logistic regression, naive Bayes, random forest, and support vector machines, of which the random forest technique performs the best. We detect that secondary educational score and the community poverty index are important predictive variables, which have not been previously reported in educational studies of this type. The dropout assessment at various levels reported here is valid for higher education institutions around the world with similar conditions to the Chilean case, where dropout rates affect the efficiency of such institutions. Having the ability to predict dropout based on student’s data enables these institutions to take preventative measures, avoiding the dropouts. In the case study, balancing the majority and minority classes improves the performance of the algorithms.	es_CL
dc.language.iso	en	es_CL
dc.rights	Atribución-NoComercial-SinDerivadas 3.0 Chile	*
dc.rights.uri	http://creativecommons.org/licenses/by-nc-nd/3.0/cl/	*
dc.source	Entropy, 23(4), 485	es_CL
dc.subject	Data analytics	es_CL
dc.subject	Databases	es_CL
dc.subject	Data science	es_CL
dc.subject	Friedman test	es_CL
dc.subject	Socioeconomic index	es_CL
dc.subject	University dropout	es_CL
dc.title	Knowledge discovery for higher education student retention based on data mining: Machine learning algorithms and case study in Chile	es_CL
dc.type	Article	es_CL
dc.ucm.facultad	Facultad de Ciencias de la Ingeniería	es_CL
dc.ucm.indexacion	Scopus	es_CL
dc.ucm.indexacion	Isi	es_CL
dc.ucm.uri	www.mdpi.com/1099-4300/23/4/485	es_CL
dc.ucm.doi	doi.org/10.3390/e23040485	es_CL