Evaluation on embeddings application for Spanish automatic text clustering

Authors

  • Anthony Wainer Cachay-Guivin Pontificia Universidad Católica del Perú

Keywords:

Model analysis, Artificial intelligence, Datasets

Abstract

The vast amount of information on the Internet, primarily composed of texts, makes clustering reliable information a complicated task. This research aims to improve the automatic clustering of Spanish texts by applying embeddings and unsupervised learning algorithms. Five datasets were used, and embedding generation techniques such as Word2Vec, FastText, Glove, BERT, and GPT-2 were applied. Models like K-means, HDBSCAN, and AutoEncoder combined with K-means were employed for clustering. The results showed that the AutoEncoder model combined with the K-means using Glove embeddings achieved superior performance with an accuracy of 0.92, NMI of 0.79, and ARI of 0.81 on the BBC News dataset. In other datasets, the results varied, but the AutoEncoder with the K-means model consistently outperformed other methods. We conclude that neural network models with AutoEncoder and K-means layer are highly effective for automatically clustering Spanish texts, especially when using high-quality embeddings like Glove.

Downloads

Download data is not yet available.

Author Biography

Anthony Wainer Cachay-Guivin, Pontificia Universidad Católica del Perú

Pontificia Universidad Católica del Perú

Escuela de Postgrado

Published

2025-01-07

How to Cite

[1]
A. W. Cachay-Guivin, “Evaluation on embeddings application for Spanish automatic text clustering”, Ingeniare, Rev. chil. ing., vol. 32, Jan. 2025.

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.