Evaluation on embeddings application for Spanish automatic text clustering
Keywords:
Model analysis, Artificial intelligence, DatasetsAbstract
The vast amount of information on the Internet, primarily composed of texts, makes clustering reliable information a complicated task. This research aims to improve the automatic clustering of Spanish texts by applying embeddings and unsupervised learning algorithms. Five datasets were used, and embedding generation techniques such as Word2Vec, FastText, Glove, BERT, and GPT-2 were applied. Models like K-means, HDBSCAN, and AutoEncoder combined with K-means were employed for clustering. The results showed that the AutoEncoder model combined with the K-means using Glove embeddings achieved superior performance with an accuracy of 0.92, NMI of 0.79, and ARI of 0.81 on the BBC News dataset. In other datasets, the results varied, but the AutoEncoder with the K-means model consistently outperformed other methods. We conclude that neural network models with AutoEncoder and K-means layer are highly effective for automatically clustering Spanish texts, especially when using high-quality embeddings like Glove.
Downloads
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2024 Anthony Wainer Cachay-Guivin

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors retain copyright of their work and grant the journal the right of first publication under the Creative Commons CC-BY Attribution License, which permits unrestricted use, distribution, and reproduction provided the original authorship and the journal’s first publication are acknowledged.


