Deep learning algorithms for speech interaction analysis
DOI:
https://doi.org/10.64966/ingeniare.v33.32Keywords:
Speaker diarization, deep learning, voice separation, collaborative environments, speech analysisAbstract
This article offers a comparative assessment of deep learning algorithms for automatic analysis of verbal interactions, drawing upon evidence from n recent scientific literature. Four widely used algorithms –LSTM, CNN, GRU, and x-vectors– are analyzed across key dimensions relevant to speech analytics in collaborative contexts: speaker diarization accuracy (DER), ability to handle overlapping speech (SDR), robustness to noise as indicated by error recognition (WER), and computational efficiency assessed via inference time. Based on previously published empirical results, this synthesis evaluates the strengths, limitations, and optimal application scenarios for each algorithm, with a focus on educational and collaborative settings. The analysis shows that LSTMs excel in temporal modeling, CNNs achieve superior performance in voice separation, GRUs reach a favorable balance between accuracy and efficiency, and x-vectors enable highly scalable speaker identification. This study establishes a comprehensive framework for selecting and integrating deep learning techniques into speech analysis systems that demonstrate robustness and context-sensitivity across educational, professional, and research applications.
Downloads
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Héctor Cornide-Reyes, José Jorquera, Diego Monsalves

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors retain copyright of their work and grant the journal the right of first publication under the Creative Commons CC-BY Attribution License, which permits unrestricted use, distribution, and reproduction provided the original authorship and the journal’s first publication are acknowledged.


