Evaluating early fusion and transformer-based models for sign language recognition using manual and non-manual features
Artykuł w czasopiśmie
MNiSW
20
Lista 2024
| Status: | |
| Autorzy: | Amangeldy Nurzada, Yerimbetova Aigerim, Miłosz Marek, Gazizova Nazerke, Tursynova Nazira, Kassymova Akmaral |
| Dyscypliny: | |
| Aby zobaczyć szczegóły należy się zalogować. | |
| Rok wydania: | 2026 |
| Wersja dokumentu: | Drukowana | Elektroniczna |
| Język: | angielski |
| Wolumen/Tom: | 9 |
| Numer artykułu: | 200527 |
| Strony: | 1 - 19 |
| Impact Factor: | 3,6 |
| Web of Science® Times Cited: | 0 |
| Scopus® Cytowania: | 0 |
| Bazy: | Web of Science | Scopus |
| Efekt badań statutowych | NIE |
| Finansowanie: | This research was funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24992875). |
| Materiał konferencyjny: | NIE |
| Publikacja OA: | TAK |
| Licencja: | |
| Sposób udostępnienia: | Witryna wydawcy |
| Wersja tekstu: | Ostateczna wersja opublikowana |
| Czas opublikowania: | W momencie opublikowania |
| Data opublikowania w OA: | 4 lipca 2026 |
| Abstrakty: | angielski |
| The paper proposes a hybrid multimodal architecture for automatic gesture recognition, which implements separate processing of manual and non-manual features, followed by their subsequent merging. To evaluate the effectiveness of the proposed approach, a comparative analysis of four architectures was conducted, which revealed that models with early feature merging and parallel processing of modalities based on LSTM demonstrated the highest accuracy (up to 0.99). Architectures with attention and Transformer achieved an accuracy of 0.97, confirming high efficiency when working with time sequences. The scientific novelty of the work lies in the comparison of several neural network architectures for the task of gesture recognition, the implementation of separate processing for manual (hand movements) and non-manual (lip articulation) components, and their subsequent merging, as well as the construction of an adaptive architecture capable of effectively recognizing gestures. Unlike most existing multimodal systems, which require specialized equipment (such as depth cameras and inertial sensors), the presented approach utilizes only a video stream, ensuring ease of implementation, flexibility in configuration, and compatibility with mass devices. The developed architecture can serve as a basis for building inexpensive, scalable systems for automatic sign language interpretation and human–machine interaction, including those in real-time. |
