Deep Learning and Large Language Models for Offline Recognition of Latin Handwritten Kazakh Text
Artykuł w czasopiśmie
MNiSW
20
Lista 2024
| Status: | |
| Autorzy: | Shormakova Assem, Mansurova Madina, Yerkegul Beibitkhan, Miłosz Marek |
| Dyscypliny: | |
| Aby zobaczyć szczegóły należy się zalogować. | |
| Rok wydania: | 2026 |
| Wersja dokumentu: | Drukowana | Elektroniczna |
| Język: | angielski |
| Numer czasopisma: | 9 |
| Wolumen/Tom: | 15 |
| Numer artykułu: | 552 |
| Strony: | 1 - 21 |
| Impact Factor: | 5,2 |
| Efekt badań statutowych | NIE |
| Finansowanie: | This research was funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24993001, Creation of a Large Language Model (LLM) to Maintain the Implementation of the Kazakh Language and Increase Technological Progress). |
| Materiał konferencyjny: | NIE |
| Publikacja OA: | TAK |
| Licencja: | |
| Sposób udostępnienia: | Witryna wydawcy |
| Wersja tekstu: | Ostateczna wersja opublikowana |
| Czas opublikowania: | W momencie opublikowania |
| Data opublikowania w OA: | 24 sierpnia 2026 |
| Abstrakty: | angielski |
| This article investigates offline recognition of handwritten Kazakh text in the Latin script using a convolutional recurrent neural network. The relevance of the study is deter- mined by the transition of the Kazakh language to the Latin alphabet and the need to automate the processing of handwritten documents. The proposed model consists of a convolutional neural network feature extractor, two bidirectional long short-term memory layers, and a Connectionist Temporal Classification decoder. The convolutional layers extract visual features from word images, the bidirectional recurrent layers model the sequential relationships between characters, and CTC enables end-to-end training without explicit character-level segmentation. A specialized dataset named KazEsim, containing 20,000 handwritten Kazakh name images, was created and divided into writer-independent training, validation, and test subsets. Experimental results showed a character accuracy rate of 96.5% and a word accuracy rate of 92.3%. Compared with a conventional CNN baseline, the proposed CRNN model improved character accuracy by 6.1 percentage points and word accuracy by 9.2 percentage points. The proposed model also outperformed the fine-tuned TrOCR-small comparative baseline while requiring fewer parameters and lower inference latency. These findings demonstrate the effectiveness of CNN–BiLSTM–CTC sequence modeling for offline recognition of handwritten Kazakh words in the Latin script. |
