Parts of Speech Tagging for Handwritten Sindhi Sentence using Deep Learning Models

Authors

  • Marya Soomro 1Department of Information Technology & Computer Science Sindh University Campus, Naushahro Feroze https://orcid.org/0009-0001-4445-1482 (unauthenticated)
  • Muhammad Ahsan Raza Mughal 1Department of Information Technology & Computer Science Sindh University Campus, Naushahro Feroze
  • Muhammad Khalid Sheikh 2Department of Computer Science, National University of Computer and Emerging Sciences (FAST-NUCES) Karachi,
  • Azhar Ali Shah Department of Information Technology & Computer Science Sindh University Campus, Naushahro Feroze

Keywords:

POS Tagging, Handwritten Sindh, Named Entity Recognition, YOLOv8, Handwriting Recognition, Machine learning

Abstract

Part-of-Speech (POS) tagging is an essential task in Natural Language Processing (NLP) that helps identify the role of each word in a sentence. For handwritten low-resource languages, this task becomes more difficult because of the lack of available datasets and differences in individual writing styles. In this study, a deep learning-based approach is developed for POS tagging of handwritten Sindhi sentences. For this purpose, a dataset of handwritten Sindhi sentences was collected and manually labeled with corresponding POS categories. The prepared dataset was then used for training and evaluating the proposed models. The study applies and compares two deep learning models, Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations from Transformers (BERT), for automatic POS tagging. The models were evaluated using accuracy, precision, recall, and F1-score metrics. Experimental results demonstrate that BERT achieved better performance compared to LSTM due to its ability to capture contextual information from sentence structures. BERT achieved an accuracy of 87.3% while LSTM achieved 85.1%. The proposed work contributes to Sindhi language processing by providing a handwritten dataset and a baseline deep learning approach for POS tagging of a low-resource language.

Author Biography

  • Marya Soomro, 1Department of Information Technology & Computer Science Sindh University Campus, Naushahro Feroze

    Maria Soomro is a graduate researcher in the Department of IT & Computer Science, University of Sindh, Pakistan. Her research focuses on NLP, POS tagging, and deep learning techniques for handwritten Sindhi text.

References

Adnan A. Memon, Saman Hina, Abdul K. Kazi, Saad Ahmed. 2024. Parts-of-speech tagger for Sindhi language using deep neural network architecture. MUET Research Journal 43, 3 (July 2024), 47-55. https://doi.org/10.22581/muet1982.2768

Mutee R. Arain. 2015. Towards Sindhi Corpus Construction. Linguistics and Literature Review 1, 1 (March 2015), 39-48. https://doi.org/10.32350/llr/11/04

Erum N. Sodhar, Hina Bhanbro, Zira H. Amur, Akktar H. Jalbani, Abdul H. Buller. 2020. Sindhi Language Processing on Online SindhiNLP Tool. University of Information and Communication Technology 4, 3 (October 2020), 139-142. https://sujo.usindh.edu.pk/index.php/USJICT/ article/view/2879

Vijay Rowtula and Praveen Krishnan. 2018. POS Tagging and Named Entity Recognition on Handwritten Documents. In Proceedings of the 15th International Conference on Natural Language Processing. NLP Association of India. 82-86. https://aclanthology.org/2018.icon-1. 11/

Farooq Zaman, Onaiza Maqbool, and Jaweria Kanwal. 2024. Leveraging Bidirectional LSTM with CRFs for Pashto Tagging. ACM Trans. Asian Low Resour. Lang. Inf. Process. 23, 4, Article 58 (April 2024), 17 pages. https://doi.org/10.1145/3649456

Wazir Ali, Zenglin Xu, and Jay Kumar. 2021. SiPOS: A Benchmark Dataset for Sindhi Partof-Speech Tagging. In Proceedings of the Student Research Workshop Associated with RANLP 2021, pages 22–30, Online. https://aclanthology.org/2021.ranlp-srw.4/

Nawaz, Ali Nawaz, Muhammad Shaikh, Noor Rajper, Samina Baber, Junaid Khalid, Muhammad. (2023). TPTS: Text Pre-processing Techniques for Sindhi Language. Pakistan Journal of Emerging Science and Technologies (PJEST). 4. 1-12. 10.58619/pjest.v4i3.89

Sagor Sarker (2021), BNLP: Natural language processing toolkit for Bengali language, CoRR, abs/2102.00405. https://arxiv.org/abs/2102.00405

Awwalu, Jamilu Abdullahi, Saleh Evwiekpaefe, Abraham. (2020). PARTS OF SPEECH TAGGING: A REVIEW OF TECHNIQUES. FUDMA Journal of Sciences. 4. 712-721. 10.33003/ fjs-2020-0402-325

Ali, Asghar. (2020). Deep learning-based isolated handwritten Sindhi character recognition. Indian Journal of Science and Technology. 13. 2565-2574. 10.17485/IJST/v13i25.914

Ali, Irfan Ali, Insaf Guriro, Subhash Khan, Asif Raza, Syed Qureshi, Basit Bhatti, Priha. (2019). Sindhi Handwritten-Digits Recognition Using Machine Learning Techniques. International Journal of Computer Network and Information Security. VOL.19. pp. 195-201.

Khan, Wahab Daud, Ali Khan, Khairullah Nasir, Jamal Basheri, Mohammed Aljohani, Naif Alotaibi, Fahd. (2019). Part of Speech Tagging in Urdu: Comparison of Machine and Deep Learning Approaches. IEEE Access. PP. 1-1. 10.1109/ACCESS.2019.2897327

A. A. Kalhoro, R. Husain, and H. Shaikh, “Deep Learning for Sugarcane Disease Detection: A Field-Validated EfficientNet-B4 Approach,” University of Sindh Journal of Information and Communication Technology, vol. 9, no. 1, pp. 28–38, 2025.

A. A. Abro, A. A. Khan, M. S. H. Talpur, I. Kayijuka, and E. Ya?ar, “Machine Learning Classifiers: A Brief Primer,” University of Sindh Journal of Information and Communication Technology, vol. 5, no. 2, pp. 63–68, 2021.

Downloads

Published

2026-07-30

How to Cite

Parts of Speech Tagging for Handwritten Sindhi Sentence using Deep Learning Models. (2026). University of Sindh Journal of Information and Communication Technology , 9(2), 88-94. https://sujo.usindh.edu.pk/index.php/USJICT/article/view/7942

Similar Articles

11-20 of 69

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)