Mixed Script Identification Using Automated DNN Hyperparameter Optimization.

Journal: Computational intelligence and neuroscience

Published Date: Dec 10, 2021

Abstract

Mixed script identification is a hindrance for automated natural language processing systems. Mixing cursive scripts of different languages is a challenge because NLP methods like POS tagging and word sense disambiguation suffer from noisy text. This study tackles the challenge of mixed script identification for mixed-code dataset consisting of Roman Urdu, Hindi, Saraiki, Bengali, and English. The language identification model is trained using word vectorization and RNN variants. Moreover, through experimental investigation, different architectures are optimized for the task associated with Long Short-Term Memory (LSTM), Bidirectional LSTM, Gated Recurrent Unit (GRU), and Bidirectional Gated Recurrent Unit (Bi-GRU). Experimentation achieved the highest accuracy of 90.17 for Bi-GRU, applying learned word class features along with embedding with GloVe. Moreover, this study addresses the issues related to multilingual environments, such as Roman words merged with English characters, generative spellings, and phonetic typing.

Authors

Muhammad Yasir

College of Oceanography and Space Informatics, China University of Petroleum, Qingdao, China.
Li Chen

Department of Endocrinology and Metabolism, Qilu Hospital, Shandong University, Jinan, China.
Amna Khatoon

Department of Information Engineering, Chang'an University, Xi'an, Shaanxi, China.
Muhammad Amir Malik

Department of Computer Science, Islamic International University, Islamabad, Pakistan.
Fazeel Abid

School of Information Science and Technology, Northwest University, Xi'an 710127, China.

Keywords

Natural Language Processing Research Design

External Resources

View on PubMed Access via DOI PubMed (34925496)

Mixed Script Identification Using Automated DNN Hyperparameter Optimization.

Abstract

Authors

Keywords

External Resources

Popular Topics

Recent Journals