A 48-class handwritten dataset for the endangered chakma language.

Journal: Data in brief
Published Date:

Abstract

This article describes a comprehensive handwritten character dataset for the endangered Chakma language, primarily spoken in the Chittagong Hill Tracts of Bangladesh. The dataset comprises 37,708 processed RGB images, standardized to 40 × 40 pixels, covering 48 distinct classes that include 38 letters and 10 numerals. Data collection involved manual input from a diverse demographic of native speakers in the Rangamati and Khagrachhari districts, utilizing standardized form to capture authentic stylistic variability. The raw handwritten samples were subsequently digitized using high-resolution scanning, followed by automated cropping and resizing scripts to generate a uniform, machine-learning-ready format. This open-access resource addresses the scarcity of digital tools for indigenous scripts and can be utilized for training Handwritten Character Recognition (HCR) models, developing synthetic fonts via generative networks, and conducting linguistic analysis of Chakma handwriting characteristics.

Authors

Keywords

No keywords available for this article.