PolCSBD: A contextual dataset of code-mixed Bengali and Banglish for political counter-speech detection.

Journal: Data in brief
Published Date:

Abstract

The detection of counter-speech on social media is heavily dependent on the conversational context, particularly in low-resource and code-mixed languages. This article presents a dataset of 10,021 contextual comment pairs extracted from political discussions on Bangladeshi YouTube channels. The dataset captures native Bengali script, fully Romanized Bengali ("Banglish"), and hybrid code-mixed sentences. Each instance consists of a parent comment and a corresponding nested reply, annotated manually by native speakers to indicate the presence or absence of counter-speech. Replies that actively dispute, correct, or challenge the parent text with a logical counter-narrative are labeled as '1' (Counter-Speech), while replies indicating agreement, off-topic remarks, or non-argumentative insults are labeled as '0' (Non-Counter Speech). The dataset underwent rigorous preprocessing, including the removal of URLs, user tags, and emojis, alongside the normalization of English characters. This dataset provides a foundational resource for training and evaluating natural language processing (NLP) models in understanding contextual political discourse in South Asian code-mixed environments. By preserving the direct link between a baseline statement and its response, the data allows researchers to move beyond isolated sentence classification and study actual dialogic structures. It specifically targets the acute shortage of resources for analyzing informal, phonetically typed regional dialects. Consequently, machine learning practitioners and social scientists can utilize this collection to benchmark text classifiers, test large language models (LLMs) on South Asian discourse, and explore algorithmic moderation strategies in highly polarized online spaces.

Authors

Keywords

No keywords available for this article.