Legal case documents: A comprehensive dataset for Arabic natural language processing research and applications.
Journal:
Data in brief
Published Date:
Jan 2, 2026
Abstract
The legal sector remains distinctive due to the complex language structure and specialized terminology of legal data. This complexity offers considerable contextual information, which demands natural language processing (NLP). The availability of high-quality and well-structured legal datasets is essential for advancing NLP research and applications within the legal field. However, a gap exists within the Arabic legal NLP owing to insufficient research and datasets. To address this gap, we aim to propose an Arabic legal case dataset containing cases, case summaries, relevant keywords, and case categories. The legal case data were obtained from the Board of Grievances website in Saudi Arabia and include 3170 cases distributed across 47 classes. The number of words in these cases varies significantly, ranging from about 100 to nearly 30,000 words per case. Moreover, the number of pages varies, ranging from one page to 80 pages per case. Therefore, this dataset supports various NLP applications, including text categorization, data extraction, sentiment analysis, and summarization, thereby improving task efficiency and decision accuracy in the legal profession.
Authors
Keywords
No keywords available for this article.