Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability
| Source: arXiv AI
Tags: Bangla NLP, sentence classification, low-resource NLP, ensemble learning, natural language processing
Researchers introduce a 10,000-sentence annotated Bangla corpus for sentence function classification across four categories, with a Double-Level Ensemble achieving 0.95 accuracy and macro-F1—establishing the first strong baselines for this task in a language spoken by ~230 million people.
Details
Automatic sentence function identification matters for dialogue systems, text-to-speech, and machine translation, but Bangla lacks benchmark resources for this specific task. This paper introduces a 10,000-sentence corpus manually annotated across four categories (declarative, interrogative, imperative, exclamatory) with strong inter-annotator agreement (Fleiss Kappa 0.82) and near-balanced class distribution.\n\nExperiments compare BoW, TF-IDF, and Word2Vec features with classical classifiers and two heterogeneous ensemble architectures. TF-IDF consistently outperforms Word2Vec, attributed to the relatively small corpus limiting Word2Vec embedding quality. The Double-Level Ensemble with TF-IDF achieves 0.95 accuracy and macro-F1, confirmed robust through cross-validation, with LIME-based interpretability providing insight into model predictions.\n\nThe work establishes baselines for Bangla NLP but does not evaluate transformer-based models or LLMs on the corpus, which remains future work. Accepted at RAAICON 2026.