Abstract
Sentiment analysis in low-resource languages such as Urdu poses unique challenges due to limited annotated data, morphological complexity, and significant class imbalance in most publicly available datasets. This study addresses these issues through two experimental strategies. First, we explore class imbalance mitigation by using instruction-tuned large language models (LLMs) to generate synthetic negative sentiment samples in Urdu. This augmentation strategy results in a more balanced dataset, which significantly improves the recall and F1-score for minority class predictions when fine-tuned using a multilingual BERT model. Second, we investigate the effectiveness of translating Urdu text into English and applying sentiment classification through a pre-trained English language model. Comparative evaluation reveals that the translation-based pipeline, using a RoBERTa model fine-tuned for English sentiment classification, achieves superior performance across major metrics. Our results suggest that LLM-based augmentation and cross-lingual transfer via translation both serve as viable approaches to overcome data scarcity and performance limitations in sentiment analysis for low-resource languages. The findings highlight the potential applicability of these approaches to other under-resourced linguistic domains.
| Original language | English |
|---|---|
| Title of host publication | EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026 |
| Editors | Mo El-Haj, Mo El-Haj, Paul Rayson, Mustafa Jarrar, Ignatius Ezeani, Saad Ezzini, Sina Ahmadi, Amal Haddad Haddad, Cynthia Amol, Ahmad Abdelali, Shadi Abudalfa |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 198-207 |
| Number of pages | 10 |
| ISBN (Electronic) | 9798891763616 |
| DOIs | |
| State | Published - 2026 |
| Event | 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026 - Rabat, Morocco Duration: 28 Mar 2026 → … |
Publication series
| Name | EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026 |
|---|
Conference
| Conference | 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026 |
|---|---|
| Country/Territory | Morocco |
| City | Rabat |
| Period | 28/03/26 → … |
Bibliographical note
Publisher Copyright:© 2026 Association for Computational Linguistics.
Keywords
- Urdu sentiment analysis
- cross-lingual transfer
- data augmentation
- large language models
- machine translation
ASJC Scopus subject areas
- Artificial Intelligence
- Linguistics and Language
- Computer Science Applications
- Signal Processing
Fingerprint
Dive into the research topics of 'Enhancing Urdu Sentiment Classification through Instruction-Tuned LLMs and Cross-Lingual Transfer'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver