Skip to main navigation Skip to search Skip to main content

Enhancing Urdu Sentiment Classification through Instruction-Tuned LLMs and Cross-Lingual Transfer

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Sentiment analysis in low-resource languages such as Urdu poses unique challenges due to limited annotated data, morphological complexity, and significant class imbalance in most publicly available datasets. This study addresses these issues through two experimental strategies. First, we explore class imbalance mitigation by using instruction-tuned large language models (LLMs) to generate synthetic negative sentiment samples in Urdu. This augmentation strategy results in a more balanced dataset, which significantly improves the recall and F1-score for minority class predictions when fine-tuned using a multilingual BERT model. Second, we investigate the effectiveness of translating Urdu text into English and applying sentiment classification through a pre-trained English language model. Comparative evaluation reveals that the translation-based pipeline, using a RoBERTa model fine-tuned for English sentiment classification, achieves superior performance across major metrics. Our results suggest that LLM-based augmentation and cross-lingual transfer via translation both serve as viable approaches to overcome data scarcity and performance limitations in sentiment analysis for low-resource languages. The findings highlight the potential applicability of these approaches to other under-resourced linguistic domains.

Original languageEnglish
Title of host publicationEACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026
EditorsMo El-Haj, Mo El-Haj, Paul Rayson, Mustafa Jarrar, Ignatius Ezeani, Saad Ezzini, Sina Ahmadi, Amal Haddad Haddad, Cynthia Amol, Ahmad Abdelali, Shadi Abudalfa
PublisherAssociation for Computational Linguistics (ACL)
Pages198-207
Number of pages10
ISBN (Electronic)9798891763616
DOIs
StatePublished - 2026
Event2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026 - Rabat, Morocco
Duration: 28 Mar 2026 → …

Publication series

NameEACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026

Conference

Conference2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026
Country/TerritoryMorocco
CityRabat
Period28/03/26 → …

Bibliographical note

Publisher Copyright:
© 2026 Association for Computational Linguistics.

Keywords

  • Urdu sentiment analysis
  • cross-lingual transfer
  • data augmentation
  • large language models
  • machine translation

ASJC Scopus subject areas

  • Artificial Intelligence
  • Linguistics and Language
  • Computer Science Applications
  • Signal Processing

Fingerprint

Dive into the research topics of 'Enhancing Urdu Sentiment Classification through Instruction-Tuned LLMs and Cross-Lingual Transfer'. Together they form a unique fingerprint.

Cite this