Skip to main navigation Skip to search Skip to main content

BERT-Based Sentiment Analysis for Low-Resourced Languages: A Case Study of Urdu Language

  • Muhammad Rehan Ashraf
  • , Yasmeen Jana
  • , Qasim Umer*
  • , M. Arfan Jaffar
  • , Sungwook Chung
  • , Waheed Yousuf Ramay
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

26 Scopus citations

Abstract

Sentiment analysis holds significant importance in research projects by providing valuable insights into public opinions. However, the majority of sentiment analysis studies focus on the English language, leaving a gap in research for other low-resourced languages or regional languages, e.g., Persian, Pashto, and Urdu. Moreover, computational linguists face the challenge of developing lexical resources for these languages. In light of this, this paper presents a deep learning-based approach for Urdu Text Sentiment Analysis (USA-BERT), leveraging Bidirectional Encoder Representations from Transformers and introduces an Urdu Dataset for Sentiment Analysis-23 (UDSA-23). USA-BERT first preprocesses the Urdu reviews by exploiting BERT-Tokenizer. Second, it creates BERT embeddings for each Urdu review. Third, given the BERT embeddings, it fine-tunes a deep learning classifier (BERT). Finally, it employs the Pareto principle on two datasets (the state-of-the-art (UCSA-21) and UDSA-23) to assess USA-BERT. The assessment results demonstrate that USA-BERT significantly surpasses the existing methods by improving the accuracy and f-measure up to 26.09% and 25.87%, respectively.

Original languageEnglish
Pages (from-to)110245-110259
Number of pages15
JournalIEEE Access
Volume11
DOIs
StatePublished - 2023
Externally publishedYes

Bibliographical note

Publisher Copyright:
© 2013 IEEE.

Keywords

  • BERT
  • Natural language processing
  • Urdu
  • classification
  • sentiment analysis

ASJC Scopus subject areas

  • General Computer Science
  • General Materials Science
  • General Engineering

Fingerprint

Dive into the research topics of 'BERT-Based Sentiment Analysis for Low-Resourced Languages: A Case Study of Urdu Language'. Together they form a unique fingerprint.

Cite this