Bridging the NLP Gap: A Transformer-Based Sentiment Analysis Framework for Code-Mixed Nigerian Pidgin

Authors

  • Adewuyi Joseph Oluwaseyi Department of Computer Science, School of Computing, Babcock University, Ilishan-Remo, Ogun State, Nigeria
  • Odiase Sophia Osariemen Department of Computer Science, School of Computing, Babcock University, Ilishan-Remo, Ogun State, Nigeria
  • Imhanzenobe Daisy Ehijie Department of Computer Science, School of Computing, Babcock University, Ilishan-Remo, Ogun State, Nigeria
  • Ibrahim Mujeeb Olayemi Department of Computer Science, School of Computing, Babcock University, Ilishan-Remo, Ogun State, Nigeria

DOI:

https://doi.org/10.70112/ajcst-2026.15.2.4459

Keywords:

Nigerian Pidgin, Sentiment Analysis, Natural Language Processing, AfriBERTa, AfroXLM-R, Transformer Models, Low-Resource Languages, Deep Learning

Abstract

This study addresses the limitations of existing Natural Language Processing (NLP) systems in analyzing Nigerian Pidgin, a widely spoken but low-resource language. Traditional sentiment analysis models, primarily designed for English, perform poorly due to challenges such as non-standardized spelling, semantic drift, and code-mixing. These limitations hinder accurate interpretation of public opinion expressed in Nigerian Pidgin. This project therefore aims to develop an effective sentiment analysis system capable of classifying Pidgin text into positive, negative, and neutral categories. A transformer-based approach was adopted using pretrained multilingual models, specifically AfriBERTa and XLM-R. The system was trained on the Nigerian Pidgin subset of the AfriSenti dataset (approximately 14,000 labeled tweets). A preprocessing pipeline involving data cleaning, normalization of spelling variations, and SentencePiece tokenization was implemented to improve model performance. The models were fine-tuned using supervised learning and evaluated using accuracy, precision, recall, and weighted F1-score. The model achieved a peak accuracy of 70.25%, with a precision of 0.6805, recall of 0.7025, and an F1-score of 0.6819. Comparative evaluation showed AfroXLM-R performed slightly better with 71.10% accuracy, emphasizing the advantage of African language-specific pretraining. These results demonstrate that transformer models can effectively capture sentiment in Nigerian Pidgin despite its linguistic complexity. In conclusion, the study confirms the viability of transformer-based NLP for low-resource languages. While challenges such as sarcasm, code-mixing, and limited datasets remain, the system provides a strong foundation for real-world applications and future improvements in indigenous language processing.

References

[1] S. Muhammad, D. I. Adelani, R. Niyongabo, J. O. Alabi, P. Ogayo, and A. Bukula, “AfriSenti: A Twitter sentiment analysis benchmark for African languages,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12345–12360.

[2] D. M. Eberhard, G. F. Simons, and C. D. Fennig, Eds., Ethnologue: Languages of the World, 26th ed. SIL International, 2023.

[3] N. G. Faraclas, Nigerian Pidgin. Routledge, 2013.

[4] C. A. Okoloegbo, U. F. Eze, G. A. Chukwudebe, and O. C. Nwokonkwo, “Multilingual cyberbullying detector (CD) application for Nigerian Pidgin and Igbo language corpus,” in Proceedings of the 2022 5th Information Technology for Education and Development (ITED), 2022, pp. 1–6. https://doi.org/10.1109/ITED56637.2022.10051345.

[5] S. Rathi, S. Pande, H. Atkare, R. Tangsali, A. Vyawahare, and D. Kadam, “Trinity at SemEval-2023 Task 12: Sentiment analysis for low-resource African languages using Twitter dataset,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), 2023, pp. 1161–1165. https://doi.org/10.18653/v1/2023.semeval-1.161.

[6] S. A. Salahudeen et al., “HausaNLP at SemEval-2023 Task 12: Leveraging African low-resource tweet data for sentiment analysis,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), 2023, pp. 42–48. https://doi.org/10.18653/v1/2023.semeval-1.6.

[7] N. Hughes, K. Baker, A. Singh, A. Singh, T. Dauda, and S. Bhattacharya, “Bhattacharya_Lab at SemEval-2023 task 12: A transformer-based language model for sentiment classification for low resource African languages: Nigerian Pidgin and Yoruba,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Toronto, Canada, Jul. 2023, pp. 1502–1507. [Online]. Available: https://aclanthology.org/2023.semeval-1.207.

[8] N. Raychawdhary, A. Das, G. Dozier, and C. D. Seals, “Seals Lab at SemEval-2023 task 12: Sentiment analysis for low-resource African languages, Hausa and Igbo,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Toronto, Canada, Jul. 2023, p. 1508.

[9] I. Shode, D. I. Adelani, J. Peng, and A. Feldman, “NollySenti: Leveraging transfer learning and machine translation for Nigerian movie sentiment classification,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 986–998. https://doi.org/10.18653/v1/2023.acl-short.85.

[10] P. Rust, J. Pfeiffer, I. Vulić, S. Ruder, and I. Gurevych, “How good is your tokenizer? On the monolingual performance of multilingual language models,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2021, pp. 3118–3135. https://doi.org/10.18653/v1/2021.acl-long.243.

Downloads

Published

16-09-2026

How to Cite

Oluwaseyi, A. J., Osariemen, O. S., Ehijie, I. D., & Olayemi, I. M. (2026). Bridging the NLP Gap: A Transformer-Based Sentiment Analysis Framework for Code-Mixed Nigerian Pidgin. Asian Journal of Computer Science and Technology , 15(2), 29–37. https://doi.org/10.70112/ajcst-2026.15.2.4459

Issue

Section

Research Article

Similar Articles

<< < 17 18 19 20 21 22 23 24 > >> 

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)