Bridging the NLP Gap: A Transformer-Based Sentiment Analysis Framework for Code-Mixed Nigerian Pidgin
DOI:
https://doi.org/10.70112/ajcst-2026.15.2.4459Keywords:
Nigerian Pidgin, Sentiment Analysis, Natural Language Processing, AfriBERTa, AfroXLM-R, Transformer Models, Low-Resource Languages, Deep LearningAbstract
This study addresses the limitations of existing Natural Language Processing (NLP) systems in analyzing Nigerian Pidgin, a widely spoken but low-resource language. Traditional sentiment analysis models, primarily designed for English, perform poorly due to challenges such as non-standardized spelling, semantic drift, and code-mixing. These limitations hinder accurate interpretation of public opinion expressed in Nigerian Pidgin. This project therefore aims to develop an effective sentiment analysis system capable of classifying Pidgin text into positive, negative, and neutral categories. A transformer-based approach was adopted using pretrained multilingual models, specifically AfriBERTa and XLM-R. The system was trained on the Nigerian Pidgin subset of the AfriSenti dataset (approximately 14,000 labeled tweets). A preprocessing pipeline involving data cleaning, normalization of spelling variations, and SentencePiece tokenization was implemented to improve model performance. The models were fine-tuned using supervised learning and evaluated using accuracy, precision, recall, and weighted F1-score. The model achieved a peak accuracy of 70.25%, with a precision of 0.6805, recall of 0.7025, and an F1-score of 0.6819. Comparative evaluation showed AfroXLM-R performed slightly better with 71.10% accuracy, emphasizing the advantage of African language-specific pretraining. These results demonstrate that transformer models can effectively capture sentiment in Nigerian Pidgin despite its linguistic complexity. In conclusion, the study confirms the viability of transformer-based NLP for low-resource languages. While challenges such as sarcasm, code-mixing, and limited datasets remain, the system provides a strong foundation for real-world applications and future improvements in indigenous language processing.
References
[1] S. Muhammad, D. I. Adelani, R. Niyongabo, J. O. Alabi, P. Ogayo, and A. Bukula, “AfriSenti: A Twitter sentiment analysis benchmark for African languages,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12345–12360.
[2] D. M. Eberhard, G. F. Simons, and C. D. Fennig, Eds., Ethnologue: Languages of the World, 26th ed. SIL International, 2023.
[3] N. G. Faraclas, Nigerian Pidgin. Routledge, 2013.
[4] C. A. Okoloegbo, U. F. Eze, G. A. Chukwudebe, and O. C. Nwokonkwo, “Multilingual cyberbullying detector (CD) application for Nigerian Pidgin and Igbo language corpus,” in Proceedings of the 2022 5th Information Technology for Education and Development (ITED), 2022, pp. 1–6. https://doi.org/10.1109/ITED56637.2022.10051345.
[5] S. Rathi, S. Pande, H. Atkare, R. Tangsali, A. Vyawahare, and D. Kadam, “Trinity at SemEval-2023 Task 12: Sentiment analysis for low-resource African languages using Twitter dataset,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), 2023, pp. 1161–1165. https://doi.org/10.18653/v1/2023.semeval-1.161.
[6] S. A. Salahudeen et al., “HausaNLP at SemEval-2023 Task 12: Leveraging African low-resource tweet data for sentiment analysis,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), 2023, pp. 42–48. https://doi.org/10.18653/v1/2023.semeval-1.6.
[7] N. Hughes, K. Baker, A. Singh, A. Singh, T. Dauda, and S. Bhattacharya, “Bhattacharya_Lab at SemEval-2023 task 12: A transformer-based language model for sentiment classification for low resource African languages: Nigerian Pidgin and Yoruba,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Toronto, Canada, Jul. 2023, pp. 1502–1507. [Online]. Available: https://aclanthology.org/2023.semeval-1.207.
[8] N. Raychawdhary, A. Das, G. Dozier, and C. D. Seals, “Seals Lab at SemEval-2023 task 12: Sentiment analysis for low-resource African languages, Hausa and Igbo,” in Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), Toronto, Canada, Jul. 2023, p. 1508.
[9] I. Shode, D. I. Adelani, J. Peng, and A. Feldman, “NollySenti: Leveraging transfer learning and machine translation for Nigerian movie sentiment classification,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 986–998. https://doi.org/10.18653/v1/2023.acl-short.85.
[10] P. Rust, J. Pfeiffer, I. Vulić, S. Ruder, and I. Gurevych, “How good is your tokenizer? On the monolingual performance of multilingual language models,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2021, pp. 3118–3135. https://doi.org/10.18653/v1/2021.acl-long.243.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Centre for Research and Innovation

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
