A Hybrid Neural-Symbolic Framework for Reliable Automated Source Code Documentation

Authors

  • Adewuyi Joseph Oluwaseyi Department of Computer Science, School of Computing, Babcock University, Ogun State, Nigeria
  • Kareem Temiloluwa Faith Department of Computer Science, School of Computing, Babcock University, Ogun State, Nigeria
  • Odelola Precious Ayooluwa Department of Computer Science, School of Computing, Babcock University, Ogun State, Nigeria
  • Oluwafemi Oluwasemilore Esther Department of Computer Science, School of Computing, Babcock University, Ogun State, Nigeria

DOI:

https://doi.org/10.70112/ajcst-2026.15.2.4452

Keywords:

Code Summarization, Automated Documentation, CodeBERT, CodeT5, Deep Learning, Transformer Models, Software Engineering, Nigeria

Abstract

Inadequate code documentation is a persistent challenge in software development. Time constraints often lead developers to neglect commenting, resulting in poorly documented codebases. This undermines knowledge sharing, reduces maintainability, and weakens team collaboration. Over time, such neglect erodes software sustainability, as critical knowledge becomes confined to individual developers rather than embedded within the code itself. This study presents CodeSage, an automated system for generating inline code comments and docstrings via a retrieval-augmented architecture. CodeBERT (125M parameters) retrieves semantically similar examples from a FAISS index of 50,000 code-documentation pairs, while CodeT5 (220M parameters) generates natural language documentation. A modular parser supports five languages: Python, JavaScript, TypeScript, Java, and C++. Evaluation used BLEU, ROUGE-L, METEOR, and BERTScore on 80 Python code-comment pairs from CodeSearchNet, alongside human evaluation across four quality criteria. The system achieved a BERTScore F1 of 0.7748, METEOR of 0.2978, ROUGE-L of 0.2585, and BLEU of 0.1583, consistent with fine-tuned code summarization benchmarks. Human evaluators rated accuracy and readability at 4.6/5.0, conciseness at 4.5/5.0, and completeness at 4.3/5.0. Mean generation time was 1.17 seconds per sample with a 100% success rate. CodeSage demonstrates that hybrid retrieval-augmented generation with domain-specific transformer models produces semantically accurate, readable documentation across multiple programming languages. It offers a practical productivity tool for Nigerian academic and professional software environments, with future work targeting domain-specific fine-tuning and expanded multilingual support.

References

[1] E. A. AlOmar, H. Alrubaye, M. W. Mkaouer, A. Ouni, and M. Kessentini, “Refactoring practices in the context of modern code review: An industrial case study at Xerox,” Empirical Software Engineering, vol. 26, no. 3, 2021.

[2] B. Wei, Y. Li, G. Li, X. Xia, and Z. Jin, “Retrieve and refine: Exemplar-based neural comment generation,” in Proc. 42nd Int. Conf. Software Engineering (ICSE), 2020.

[3] Z. Feng et al., “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of EMNLP, 2020, pp. 1536–1547.

[4] D. Guo et al., “GraphCodeBERT: Pre-training code representations with data flow,” arXiv preprint arXiv:2009.08366, 2021.

[5] Y. Wang, W. Wang, S. Joty, and S. C. H. Hoi, “CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” arXiv preprint arXiv:2109.00859, 2021.

[6] T. P. T. Le, A. M. T. Bui, H. N. D. Pham, A. Bucaioni, and P. T. Nguyen, “When retriever meets generator: A joint model for code comment generation,” arXiv preprint arXiv:2507.12558, 2025.

[7] A. LeClair, S. Jiang, and C. McMillan, “A neural model for generating natural language summaries of program subroutines,” arXiv preprint arXiv:1902.01954, 2019.

[8] A. Mastropaolo, E. Aghajani, L. Pascarella, and G. Bavota, “An empirical study on code comment completion,” arXiv preprint arXiv:2107.10544, 2021.

[9] D. Prasad Ghale and M. Dabbagh, “Automated code comments generation using large language models: Empirical evaluation of T5 and BART,” IEEE Access, vol. 13, pp. 141420–141433, 2025.

[10]H. Husain, H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt, “CodeSearchNet Challenge: Evaluating the state of semantic code search,” arXiv preprint arXiv:1909.09436, 2020.

[11]X. Song, H. Sun, X. Wang, and J. Yan, “A survey of automatic generation of source code comments: Algorithms and techniques,” IEEE Access, vol. 7, pp. 111411–111428, 2019.

[12]E. Shi et al., “On the evaluation of neural code summarization,” in Proc. 44th Int. Conf. Software Engineering (ICSE), 2022, pp. 1597–1608.

[13]W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Proc. NAACL-HLT, 2021, pp. 2655–2668.

Downloads

Published

16-08-2026

How to Cite

Oluwaseyi, A. J., Faith, K. T., Ayooluwa, O. P., & Esther, O. O. (2026). A Hybrid Neural-Symbolic Framework for Reliable Automated Source Code Documentation. Asian Journal of Computer Science and Technology , 15(2), 22–28. https://doi.org/10.70112/ajcst-2026.15.2.4452

Issue

Section

Research Article

Similar Articles

<< < 22 23 24 25 26 27 28 29 > >> 

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)