πŸ“ž +91-7667918914 | βœ‰οΈ ijarcce@gmail.com
International Journal of Advanced Research in Computer and Communication Engineering
International Journal of Advanced Research in Computer and Communication Engineering A monthly Peer-reviewed & Refereed journal
ISSN Online 2278-1021ISSN Print 2319-5940Since 2012
IJARCCE adheres to the suggestive parameters outlined by the University Grants Commission (UGC) for peer-reviewed journals, upholding high standards of research quality, ethical publishing, and academic excellence.
← Back to VOLUME 15, ISSUE 7, JULY 2026

A Scalable Extractive Text Summarisation Model for Hausa Language Documents Using K-Nearest Neighbours and Sparse Matrix Representation

Musa Tanimu Karatu, Abdulsamad Muazu, Ibrahim Saidu

πŸ‘ 5 viewsπŸ“₯ 1 download
Share: 𝕏 f in ✈ βœ‰
Abstract: Extractive text summarisation has become an essential Natural Language Processing (NLP) task for managing the rapid growth of digital textual information. However, conventional graph-based summarisation methods often suffer from high computational complexity and memory consumption due to exhaustive pairwise sentence similarity computations, limiting their applicability to large-scale datasets and low-resource languages such as Hausa. This study proposes a scalable extractive text summarisation model for Hausa language documents using K-Nearest Neighbours (KNN) and Sparse Matrix Representation to improve computational efficiency while preserving summary quality. The proposed approach preprocesses Hausa news articles through tokenisation, stop-word removal, and TF-IDF vectorisation, after which KNN identifies the nearest neighbouring sentences to construct a sparse similarity matrix. This reduces computational complexity from O(NΒ²) to O(N log N) by eliminating unnecessary similarity computations while preserving the most informative sentence relationships for graph-based ranking. The proposed model was evaluated on a corpus of Hausa news articles and compared with a Graph Convolutional Network–Recurrent Neural Network (GCN– RNN) model using ROUGE-1, ROUGE-2, ROUGE-L, execution time, memory consumption, compression ratio, and computational complexity. Experimental results demonstrate that the proposed model completed summarisation in 32 s while consuming only 4.79 MB of memory, compared with 104 s and 112.96 MB for the GCN–RNN model, representing reductions of 69.23% in execution time and 95.76% in memory usage. It also achieved ROUGE-1, ROUGE-2, and ROUGE-L F1-scores of 85.36%, 75.64%, and 76.77%, respectively, outperforming the baseline while maintaining superior computational efficiency. These findings demonstrate that the proposed KNN with Sparse Matrix Representation provides an efficient, scalable, and practical framework for extractive text summarisation of Hausa language documents and offers a promising solution for other low-resource languages.

Keywords: Extractive Text Summarisation, Hausa Language, Natural Language Processing, K-Nearest Neighbours, Sparse Matrix Representation, TF-IDF, ROUGE Evaluation, Scalable Text Summarisation, Computational Efficiency.

How to Cite:

[1] Musa Tanimu Karatu, Abdulsamad Muazu, Ibrahim Saidu, β€œA Scalable Extractive Text Summarisation Model for Hausa Language Documents Using K-Nearest Neighbours and Sparse Matrix Representation,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15713

Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License.