← Back to VOLUME 15, ISSUE 7, JULY 2026
This work is licensed under a Creative Commons Attribution 4.0 International License.
A Scalable Extractive Text Summarisation Model for Hausa Language Documents Using K-Nearest Neighbours and Sparse Matrix Representation
Musa Tanimu Karatu, Abdulsamad Muazu, Ibrahim Saidu
π 5 viewsπ₯ 1 download
Abstract: Extractive text summarisation has become an essential Natural Language Processing (NLP) task for managing the rapid growth of digital textual information. However, conventional graph-based summarisation methods often suffer from high computational complexity and memory consumption due to exhaustive pairwise sentence similarity computations, limiting their applicability to large-scale datasets and low-resource languages such as Hausa. This study proposes a scalable extractive text summarisation model for Hausa language documents using K-Nearest Neighbours (KNN) and Sparse Matrix Representation to improve computational efficiency while preserving summary quality. The proposed approach preprocesses Hausa news articles through tokenisation, stop-word removal, and TF-IDF vectorisation, after which KNN identifies the nearest neighbouring sentences to construct a sparse similarity matrix. This reduces computational complexity from O(NΒ²) to O(N log N) by eliminating unnecessary similarity computations while preserving the most informative sentence relationships for graph-based ranking. The proposed model was evaluated on a corpus of Hausa news articles and compared with a Graph Convolutional NetworkβRecurrent Neural Network (GCNβ RNN) model using ROUGE-1, ROUGE-2, ROUGE-L, execution time, memory consumption, compression ratio, and computational complexity. Experimental results demonstrate that the proposed model completed summarisation in 32 s while consuming only 4.79 MB of memory, compared with 104 s and 112.96 MB for the GCNβRNN model, representing reductions of 69.23% in execution time and 95.76% in memory usage. It also achieved ROUGE-1, ROUGE-2, and ROUGE-L F1-scores of 85.36%, 75.64%, and 76.77%, respectively, outperforming the baseline while maintaining superior computational efficiency. These findings demonstrate that the proposed KNN with Sparse Matrix Representation provides an efficient, scalable, and practical framework for extractive text summarisation of Hausa language documents and offers a promising solution for other low-resource languages.
Keywords: Extractive Text Summarisation, Hausa Language, Natural Language Processing, K-Nearest Neighbours, Sparse Matrix Representation, TF-IDF, ROUGE Evaluation, Scalable Text Summarisation, Computational Efficiency.
Keywords: Extractive Text Summarisation, Hausa Language, Natural Language Processing, K-Nearest Neighbours, Sparse Matrix Representation, TF-IDF, ROUGE Evaluation, Scalable Text Summarisation, Computational Efficiency.
How to Cite:
[1] Musa Tanimu Karatu, Abdulsamad Muazu, Ibrahim Saidu, βA Scalable Extractive Text Summarisation Model for Hausa Language Documents Using K-Nearest Neighbours and Sparse Matrix Representation,β International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15713
