πŸ“ž +91-7667918914 | βœ‰οΈ ijarcce@gmail.com
International Journal of Advanced Research in Computer and Communication Engineering
International Journal of Advanced Research in Computer and Communication Engineering A monthly Peer-reviewed & Refereed journal
ISSN Online 2278-1021ISSN Print 2319-5940Since 2012
IJARCCE adheres to the suggestive parameters outlined by the University Grants Commission (UGC) for peer-reviewed journals, upholding high standards of research quality, ethical publishing, and academic excellence.
← Back to VOLUME 15, ISSUE 9, SEPTEMBER 2026

A Machine Learning and NLP Framework for Automatic Title Similarity Detection and Preliminary Verification

Miss. Suvarna Bhoi, Dr. Sarkarsinha H. Rajput, Dr. Shital A. Patil, Dr. Dinesh D. Puri

πŸ‘ 15 viewsπŸ“₯ 10 downloads
Share: 𝕏 f in ✈ βœ‰
Abstract: Title verification becomes difficult when a large collection of publication titles has to be checked for similarity before a new title is accepted. Exact string matching is not sufficient because titles may contain spelling variations, reordered words, prefixes or suffixes, periodicity changes, phonetic variations, or different words with similar meanings. This research proposes a practical Natural Language Processing (NLP) and Machine Learning framework for automatic preliminary title verification. The proposed approach first normalizes and preprocesses title text and then represents titles using Term Frequency–Inverse Document Frequency (TF-IDF). Cosine Similarity is used as the main lexical similarity measure to rank existing titles against a newly submitted title. Character-level and phonetic features can be added to improve detection of spelling and sound-based variations. Logistic Regression, Support Vector Machine (SVM), and Random Forest are considered for classifying title pairs as similar or non-similar. The system will return the most similar existing titles, similarity scores, and a verification recommendation for human review. Evaluation will consider accuracy, precision, recall, F1-score, false-positive and false-negative behaviour, and processing time. The study aims to provide an interpretable, scalable, and practical approach that reduces repeated manual comparison while keeping final verification under human control.

Keywords: Title Similarity Detection, Natural Language Processing (NLP), TF-IDF, Cosine Similarity, SBERT, Machine Learning, Semantic Similarity, Preliminary Verification.

How to Cite:

[1] Miss. Suvarna Bhoi, Dr. Sarkarsinha H. Rajput, Dr. Shital A. Patil, Dr. Dinesh D. Puri, β€œA Machine Learning and NLP Framework for Automatic Title Similarity Detection and Preliminary Verification,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15927

Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License.