← Back to VOLUME 15, ISSUE 10, OCTOBER 2026
This work is licensed under a Creative Commons Attribution 4.0 International License.
AI-BASED RANSOMWARE DETECTION FRAMEWORK USING MACHINE LEARNING
Kajal Shildar Padvi, Dr.Rucha Mohite
Downloads: Download PDF
π 12 viewsπ₯ 6 downloads
Abstract: Ransomware remains one of the major cybersecurity threats because modern ransomware variants continuously evolve and may differ considerably from previously identified samples. Traditional signature-based detection approaches are effective against known threats but may have limited capability against previously unseen variants. Machine learning provides an alternative approach by learning discriminative patterns from malicious and benign executable samples. However, high classification performance under a conventional random train-test split does not necessarily demonstrate generalization to ransomware families that were not represented during training.
This research proposes an AI-Based Ransomware Detection Framework Using Machine Learning based on static characteristics of Windows Portable Executable (PE) files. The framework extracts a compact set of eight static features consisting of API count, DLL count, unique API count, unique DLL count, file size, packed status, entropy, and PE timestamp year. Five machine learning algorithmsβLogistic Regression, Decision Tree, Random Forest, XGBoost, and LightGBMβare trained and compared using accuracy, precision, recall, F1score, ROC-AUC, false positive rate, and false negative rate.
The experimental framework additionally incorporates validation-based probability threshold analysis, ransomware- family-disjoint evaluation, family-level generalization analysis, Isolation Forest-based novelty detection, and SHAP- based explainability. Random Forest achieved the best overall standard-test performance with 95.42% accuracy, 96.22% precision, 94.28% recall, 95.24% F1-score, and 98.98% ROC-AUC. When evaluated on ransomware families excluded from training, the Random Forest model achieved 90.20% accuracy, 98.88% precision, 87.02% recall, 92.57% F1-score, and 97.72% ROC-AUC. The family-disjoint experiment demonstrates that detection becomes more challenging when the model encounters ransomware families not represented during training. Isolation Forest additionally identified approximately 8.04% of unseen ransomware samples as anomalous relative to the known ransomware distribution.
The results demonstrate that a compact static PE representation combined with ensemble machine learning can provide strong ransomware classification performance while enabling additional analysis of generalization, novelty, and interpretability.
Keywords: Ransomware Detection, Machine Learning, Static Analysis, Portable Executable, Random Forest, Explainable AI, SHAP, Unseen Ransomware Families, Anomaly Detection, Malware Detection.
This research proposes an AI-Based Ransomware Detection Framework Using Machine Learning based on static characteristics of Windows Portable Executable (PE) files. The framework extracts a compact set of eight static features consisting of API count, DLL count, unique API count, unique DLL count, file size, packed status, entropy, and PE timestamp year. Five machine learning algorithmsβLogistic Regression, Decision Tree, Random Forest, XGBoost, and LightGBMβare trained and compared using accuracy, precision, recall, F1score, ROC-AUC, false positive rate, and false negative rate.
The experimental framework additionally incorporates validation-based probability threshold analysis, ransomware- family-disjoint evaluation, family-level generalization analysis, Isolation Forest-based novelty detection, and SHAP- based explainability. Random Forest achieved the best overall standard-test performance with 95.42% accuracy, 96.22% precision, 94.28% recall, 95.24% F1-score, and 98.98% ROC-AUC. When evaluated on ransomware families excluded from training, the Random Forest model achieved 90.20% accuracy, 98.88% precision, 87.02% recall, 92.57% F1-score, and 97.72% ROC-AUC. The family-disjoint experiment demonstrates that detection becomes more challenging when the model encounters ransomware families not represented during training. Isolation Forest additionally identified approximately 8.04% of unseen ransomware samples as anomalous relative to the known ransomware distribution.
The results demonstrate that a compact static PE representation combined with ensemble machine learning can provide strong ransomware classification performance while enabling additional analysis of generalization, novelty, and interpretability.
Keywords: Ransomware Detection, Machine Learning, Static Analysis, Portable Executable, Random Forest, Explainable AI, SHAP, Unseen Ransomware Families, Anomaly Detection, Malware Detection.
How to Cite:
[1] Kajal Shildar Padvi, Dr.Rucha Mohite, βAI-BASED RANSOMWARE DETECTION FRAMEWORK USING MACHINE LEARNING,β International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE)
