Home > Published Issues > 2026 > Volume 17, No. 7, 2026 >
JAIT 2026 Vol.17(7): 1310-1320
doi: 10.12720/jait.17.7.1310-1320

Cross-Dataset Generalization Framework for Cybercrime Detection Using CIC-IDS2017, CIC-Phishing2019, and Malicious Uniform Resource Locator (URL) Data

R. Adinarayana 1,* and G. Vamsi Krishna 2
1. Department of Computer Science and Systems Engineering, Andhra University, Visakhapatnam, India
2. Department of Computer Science and Engineering, Dr. Lankapalli Bullayya College of Engineering (A), Visakhapatnam, India
Email: aaditeresa@gmail.com (R.A.); vamsikrishna527@gmail.com (G.V.K.)
*Corresponding author

Manuscript received February 8, 2026; revised April 8, 2026; accepted May 6, 2026; published July 23, 2026.

Abstract—Proposed cybercrime detection models demonstrate satisfactory performance on specific benchmark datasets; however, they are not always robust across diverse, practical scenarios. This paper examines the generalization performance of a cross-dataset measure for cybercrime detection across network traffic, phishing email, and malicious Uniform Resource Locator (URL) domains. The work combines three popular security benchmarks Canadian Institute for Cybersecurity Intrusion Detection System Dataset2017 (CIC-IDS2017), Canadian Institute for Cybersecurity Phishing Dataset2019 (CIC-Phishing2019), and Malicious URL 2020 in a single preprocessing and learning pipeline to avoid dataset-specific bias. One or more datasets are used to train models, which are then directly evaluated on previously unseen datasets to verify transferability across distribution shifts. We will discuss performance while considering accuracy, F1−Score, Receiver Operating Characteristic (ROC), and fold-to-fold stability. Experiments demonstrate that direct training with a single random source results in significant performance deterioration, with a 23% decrease in F1−Score when applied to unseen datasets. In contrast, the degradation caused by the proposed framework is kept below 10% and maintains Receiver Operating Characteristic–Area Under the Curve (ROC–AUC) values consistently above 0.90. Paired significance testing demonstrates that the gains in robustness are highly significant (p < 0.01). The results indicate that cross-dataset evaluation is essential for the development of deployable cybercrime detection systems, providing empirical insights for developing generalization-oriented security analytics.
 
Keywords—cross-dataset generalization, cybercrime detection, network traffic analysis, phishing email detection, malicious Uniform Resource Locator (URL) classification, security benchmark datasets

Cite: R. Adinarayana and G. Vamsi Krishna, "Cross-Dataset Generalization Framework for Cybercrime Detection Using CIC-IDS2017, CIC-Phishing2019, and Malicious Uniform Resource Locator (URL) Data," Journal of Advances in Information Technology, Vol. 17, No. 7, pp. 1310-1320, 2026. doi: 10.12720/jait.17.7.1310-1320

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions