Home > Published Issues > 2026 > Volume 17, No. 9, 2026 >
JAIT 2026 Vol.17(9): 1821-1830
doi: 10.12720/jait.17.9.1821-1830

Sentiment Classification on IMDb Movie Reviews: A Controlled Comparative Study of Classical, Deep-learning, Feature-fusion, and Fine-tuned Transformer Models

Santosh Kumar Banbhrani 1,2,3,*, Wang Shaojie 1,*, Liu Baiyang 1, Chen Wei 1, Pir Dino Soomro 4,
Ali Raza Rang 2,3, Dhani Bux Talpur 2,3, and Asad Ullah Shah 3
1. School of Electrical Engineering, Shaoyang University, Shaoyang, China
2. Department of Information and Computing, University of Sufism and Modern Science, Bhitshah, Matiari, Sindh, Pakistan
3. Kulliyyah of Information and Communication Technology, International Islamic University Malaysia, Jalan Gombak, Kuala Lumpur, Selangor, Malaysia
4. School of Information Science and Technology, Dalian Maritime University, Dalian, China
Email: santosh.kumar@hnsyu.edu.cn and banbhrani@gmail.com (S.K.B.); shaojiew@hnsyu.edu.cn (W.S.); 806717285@qq.com (L.B.); 2569@hnsyu.edu.cn (C.W.); pirdinosoomro@dlmu.edu.cn (P.D.S.); alirazarang@gmail.com (A.R.R.); dhanibux@hotmail.com (D.B.T.); asadullah@iium.edu.my (A.U.S.)
*Corresponding author

Manuscript received February 23, 2026; revised March 10, 2026; accepted July 10, 2026; published September 24, 2026.

Abstract—Sentiment analysis is one of the most important tasks in natural language processing, as it allows opinions expressed in text to be automatically categorized. This study presents a systematically controlled comparative analysis of classical machine-learning, deep-learning, feature-fusion, and fully fine-tuned transformer models for binary sentiment classification on the Internet Movie Database (IMDb) Movie Reviews dataset. While the individual components of the framework, including Term Frequency-Inverse Document Frequency (TF-IDF), Bidirectional Encoder Representations from Transformers (BERT) embeddings, feature fusion, and Multi-Layer Perceptron (MLP) classification, are established techniques, this study contributes a comprehensive empirical evaluation of their combined effectiveness for the target classification task, providing insights into the practical benefits of integrating lexical and contextual representations. Under a single experimental framework, all models are evaluated on a commonly held-out test set with 95% bootstrap confidence intervals, paired McNemar significance testing, confusion-matrix and error analysis, and a computational-complexity comparison. The fully fine-tuned transformers achieve the highest performance Robustly Optimized BERT Pretraining Approach (RoBERTa), F1-Score = 0.9427. The TF-IDF and BERT feature-fusion model is significantly more accurate than the classical, deep-learning, and BERT-only baselines (all p < 0.001), but significantly less accurate than the fine-tuned transformers (all p < 0.01); furthermore, concatenating the [CLS] token and mean-pooled representations does not significantly improve upon mean pooling alone (p = 0.48). The study provides a transparent, evidence-based account of the accuracy and cost trade-offs among these model families.
 
Keywords—sentiment analysis, Internet Movie Database (IMDb) dataset, machine learning, deep learning, hybrid models, Term Frequency-Inverse Document Frequency (TF-IDF), Bidirectional Encoder Representations from Transformers (BERT)
 
Cite: Santosh Kumar Banbhrani, Wang Shaojie, Liu Baiyang, Chen Wei, Pir Dino Soomro, Ali Raza Rang, Dhani Bux Talpu, and Asad Ullah Shah, "Sentiment Classification on IMDb Movie Reviews: A Controlled Comparative Study of Classical, Deep-learning, Feature-fusion, and Fine-tuned Transformer Models," Journal of Advances in Information Technology, Vol. 17, No. 9, pp. 1821-1830, 2026. doi: 10.12720/jait.17.9.1821-1830

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions