Home > Published Issues > 2026 > Volume 17, No. 7, 2026 >
JAIT 2026 Vol.17(7): 1397-1408
doi: 10.12720/jait.17.7.1397-1408

Ultra-light Multi-scale Human Parsing with Attention-guided Grouping

Abderrahim Ouza 1,*, Mohamed El Ghmary 2, and Ali Choukri 1
1. Faculty of Science, Ibn Tofail University, Morocco
2. ENSAM, Mohammed V University in Rabat, Morocco
Email: Abderrahim.ouza@uit.ac.ma (A.O.); mohamed.elghmary@um5.ac.ma (M.E.G.); ali.choukri@uit.ac.ma (A.C.)
*Corresponding author

Manuscript received December 28, 2025; revised April 23, 2026; accepted May 12, 2026; published July 28, 2026.

Abstract—Human parsing in large and dynamic environments is commonly handled by large-scale models that deliver high segmentation accuracy but require hundreds of millions of parameters and substantial computational resources. This level of complexity makes such approaches difficult to deploy on power- and resource-constrained devices, particularly in real-time edge Artificial Intelligence (AI) settings. In this paper, we introduce Fast Dilated Spatial Pyramid Pooling + Parsing Grouping Network + Attention (Fast DSPP + PGN + ATTN), an efficient human parsing framework designed to balance segmentation performance with computational efficiency. The model employs a MobileNetV2 encoder with DSPP module for contextual modeling and a PGN-style grouping decoder integrated with light-weight spatial attention and Squeeze-and-Excitation (SE) attention. With 2.14 M parameters and 5.70 Giga Floating Point Operations (GFLOPs), it attains 40.67% mean Intersection over Union (mIoU). In ablation results, we observe that DSPP has the most significant effect (0.0151 mIoU), while spatial attention and SE yield smaller improvements. The accuracy of the grouping decoder is retained while the number of computations can be reduced. The model achieves 0.2490 mIoU on Look into Person (LIP) dataset, with variants to 0.2640, thus confirming robustness over domain shift. This indicates that integrating structured pixel grouping and efficient contextual aggregation together can solve the task of human parsing, even with severe computational constraints, which could render the proposed method particularly appealing to real-world applications such as embedded vision systems/mobile perception and real-time surveillance.
 
Keywords—human parsing, efficient architectures, contextual modeling, edge Artificial Intelligence (AI), multi-scale feature

Cite: Abderrahim Ouza, Mohamed El Ghmary, and Ali Choukri, "Ultra-light Multi-scale Human Parsing with Attention-guided Grouping," Journal of Advances in Information Technology, Vol. 17, No. 7, pp. 1397-1408, 2026. doi: 10.12720/jait.17.7.1397-1408

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions