Home > Published Issues > 2026 > Volume 17, No. 9, 2026 >
JAIT 2026 Vol.17(9): 1716-1723
doi: 10.12720/jait.17.9.1716-1723

Design and Development of a Controlled Multi-agent Conversational System Using Large Language Models for Public Services: A Case Study

Zaid Rustamov 1, Mir Amir Pashayev 1, Sanan Jafarov 1, Ulvi Aslanli 1, Aghasalim Mammadov 1,
Alakbar Valizada 1, and Samir Rustamov 2,*
1. AI Laboratory, Megasec LLC, Baku, Azerbaijan
2. School of IT and Engineering, ADA University, Baku, Azerbaijan
Email: zaid.r@megasec.ai (Z.R.); amir.p@megasec.ai (M.A.P.); sanan.j@megasec.az (S.J.); ulvi.a@megasec.ai (U.A.); aghasalim.m@megasec.ai (A.M.); alakbar.v@megasec.ai (A.V.); srustamov@ada.edu.az (S.R.)
*Corresponding author

Manuscript received March 5, 2026; revised May 15, 2026; accepted June 9, 2026; published September 14, 2026.

Abstract—Large Language Model (LLM)-based conversational systems are increasingly adopted as front-facing interfaces for public and institutional services. However, end-to-end and tool-augmented approaches often rely on implicit, prompt-driven decision making, which limits controllability, transparency, and safety in regulated environments. This paper presents a case study of a controlled multi-agent conversational architecture deployed within a single institutional public-service environment, explicitly separating intent classification, service selection, and response generation into independent and auditable decision stages. The proposed design introduces an intent-driven control mechanism and a deterministic service routing process that constrains response generation to supported service contexts while enabling systematic handling of ambiguity and out-of-scope queries. To support reliable deployment, we propose an iterative development workflow in which prompts are treated as evolving behavioral specifications validated through structured testing and monitoring. The system is evaluated on 1255 real conversational queries spanning multi-turn interactions collected from the deployed system. Experimental results demonstrate high reliability across all modules, achieving 99.2% accuracy in intent classification, 97.7% accuracy in single-service routing, and 98.8% correctness for service-grounded responses. Overall, the system produces correct responses for 97.6% of evaluated queries while maintaining deterministic handling of unsupported and ambiguous requests. Within the scope of this single-institution case study, the findings highlight the importance of explicit decision boundaries and modular design for deploying LLM-based assistants in safety-critical and public-service domains and provide practical guidance for building controllable, service-grounded conversational systems.
 
Keywords—conversational AI, intent classification, multi-agent systems, public-service chatbots, safety-critical AI, tool-augmented systems
 
Cite: Zaid Rustamov, Mir Amir Pashayev, Sanan Jafarov, Ulvi Aslanli, Aghasalim Mammadov, Alakbar Valizada, and Samir Rustamov, "Design and Development of a Controlled Multi-agent Conversational System Using Large Language Models for Public Services: A Case Study," Journal of Advances in Information Technology, Vol. 17, No. 9, pp. 1716-1723, 2026. doi: 10.12720/jait.17.9.1716-1723

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions