Adaptive Verification and Confidence-Aware Routing for Reliable Multi-Agent Large Language Model Reasoning
Keywords:
large language models; multi-agent systems; confidence calibration; uncertainty estimation; adaptive verification; dynamic routing; reliable artificial intelligenceAbstract
Large-language-model (LLM) multi-agent systems can improve difficult reasoning by decomposing a problem, assigning specialized roles, and checking intermediate outputs. Most existing systems, however, invoke verification according to a fixed workflow and route subtasks using static role descriptions or uncalibrated self-reports. They consequently spend computation on low-risk states while allowing confidently wrong, high-impact states to propagate. We propose CAVR-MAS (Confidence-Aware Verification and Routing for Multi-Agent Systems), a control framework that couples feature-augmented confidence calibration, domain-conditioned reliability memory, uncertainty-aware routing, and multi-level verification. For every node in a subtask dependency graph, CAVR- MAS estimates calibrated correctness, path disagreement, downstream propagation risk, task criticality, and verification cost. A budget-adaptive controller selects among no check, critique, cross-agent checking, evidence- grounded checking, and re-decomposition by maximizing estimated error reduction per unit cost. Routing and final fusion use conservative posterior estimates of agent reliability rather than raw confidence or majority voting. We prove one-step Bayes optimality under correct estimators, a finite-error routing-regret bound, monotone verification workload with respect to the trigger threshold, a cumulative budget-violation bound, and consistency of nondiscounted reliability memory. We also define an executable evaluation protocol on GSM8K, MATH, PlanBench, MuSiQue, and HotpotQA, with controlled baselines, ablations, calibration analysis, and paired statistical tests. The result is a technically complete reliability-control formulation whose empirical claims are explicitly falsifiable.
References
[1] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, vol. 35, pp. 24824–24837, 2022.
[2] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in International Conference on Learning Representations, 2023.
[3] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in International Conference on Learning Representations, 2023.
[4] S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, “Tree of Thoughts: Deliberate problem solving with large language models,” in Advances in Neural Information Processing Systems, vol. 36, 2023.
[5] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegrefe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al., “Self-Refine: Iterative refinement with self-feedback,” in Advances in Neural Information Processing Systems, vol. 36, pp. 46534–46594, 2023.
[6] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” in Advances in Neural Information Processing Systems, vol. 36, 2023.
[7] Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving factuality and reasoning in language models through multiagent debate,” in Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pp. 11733–11763, 2024.
[8] G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “CAMEL: Communicative agents for ‘mind’ exploration of large scale language model society,” in Advances in Neural Information Processing Systems, vol. 36, pp. 51991–52008, 2023.
[9] S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, et al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” in International Conference on Learning Representations, 2024.
[10] W. Chen, Y. Su, J. Zuo, C. Yang, C. Yuan, C. Qian, C.-M. Chan, Y. Qin, Y. Lu, R. Xie, et al., “AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors,” in International Conference on Learning Representations, 2024.
[11] Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation,” arXiv:2308.08155, 2023.
[12] J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou, “Mixture-of-Agents enhances large language model capabilities,” in International Conference on Learning Representations, 2025.
[13] G. Zhang, Y. Yue, Z. Li, S. Yun, G. Wan, K. Wang, D. Cheng, J. X. Yu, and T. Chen, “Cut the Crap: An economical communication pipeline for LLM-based multi-agent systems,” in International Conference on Learning Representations, 2025.
[14] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” in International Conference on Learning Representations, 2025.
[15] J. Dekoninck, M. Baader, and M. Vechev, “A unified approach to routing and cascading for LLMs,” in Proceedings of the 42nd International Conference on Machine Learning, 2025.
[16] C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, et al., “ChatDev: Communicative agents for software development,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 15174–15186, 2024.
[17] Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.
[18] J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins, “Solving math word problems with process- and outcome-based feedback,” arXiv:2211.14275, 2022.
[19] H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” in International Conference on Learning Representations, 2024.
[20] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proceedings of the 34th International Conference on Machine Learning, PMLR 70, pp. 1321–1330, 2017.
[21] K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. Manning, “Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback,” in Proceedings of EMNLP, pp. 5433–5442, 2023, doi:10.18653/v1/2023.emnlp-main.330.
[22] L. Kuhn, Y. Gal, and S. Farquhar, “Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation,” in International Conference on Learning Representations, 2023.
[23] J. Geng, F. Cai, Y. Wang, H. Koeppl, P. Nakov, and I. Gurevych, “A survey of confidence estimation and calibration in large language models,” in Proceedings of NAACL-HLT, pp. 6577–6595, 2024.
[24] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al., “Training verifiers to solve math word problems,” arXiv:2110.14168, 2021.
[25] D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” in NeurIPS Datasets and Benchmarks Track, 2021.
[26] K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati, “PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change,” in Advances in Neural Information Processing Systems, vol. 36, pp. 38975–38987, 2023.
[27] H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “MuSiQue: Multihop questions via single-hop question composition,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 539–554, 2022, doi:10.1162/tacl__a__00475.
[28] Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning, “HotpotQA: A dataset for diverse, explainable multi-hop question answering,” in Proceedings of EMNLP, pp. 2369–2380, 2018, doi:10.18653/v1/D18- 1259.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Artificial Intelligence Engineering and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.