Federated Evaluation of Medical Vision Language Models Across Middle Eastern Hospital Data Residency Settings

Authors

  • Ahmad Alzoubi Al-Hussein Bin Talal University, Faculty of Information Technology, Department of Computer Information Systems, Queen Rania Al Abdullah Street, Ma'an 71111, Jordan Author
  • Omar Khoury Tafila Technical University, Faculty of Business Administration, Department of Management Information Systems, Engineering College Street, Tafila 66110, Jordan Author
  • Ziad Haddad Irbid National University, Faculty of Information Technology, Department of Computer Science, University Street, Irbid 21110, Jordan Author

Abstract

Medical vision-language models are  being assessed on centralized benchmarks, yet many hospital systems cannot pool clinical images across borders or institutions because of data residency requirements, governance rules, and local privacy expectations. This paper presents an empirical study of federated evaluation for binary medical image question answering across Middle Eastern hospital settings. The study uses a simulated multi-institutional network spanning Gulf, Levant, and North African clinical environments, with 28,800 image-question cases from chest radiography, abdominal ultrasound, musculoskeletal radiography, and retinal imaging. Each hospital node keeps images and labels locally while sharing only aggregate evaluation statistics, uncertainty summaries, and model behavior signatures. We compare centralized evaluation, single-site evaluation, pooled-site reporting without harmonization, and a proposed federated audit protocol that standardizes task definitions, site-level confidence intervals, subgroup reporting, and disagreement tracing. The proposed protocol identifies performance instability that is hidden by pooled reporting, including a 13.4 percentage point gap between the best and weakest hospital nodes for rare thoracic findings and a 9.1 percentage point gap for Arabic-English question variants. It also reduces misleading model selection decisions by requiring site-consistency thresholds rather than relying on overall accuracy alone. The findings suggest that federated evaluation can make medical vision-language assessment more realistic for Middle Eastern health systems where cross-border data sharing is limited but coordinated model governance is still needed.

Downloads

Download data is not yet available.

Downloads

Published

2026-05-04

How to Cite

Alzoubi, Ahmad, et al. “Federated Evaluation of Medical Vision Language Models Across Middle Eastern Hospital Data Residency Settings”. Journal of Data, Models, and Decision Making for Intelligent Systems and Society, vol. 16, no. 5, May 2026, pp. 1-14, https://scidataconsortium.com/index.php/J-DMDMIS/article/view/Federated-Evaluation-of.