Federated Evaluation of Medical Vision Language Models Across Middle Eastern Hospital Data Residency Settings
Abstract
Medical vision-language models are being assessed on centralized benchmarks, yet many hospital systems cannot pool clinical images across borders or institutions because of data residency requirements, governance rules, and local privacy expectations. This paper presents an empirical study of federated evaluation for binary medical image question answering across Middle Eastern hospital settings. The study uses a simulated multi-institutional network spanning Gulf, Levant, and North African clinical environments, with 28,800 image-question cases from chest radiography, abdominal ultrasound, musculoskeletal radiography, and retinal imaging. Each hospital node keeps images and labels locally while sharing only aggregate evaluation statistics, uncertainty summaries, and model behavior signatures. We compare centralized evaluation, single-site evaluation, pooled-site reporting without harmonization, and a proposed federated audit protocol that standardizes task definitions, site-level confidence intervals, subgroup reporting, and disagreement tracing. The proposed protocol identifies performance instability that is hidden by pooled reporting, including a 13.4 percentage point gap between the best and weakest hospital nodes for rare thoracic findings and a 9.1 percentage point gap for Arabic-English question variants. It also reduces misleading model selection decisions by requiring site-consistency thresholds rather than relying on overall accuracy alone. The findings suggest that federated evaluation can make medical vision-language assessment more realistic for Middle Eastern health systems where cross-border data sharing is limited but coordinated model governance is still needed.