The Best Paper Award of the 7th International Conference on Human-Centered Design, Operation and Evaluation of Mobile Communications
has been conferred to
Mahmoud Khalil and Umair Rehman
(Western University, Canada)
for the paper entitled
"Can Multimodal LLMs Model Expert Ratings in Mobile UI Usability Evaluation?"

Mahmoud Khalil
(presenter)

Best Paper Award for the 7th International Conference on Human-Centered Design, Operation and Evaluation of Mobile Communications, in the context of HCI International 2026, Montreal, Canada, 26 - 31 July 2026

Certificate for Best Paper Award of the 7th International Conference on Human-Centered Design, Operation and Evaluation of Mobile Communications presented in the context of HCI International 2026, Montreal, Canada, 26 - 31 July 2026
Paper Abstract
Multimodal Large Language Models (MLLMs) offer a promising path to scaling usability evaluation, yet prior studies have been limited to small datasets or single-model assessments, leaving their reliability at scale unclear. This study addresses this gap by benchmarking four state-of-the-art models—GPT-5, Claude 4 Sonnet, Gemini 2.5 Pro, and Qwen-3 VL— against the UICrit dataset of 1,000 expert-annotated mobile UI screens. Models predicted expert ratings across five dimensions under zero-shot and few-shot prompting; then alignment was assessed using Brennan–Prediger Kappa (KBP) and Wilcoxon signed-rank tests. Results reveal a significant performance dichotomy. On visual dimensions (Aesthetics, Design Quality), GPT-5, Claude 4, and Qwen-3 achieved almost perfect agreement (KBP > = 0.81) under zero-shot prompting, with median rating differences near zero. However, on functional dimensions (Usability, Learnability, Efficiency), models exhibited a systematic overestimation, with median differences ranging from +1.00 to +3.00 points (p < 0.001). Few-shot prompting generally worsened agreement, decreasing by 0.01 to 0.27. Overall, GPT-5 demonstrated the strongest alignment ( KBP > 0.60 for most dimensions), whereas Gemini 2.5 Pro consistently showed the weakest alignment ( KBP < 0.42 on functional dimensions). These findings quantify current limitations of MLLMs: they can approximate expert ratings on visual dimensions, but remain unreliable for interaction-dependent usability judgments from static screenshots.
The full paper is available through SpringerLink, provided that you have proper access rights.


