tiiuae researchers introduced Alyah, a benchmark of 1,173 multiple-choice samples designed to evaluate Arabic LLMs on Emirati dialect comprehension. The dataset was manually collected from native speakers and covers culturally grounded expressions, greetings, poetry, and anecdotes. The team evaluated base and instruction-tuned models across several families, finding that dialectal performance remains challenging, particularly for culturally embedded and pragmatic language. The work was contributed by Omar saif alkaabi, Ahmed Alzubaidi, Hamza Alobeidli, Shaikha Alsuwaidi, Mohammed Alyafeai, Leen AlQadi, Basma Boussaha, and Hakim Hacid.
No score is assigned. Sources and their independence are shown in the citation chain below.