You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This paper introduces BloomBench, a new bilingual (English-Arabic), cognitively-informed benchmark based on Bloom's Taxonomy to systematically evaluate the reas…
Evaluation and inference scripts for DisasterVQA — a benchmark dataset for assessing Vision-Language Models (VLMs) on disaster-response visual question answerin…
Nahw is a comprehensive benchmark for evaluating Arabic grammar understanding in large language models. It includes multiple-choice grammar questions (Nahw-MCQ)…
This repository provides all datasets released with the EMNLP 2025 paper Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art …