SemEval 2027 · Shared Task
MMCultureQA
Given an image and a question, spoken or written, systems produce a short open-ended answer. Answering takes local cultural knowledge, not just what is visible in the picture.
- Speech and text questions
- Different language tracks
- Ranked by BERTScore F1
- Built on the OASIS dataset
The task
Questions that need local knowledge
Questions cover food, places, customs, and everyday objects, and the answer depends on cultural knowledge as much as on the image. Every question comes in two forms, audio and text:
Task 1
Spoken Visual QA
Image + spoken questionShort text answer
Answer a question asked as audio, testing multimodal understanding directly from speech.
Task details →Task 2
Textual Visual QA
Image + written questionShort text answer
The same question, without the speech step.
Task details →Each task runs as a separate track per language. See Tasks & Evaluation and the dataset.
Languages
Language tracks
-
العربية الفصحى Modern Standard Arabic
-
العربية الشامية Levantine Arabic
-
العربية المصرية Egyptian Arabic
-
English English
-
বাংলা Bangla
-
اردو Urdu
-
हिन्दी Hindi
-
অসমীয়া Assamese
-
ગુજરાતી Gujarati
-
मराठी Marathi
-
Italiano Italian
-
Türkçe Turkish
-
አማርኛ Amharic
-
Afaan Oromoo Oromo
-
Soomaali Somali
-
ትግርኛ Tigrinya
Enter any subset. Scoring is on Tasks & Evaluation.
Timeline
Important dates
Dates are tentative.
-
15 July 2026 · Sample data release
A small sample set to preview format and build a pipeline.
-
23 September 2026 · Training data releaseAvailable →
Training and development data for the MENA region, in English and three Arabic varieties.
-
To be announced · More language tracksNext
The remaining language tracks, and a second development set.
-
10 January 2027 · Evaluation phase begins
Test inputs released; submissions open on CodaBench.
-
31 January 2027 · Evaluation phase ends
Final deadline for system submissions.
-
February 2027 · System description papers due
Submit your system description paper.
-
March 2027 · Notification to authors
-
April 2027 · Camera-ready papers due
-
Summer 2027 · SemEval 2027 workshop
Co-located with a major NLP conference.
News
Latest updates
-
23 September 2026
Training and development data are out
The MENA region is released: 10,000 training and 1,000 development items for every track, in English, Modern Standard Arabic, Egyptian Arabic, and Levantine Arabic, with each question available as both text and audio. More language tracks follow. Get the data.
-
14 September 2026
Eighteen language tracks planned
Track coverage now spans Arabic varieties, English, six South Asian languages, Italian, Spanish, Portuguese, Turkish, and four languages of the Horn of Africa. The organizing team has grown to twenty researchers across nine institutions.
-
14 September 2026
Training data release date
The training and development release is being finalised and the date will be announced here and on Slack.
-
30 June 2026
Join us on Slack
We now have a community Slack for announcements, questions, and discussion. Everyone is welcome. Join the Slack.
-
14 June 2026
Registration is open
Teams can now sign up to take part in MMCultureQA. Register here.
-
13 June 2026
Shared task website is live
Explore the task definition, the two tasks, the OASIS dataset, the evaluation plan, important dates, and the organizers.
-
13 June 2026
Save the date: sample data on 15 July 2026
A sample set will be released first so teams can preview the format. Training data follows on 1 September 2026.