Tasks & Evaluation
Open-ended answers to questions about images, in the language the question was asked in.
Task definition
Given an image and a question about it, a system produces a short, open-ended answer in the language of the question. Answers are free text, so they must be semantically correct, grounded in the image, and right for its cultural context.
The two tasks
Every question exists in two forms, audio and text. Both tasks use the same images and questions, and differ only in how the question arrives.
| Task | Input | Output |
|---|---|---|
| Task 1: Spoken Visual QA | Image + spoken question (audio) | Open-ended answer (text) |
| Task 2: Textual Visual QA | Image + written question (text) | Open-ended answer (text) |
Language tracks
Each task is organized as a separate track for each language, with systems ranked independently within each track. Participants may enter any subset of the available language tracks.
- Modern Standard Arabic
- Levantine Arabic
- Egyptian Arabic
- English
- Bangla
- Urdu
- Hindi
- Assamese
- Gujarati
- Marathi
- Italian
- Turkish
- Amharic
- Oromo
- Somali
- Tigrinya
Evaluation
Because answers are open-ended, scoring rewards meaning rather than exact wording.
BERTScore F1
Semantic similarity to the reference, so the wording can differ.
BLEU & ROUGE
Lexical overlap, reported for context. Not used for the ranking.
LLM-based analysis
May be reported as extra analysis. Not used for the ranking.
The official evaluation script ships with the data, so you can reproduce the ranking locally. Submissions run on CodaBench; see Participate for the workflow.
Ready to take part?
Registration is free, and the Slack community is where news lands first.