Tasks & Evaluation

Open-ended answers to questions about images, in the language the question was asked in.

Task definition

Given an image and a question about it, a system produces a short, open-ended answer in the language of the question. Answers are free text, so they must be semantically correct, grounded in the image, and right for its cultural context.

The two tasks

Every question exists in two forms, audio and text. Both tasks use the same images and questions, and differ only in how the question arrives.

Task Input Output
Task 1: Spoken Visual QA Image + spoken question (audio) Open-ended answer (text)
Task 2: Textual Visual QA Image + written question (text) Open-ended answer (text)

Language tracks

Each task is organized as a separate track for each language, with systems ranked independently within each track. Participants may enter any subset of the available language tracks.

  • Modern Standard Arabic
  • Levantine Arabic
  • Egyptian Arabic
  • English
  • Bangla
  • Urdu
  • Hindi
  • Assamese
  • Gujarati
  • Marathi
  • Italian
  • Turkish
  • Amharic
  • Oromo
  • Somali
  • Tigrinya

Evaluation

Because answers are open-ended, scoring rewards meaning rather than exact wording.

Official ranking

BERTScore F1

Semantic similarity to the reference, so the wording can differ.

Auxiliary

BLEU & ROUGE

Lexical overlap, reported for context. Not used for the ranking.

Supplementary

LLM-based analysis

May be reported as extra analysis. Not used for the ranking.

The official evaluation script ships with the data, so you can reproduce the ranking locally. Submissions run on CodaBench; see Participate for the workflow.

Ready to take part?

Registration is free, and the Slack community is where news lands first.

Register your team Join the Slack