HyperAIHyperAI

Command Palette

Search for a command to run...

Visual Question Answering (VQA)

Visual Question Answering (VQA) is a task in the field of computer vision that aims to answer questions about images using natural language. The core objective of this task is to enable machines to understand the content of images and provide answers in an accurate and coherent linguistic form. VQA has significant application value in human-computer interaction, intelligent assistance, and content understanding, significantly enhancing the visual cognitive abilities of machines.

Leaderboard

100 models total

Benchmarks

GQA Test2019
VQA v2 test-dev
Oscar
OK-VQA
Prophet
VQA v2 test-std
BEiT-3
MSVD-QA
HCRN
DocVQA test
Human
MSRVTT-QA
HCRN
GQA test-dev
InfographicVQA
Gemini Ultra (pixel only)
A-OKVQA
VizWiz 2020 VQA
InfiMM-Eval
CLEVR
NS-VQA (1K programs)
IconQA
Patch-TRM
VQA v2 val
BLIP-2 ViT-G FlanT5 XXL (zero-shot)
TextVQA test-standard
COCO Visual Question Answering (VQA) real images 1.0 open ended
IllusionVQA
VCR (Q-A) test
COCO Visual Question Answering (VQA) real images 1.0 multiple choice
MCB 7 att.
VCR (QA-R) test
VQA-CP
CSS
InfoSeek
VCR (Q-AR) test
GPT4RoI
VizWiz 2018
LXR955, No Ensemble
VLM2-Bench
VQA-CE
RandImg
GQA test-std
ProTo
WHOOPS!
AutoHallusion
VQA v1 test-dev
SAAA (ResNet)
VQA v1 test-std
SAAA (ResNet)
CLEVR-Humans
HallusionBench
mPLUG-Owl
PMC-VQA
POPE
VizWiz 2020 Answerability
COCO Visual Question Answering (VQA) real images 2.0 open ended
HDU-USYD-UNCC
QLEVR
MAC
Visual7W
CMN
AI2D
COCO Visual Question Answering (VQA) abstract images 1.0 open ended
COCO Visual Question Answering (VQA) abstract 1.0 multiple choice
GQA
PEVL+
GRIT
MMVM
PlotQA-D1
PlotQA-D2
TextVQA
VCR (Q-A) dev
VL-BERTLARGE
VCR (Q-AR) dev
VL-BERTLARGE
VCR (QA-R) dev
VL-BERTLARGE
AMBER
Asclepius
ChartQA
F-VQA
ZS-F-VQA
FigureQA - test 1
PReFIL
ScreenSpot
CORE-MM
DocVQA
DocVQA val
BERT LARGE Baseline
MSCOCO
NuPlanQA-Eval
OVAD benchmark
SITE
TDIUC
Accuracy
TGIF-QA
VQA-X
ActivityNet
BLIP-2 T5
Adversarial
ArtQuest
PrefixLM with CLIP and T5
CHAIR
ChartVA-AITQA
ChartVA-ChartQA
ChartVA-PlotQA
COCO
DeepForm
DVQA test-familiar
PReFIL (Oracle OCR)
EgoSchema
Lyra-Pro
HMDB51-VISPR
HRScene
ImageNet
InstructBLIP / Max Token 128
InstructBLIP / Max Token 64
LLaVA-1.5 / Max Token 128
LLaVA-1.5 / Max Token 64
LongDocURL
MM-AlignBench
MM-Vet
MME
MME-Hallucination
MVBench
Popular
Quilt-conversation
Qwen-VL / Max Token 128
Qwen-VL / Max Token 64
Rad-VQA
Random
RetVQA
MI-BART
SLAKE
UCF101-VISPR
Video MME
Visual Genome (pairs)
CMN
Visual Genome (subjects)
VizWiz 2018 Answerability
VQ-FocusAmbiguity
WebSRC
ZS-F-VQA
SAN † - hard mask