Visual Question Answering (VQA)

758 papers with code • 62 benchmarks • 112 datasets

Visual Question Answering (VQA) is a task in computer vision that involves answering questions about an image. The goal of VQA is to teach machines to understand the content of an image and answer questions about it in natural language.

Image Source: visualqa.org

Benchmarks

Add a Result

These leaderboards are used to track progress in Visual Question Answering (VQA)

Dataset	Best Model	Compare
VQA v2 test-dev	PaLI	See all
VQA v2 test-std	BEiT-3	See all
OK-VQA	PaLI-X-VPD	See all
MSVD-QA	VLAB	See all
DocVQA test	Human	See all
MSRVTT-QA	VLAB	See all
InfographicVQA	Gemini Ultra (pixel only)	See all
COCO Visual Question Answering (VQA) real images 1.0 open ended	SAN	See all
CLEVR	NS-VQA (1K programs)	See all
GQA test-dev	CFR	See all
InfiMM-Eval	GPT-4V	See all
A-OKVQA	SMoLA-PaLI-X Specialist Model	See all
IconQA	ViLT	See all
VCR (Q-A) test	GPT4RoI	See all
VQA v2 val	BLIP-2 ViT-G FlanT5 XXL (zero-shot)	See all
COCO Visual Question Answering (VQA) real images 1.0 multiple choice	MCB 7 att.	See all
VQA-CP	CSS	See all
VQA-CE	RandImg	See all
VCR (QA-R) test	GPT4RoI	See all
VCR (Q-AR) test	GPT4RoI	See all
GQA test-std	NSM	See all
VQA v1 test-dev	SAAA (ResNet)	See all
IllusionVQA	GPT4-Vision	See all
VQA v1 test-std	RAU (ResNet)	See all
GQA Test2019	TRRNet (Ensemble)	See all
WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images	BLIP2 FlanT5-XXL (Fine-tuned)	See all
InfoSeek	RA-VQAv2 w/ PreFLMR	See all
CLEVR-Humans	MDETR	See all
QLEVR	MAC	See all
COCO Visual Question Answering (VQA) abstract images 1.0 open ended	Graph VQA	See all
COCO Visual Question Answering (VQA) abstract 1.0 multiple choice	Graph VQA	See all
Visual7W	CMN	See all
COCO Visual Question Answering (VQA) real images 2.0 open ended	HDU-USYD-UNCC	See all
PMC-VQA	MedVInT	See all
AI2D	SMoLA-PaLI-X Specialist Model	See all
VCR (Q-A) dev	VL-BERTLARGE	See all
VCR (QA-R) dev	VL-BERTLARGE	See all
VCR (Q-AR) dev	VL-BERTLARGE	See all
VizWiz 2018	Colin	See all
VizWiz 2020 VQA	PaLI	See all
PlotQA-D1	MatCha	See all
FigureQA - test 1	PReFIL	See all
PlotQA-D2	MatCha	See all
F-VQA	ZS-F-VQA	See all
HallusionBench	GPT-4V	See all
TDIUC	Accuracy	See all
VizWiz 2020 Answerability	CLIP-Ensemble	See all
TextVQA test-standard	TAP	See all
GQA	RelViT	See all
GRIT	Unified-IOXL	See all
TGIF-QA	HiTeA	See all
VQA-X	OFA-X-MT	See all
Visual Genome (pairs)	CMN	See all
Visual Genome (subjects)	CMN	See all
DocVQA val	BERT LARGE Baseline	See all
ZS-F-VQA	SAN † - hard mask	See all
WebSRC	DUBLIN	See all
DeepForm	DUBLIN	See all
SciGraphQA	SciGraphQA-baseline	See all
DVQA test-familiar	PReFIL (Oracle OCR)	See all
CORE-MM	GPT-4V	See all
VizWiz 2018 Answerability	ensemble_two_best	See all

Show all 62 benchmarks

Collapse benchmarks

Libraries

Use these libraries to find Visual Question Answering (VQA) models and implementations

huggingface/transformers

13 papers

124,527

salesforce/lavis

7 papers

8,685

ZephyrZhuQi/ssbaseline

5 papers

gabegrand/adversarial-vqa

5 papers

See all 15 libraries.

Datasets

Subtasks

Embodied Question Answering

3D Question Answering (3D-QA)

Generative Visual Question Answering

Factual Visual Question Answering

Most implemented papers

Most implemented Social Latest No code

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

facebookresearch/vilbert-multi-task • • NeurIPS 2019

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language.

Paper
Code

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

akirafukui/vqa-mcb • • EMNLP 2016

Approaches to multimodal pooling include element-wise product or sum, as well as concatenation of the visual and textual representations.

Paper
Code

Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge

peteanderson80/bottom-up-attention • CVPR 2018

This paper presents a state-of-the-art model for visual question answering (VQA), which won the first place in the 2017 VQA Challenge.

Paper
Code

Compositional Attention Networks for Machine Reasoning

stanfordnlp/mac-network • • ICLR 2018

We present the MAC network, a novel fully differentiable neural network architecture, designed to facilitate explicit and expressive reasoning.

Paper
Code

Hierarchical Question-Image Co-Attention for Visual Question Answering

jiasenlu/HieCoAttenVQA • • NeurIPS 2016

In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN).

Paper
Code

Pythia v0.1: the Winning Entry to the VQA Challenge 2018

facebookresearch/pythia • • 26 Jul 2018

We demonstrate that by making subtle but important changes to the model architecture and the learning rate schedule, fine-tuning image features, and adding data augmentation, we can significantly improve the performance of the up-down model on VQA v2. 0 dataset -- from 65. 67% to 70. 22%.

Paper
Code

LXMERT: Learning Cross-Modality Encoder Representations from Transformers

airsplay/lxmert • • IJCNLP 2019

In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder.

Paper
Code

GPT-4 Technical Report

openai/evals • Preprint 2023

We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.

Paper
Code

Hadamard Product for Low-rank Bilinear Pooling

jnhwkim/MulLowBiVQA • • 14 Oct 2016

Bilinear models provide rich representations compared with linear models.

Paper
Code

Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

peteanderson80/Matterport3DSimulator • • CVPR 2018

This is significant because a robot interpreting a natural-language navigation instruction on the basis of what it sees is carrying out a vision and language process that is similar to Visual Question Answering.

Paper
Code

Visual Question Answering (VQA)

Benchmarks Add a Result

Libraries

Datasets

Subtasks

Most implemented papers

Content

Benchmarks

Add a Result