Visual Question Answering (VQA)

767 papers with code • 62 benchmarks • 112 datasets

Visual Question Answering (VQA) is a task in computer vision that involves answering questions about an image. The goal of VQA is to teach machines to understand the content of an image and answer questions about it in natural language.

Image Source: visualqa.org

Benchmarks

Add a Result

These leaderboards are used to track progress in Visual Question Answering (VQA)

Dataset	Best Model	Compare
VQA v2 test-dev	PaLI	See all
VQA v2 test-std	BEiT-3	See all
OK-VQA	PaLI-X-VPD	See all
MSVD-QA	VLAB	See all
DocVQA test	Human	See all
MSRVTT-QA	VLAB	See all
InfographicVQA	Gemini Ultra (pixel only)	See all
COCO Visual Question Answering (VQA) real images 1.0 open ended	SAN	See all
CLEVR	NS-VQA (1K programs)	See all
GQA test-dev	CFR	See all
InfiMM-Eval	GPT-4V	See all
A-OKVQA	SMoLA-PaLI-X Specialist Model	See all
IconQA	ViLT	See all
VCR (Q-A) test	GPT4RoI	See all
VQA v2 val	BLIP-2 ViT-G FlanT5 XXL (zero-shot)	See all
COCO Visual Question Answering (VQA) real images 1.0 multiple choice	MCB 7 att.	See all
VQA-CP	CSS	See all
VQA-CE	RandImg	See all
VCR (QA-R) test	GPT4RoI	See all
VCR (Q-AR) test	GPT4RoI	See all
GQA test-std	NSM	See all
VQA v1 test-dev	SAAA (ResNet)	See all
IllusionVQA	GPT4-Vision	See all
VQA v1 test-std	RAU (ResNet)	See all
GQA Test2019	TRRNet (Ensemble)	See all
WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images	BLIP2 FlanT5-XXL (Fine-tuned)	See all
InfoSeek	RA-VQAv2 w/ PreFLMR	See all
CLEVR-Humans	MDETR	See all
QLEVR	MAC	See all
COCO Visual Question Answering (VQA) abstract images 1.0 open ended	Graph VQA	See all
COCO Visual Question Answering (VQA) abstract 1.0 multiple choice	Graph VQA	See all
Visual7W	CMN	See all
COCO Visual Question Answering (VQA) real images 2.0 open ended	HDU-USYD-UNCC	See all
PMC-VQA	MedVInT	See all
AI2D	SMoLA-PaLI-X Specialist Model	See all
VCR (Q-A) dev	VL-BERTLARGE	See all
VCR (QA-R) dev	VL-BERTLARGE	See all
VCR (Q-AR) dev	VL-BERTLARGE	See all
VizWiz 2018	Colin	See all
VizWiz 2020 VQA	PaLI	See all
PlotQA-D1	MatCha	See all
FigureQA - test 1	PReFIL	See all
PlotQA-D2	MatCha	See all
F-VQA	ZS-F-VQA	See all
HallusionBench	GPT-4V	See all
TDIUC	Accuracy	See all
VizWiz 2020 Answerability	CLIP-Ensemble	See all
TextVQA test-standard	TAP	See all
GQA	RelViT	See all
GRIT	Unified-IOXL	See all
TGIF-QA	HiTeA	See all
VQA-X	OFA-X-MT	See all
Visual Genome (pairs)	CMN	See all
Visual Genome (subjects)	CMN	See all
DocVQA val	BERT LARGE Baseline	See all
ZS-F-VQA	SAN † - hard mask	See all
WebSRC	DUBLIN	See all
DeepForm	DUBLIN	See all
SciGraphQA	SciGraphQA-baseline	See all
DVQA test-familiar	PReFIL (Oracle OCR)	See all
CORE-MM	GPT-4V	See all
VizWiz 2018 Answerability	ensemble_two_best	See all

Show all 62 benchmarks

Collapse benchmarks

Libraries

Use these libraries to find Visual Question Answering (VQA) models and implementations

huggingface/transformers

10 papers

125,385

salesforce/lavis

7 papers

8,774

ZephyrZhuQi/ssbaseline

5 papers

gabegrand/adversarial-vqa

5 papers

See all 15 libraries.

Datasets

Subtasks

Embodied Question Answering

3D Question Answering (3D-QA)

Generative Visual Question Answering

Factual Visual Question Answering

Latest papers

Most implemented Social Latest No code

OmniFusion Technical Report

airi-institute/omnifusion • • 9 Apr 2024

We propose an \textit{OmniFusion} model based on a pretrained LLM and adapters for visual modality.

191

09 Apr 2024

Paper
Code

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

boheumd/MA-LMM • • 8 Apr 2024

However, existing LLM-based large multimodal models (e. g., Video-LLaMA, VideoChat) can only take in a limited number of frames for short video understanding.

118

08 Apr 2024

Paper
Code

Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

faceonlive/ai-research • 6 Apr 2024

In this paper, we present a novel approach, Joint Visual and Text Prompting (VTPrompt), that employs fine-grained visual information to enhance the capability of MLLMs in VQA, especially for object-oriented perception.

186

06 Apr 2024

Paper
Code

Evaluating Text-to-Visual Generation with Image-to-Text Generation

linzhiqiu/t2v_metrics • • 1 Apr 2024

For instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce reliable scores for complex prompts involving compositions of objects, attributes, and relations.

01 Apr 2024

Paper
Code

Unsolvable Problem Detection: Evaluating Trustworthiness of Vision Language Models

atsumiyai/upd • • 29 Mar 2024

This paper introduces a novel and significant challenge for Vision Language Models (VLMs), termed Unsolvable Problem Detection (UPD).

29 Mar 2024

Paper
Code

A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions

riken-grp/gazevqa • 26 Mar 2024

Such ambiguities in questions are often clarified by the contexts in conversational situations, such as joint attention with a user or user gaze information.

26 Mar 2024

Paper
Code

Intrinsic Subgraph Generation for Interpretable Graph based Visual Question Answering

digitalphonetics/intrinsic-subgraph-generation-for-vqa • 26 Mar 2024

In this work, we introduce an interpretable approach for graph-based VQA and demonstrate competitive performance on the GQA dataset.

26 Mar 2024

Paper
Code

IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

csebuetnlp/illusionvqa • • 23 Mar 2024

GPT4V, the best-performing VLM, achieves 62. 99% accuracy (4-shot) on the comprehension task and 49. 7% on the localization task (4-shot and Chain-of-Thought).

23 Mar 2024

Paper
Code

MedPromptX: Grounded Multimodal Prompting for Chest X-ray Diagnosis

biomedia-mbzuai/medpromptx • • 22 Mar 2024

Chest X-ray images are commonly used for predicting acute and chronic cardiopulmonary conditions, but efforts to integrate them with structured clinical data face challenges due to incomplete electronic health records (EHR).

22 Mar 2024

Paper
Code

Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering

bowen-upenn/Multi-Agent-VQA • • 21 Mar 2024

This work explores the zero-shot capabilities of foundation models in Visual Question Answering (VQA) tasks.

21 Mar 2024

Paper
Code

Visual Question Answering (VQA)

Benchmarks Add a Result

Libraries

Datasets

Subtasks

Latest papers

Content

Benchmarks

Add a Result