TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Visual Dialog	VisDial v0.9 val	MVAN	MRR	0.6765	# 11
Visual Dialog	VisDial v0.9 val	MVAN	Mean Rank	3.73	# 2
Visual Dialog	VisDial v0.9 val	MVAN	R@1	54.65	# 3
Visual Dialog	VisDial v0.9 val	MVAN	R@10	91.47	# 3
Visual Dialog	VisDial v0.9 val	MVAN	R@5	83.85	# 2
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	NDCG (x 100)	59.37	# 41
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	MRR (x 100)	64.84	# 15
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	R@1	51.45	# 16
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	R@5	81.12	# 18
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	R@10	90.65	# 16
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	Mean	3.97	# 64

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/multi-view-attention-networks-for-visual/visual-dialog-on-visdial-v09-val)](https://paperswithcode.com/sota/visual-dialog-on-visdial-v09-val?p=multi-view-attention-networks-for-visual)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/multi-view-attention-networks-for-visual/visual-dialog-on-visual-dialog-v1-0-test-std)](https://paperswithcode.com/sota/visual-dialog-on-visual-dialog-v1-0-test-std?p=multi-view-attention-networks-for-visual)`

Multi-View Attention Network for Visual Dialog

29 Apr 2020 · Sungjin Park, Taesun Whang, Yeochan Yoon, Heuiseok Lim ·

Visual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered. To resolve the visual dialog task, a high-level understanding of various multimodal inputs (e.g., question, dialog history, and image) is required. Specifically, it is necessary for an agent to 1) determine the semantic intent of question and 2) align question-relevant textual and visual contents among heterogeneous modality inputs. In this paper, we propose Multi-View Attention Network (MVAN), which leverages multiple views about heterogeneous inputs based on attention mechanisms. MVAN effectively captures the question-relevant information from the dialog history with two complementary modules (i.e., Topic Aggregation and Context Matching), and builds multimodal representations through sequential alignment processes (i.e., Modality Alignment). Experimental results on VisDial v1.0 dataset show the effectiveness of our proposed model, which outperforms the previous state-of-the-art methods with respect to all evaluation metrics.

PDF Abstract

Code

Add Remove Mark official

taesunwhang/MVAN-VisDial official

Tasks

Add Remove

Visual Dialog

Datasets

VisDial

Results from the Paper

Edit

Ranked #11 on Visual Dialog on VisDial v0.9 val

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Visual Dialog	VisDial v0.9 val	MVAN	MRR	0.6765	# 11	Compare
			Mean Rank	3.73	# 2	Compare
			R@1	54.65	# 3	Compare
			R@10	91.47	# 3	Compare
			R@5	83.85	# 2	Compare
Visual Dialog	Visual Dialog v1.0 test-std	MVAN	NDCG (x 100)	59.37	# 41	Compare
			MRR (x 100)	64.84	# 15	Compare
			R@1	51.45	# 16	Compare
			R@5	81.12	# 18	Compare
			R@10	90.65	# 16	Compare
			Mean	3.97	# 64	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

Multi-View Attention Network for Visual Dialog

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove