TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Cross-Modal Retrieval	COCO 2014	ViSTA	Image-to-text R@1	68.9	# 19
Cross-Modal Retrieval	COCO 2014	ViSTA	Image-to-text R@10	95.4	# 17
Cross-Modal Retrieval	COCO 2014	ViSTA	Image-to-text R@5	90.1	# 18
Cross-Modal Retrieval	COCO 2014	ViSTA	Text-to-image R@1	52.6	# 21
Cross-Modal Retrieval	COCO 2014	ViSTA	Text-to-image R@10	87.6	# 19
Cross-Modal Retrieval	COCO 2014	ViSTA	Text-to-image R@5	79.6	# 20
Cross-Modal Retrieval	Flickr30k	ViSTA	Image-to-text R@1	89.5	# 10
Cross-Modal Retrieval	Flickr30k	ViSTA	Image-to-text R@10	99.6	# 10
Cross-Modal Retrieval	Flickr30k	ViSTA	Image-to-text R@5	98.4	# 10
Cross-Modal Retrieval	Flickr30k	ViSTA	Text-to-image R@1	75.8	# 12
Cross-Modal Retrieval	Flickr30k	ViSTA	Text-to-image R@10	96.9	# 11
Cross-Modal Retrieval	Flickr30k	ViSTA	Text-to-image R@5	94.2	# 11

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/vista-vision-and-scene-text-aggregation-for/cross-modal-retrieval-on-flickr30k)](https://paperswithcode.com/sota/cross-modal-retrieval-on-flickr30k?p=vista-vision-and-scene-text-aggregation-for)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/vista-vision-and-scene-text-aggregation-for/cross-modal-retrieval-on-coco-2014)](https://paperswithcode.com/sota/cross-modal-retrieval-on-coco-2014?p=vista-vision-and-scene-text-aggregation-for)`

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

CVPR 2022 · Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Kun Yao, Jie Chen, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang ·

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single $\textbf{Vi}$sion and $\textbf{S}$cene $\textbf{T}$ext $\textbf{A}$ggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least $\bf{8.4}\%$ at Recall@1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.

PDF Abstract CVPR 2022 PDF CVPR 2022 Abstract

Code

Add Remove Mark official

No code implementations yet. Submit your code now

Tasks

Add Remove

Contrastive Learning

Cross-Modal Retrieval

Retrieval

Datasets

MS COCO

Visual Genome

Flickr30k

TextVQA CTC

Results from the Paper

Edit

Ranked #10 on Cross-Modal Retrieval on Flickr30k (using extra training data)

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Cross-Modal Retrieval	COCO 2014	ViSTA	Image-to-text R@1	68.9	# 19	Compare
			Image-to-text R@10	95.4	# 17	Compare
			Image-to-text R@5	90.1	# 18	Compare
			Text-to-image R@1	52.6	# 21	Compare
			Text-to-image R@10	87.6	# 19	Compare
			Text-to-image R@5	79.6	# 20	Compare
Cross-Modal Retrieval	Flickr30k	ViSTA	Image-to-text R@1	89.5	# 10	Compare
			Image-to-text R@10	99.6	# 10	Compare
			Image-to-text R@5	98.4	# 10	Compare
			Text-to-image R@1	75.8	# 12	Compare
			Text-to-image R@10	96.9	# 11	Compare
			Text-to-image R@5	94.2	# 11	Compare

Methods

Add Remove

AWARE • Contrastive Learning

Edit Social Preview

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove