TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Cross-Modal Retrieval	COCO 2014	ALADIN	Image-to-text R@1	64.9	# 20
Cross-Modal Retrieval	COCO 2014	ALADIN	Image-to-text R@10	94.5	# 18
Cross-Modal Retrieval	COCO 2014	ALADIN	Image-to-text R@5	88.6	# 20
Cross-Modal Retrieval	COCO 2014	ALADIN	Text-to-image R@1	51.3	# 22
Cross-Modal Retrieval	COCO 2014	ALADIN	Text-to-image R@10	87.5	# 20
Cross-Modal Retrieval	COCO 2014	ALADIN	Text-to-image R@5	79.2	# 21

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/aladin-distilling-fine-grained-alignment/cross-modal-retrieval-on-coco-2014)](https://paperswithcode.com/sota/cross-modal-retrieval-on-coco-2014?p=aladin-distilling-fine-grained-alignment)`

ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval

29 Jul 2022 · Nicola Messina, Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Fabrizio Falchi, Giuseppe Amato, Rita Cucchiara ·

Image-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images and texts. Nonetheless, it has a direct downstream application: cross-modal retrieval, which consists in finding images related to a given query text or vice-versa. Solving this task is of critical importance in cross-modal search engines. Many recent methods proposed effective solutions to the image-text matching problem, mostly using recent large vision-language (VL) Transformer networks. However, these models are often computationally expensive, especially at inference time. This prevents their adoption in large-scale cross-modal retrieval scenarios, where results should be provided to the user almost instantaneously. In this paper, we propose to fill in the gap between effectiveness and efficiency by proposing an ALign And DIstill Network (ALADIN). ALADIN first produces high-effective scores by aligning at fine-grained level images and texts. Then, it learns a shared embedding space - where an efficient kNN search can be performed - by distilling the relevance scores obtained from the fine-grained alignments. We obtained remarkable results on MS-COCO, showing that our method can compete with state-of-the-art VL Transformers while being almost 90 times faster. The code for reproducing our results is available at https://github.com/mesnico/ALADIN.

PDF Abstract

Code

Add Remove Mark official

mesnico/aladin official

Tasks

Add Remove

Cross-Modal Retrieval

Image-text matching

Retrieval

Text Matching

Datasets

MS COCO

Results from the Paper

Edit

Ranked #22 on Cross-Modal Retrieval on COCO 2014

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Cross-Modal Retrieval	COCO 2014	ALADIN	Image-to-text R@1	64.9	# 20	Compare
			Image-to-text R@10	94.5	# 18	Compare
			Image-to-text R@5	88.6	# 20	Compare
			Text-to-image R@1	51.3	# 22	Compare
			Text-to-image R@10	87.5	# 20	Compare
			Text-to-image R@5	79.2	# 21	Compare

Methods

Add Remove

Absolute Position Encodings • Adam • ALIGN • BPE • Dense Connections • Dropout • Label Smoothing • Layer Normalization • Linear Layer • Multi-Head Attention • Position-Wise Feed-Forward Layer • Residual Connection • Scaled Dot-Product Attention • Softmax • Transformer

Edit Social Preview

ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove