TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Metric Learning	In-Shop	STIR	R@1	95	# 2
Metric Learning	In-Shop	ViT-Triplet	R@1	92.1	# 9
Metric Learning	Stanford Online Products	STIR	R@1	88.3	# 2
Metric Learning	Stanford Online Products	ViT-Triplet	R@1	86.5	# 4

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stir-siamese-transformer-for-image-retrieval/metric-learning-on-in-shop-1)](https://paperswithcode.com/sota/metric-learning-on-in-shop-1?p=stir-siamese-transformer-for-image-retrieval)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stir-siamese-transformer-for-image-retrieval/metric-learning-on-stanford-online-products-1)](https://paperswithcode.com/sota/metric-learning-on-stanford-online-products-1?p=stir-siamese-transformer-for-image-retrieval)`

STIR: Siamese Transformer for Image Retrieval Postprocessing

26 Apr 2023 · Aleksei Shabanov, Aleksei Tarasov, Sergey Nikolenko ·

Current metric learning approaches for image retrieval are usually based on learning a space of informative latent representations where simple approaches such as the cosine distance will work well. Recent state of the art methods such as HypViT move to more complex embedding spaces that may yield better results but are harder to scale to production environments. In this work, we first construct a simpler model based on triplet loss with hard negatives mining that performs at the state of the art level but does not have these drawbacks. Second, we introduce a novel approach for image retrieval postprocessing called Siamese Transformer for Image Retrieval (STIR) that reranks several top outputs in a single forward pass. Unlike previously proposed Reranking Transformers, STIR does not rely on global/local feature extraction and directly compares a query image and a retrieved candidate on pixel level with the usage of attention mechanism. The resulting approach defines a new state of the art on standard image retrieval datasets: Stanford Online Products and DeepFashion In-shop. We also release the source code at https://github.com/OML-Team/open-metric-learning/tree/main/pipelines/postprocessing/ and an interactive demo of our approach at https://dapladoc-oml-postprocessing-demo-srcappmain-pfh2g0.streamlit.app/

PDF Abstract

Code

Add Remove Mark official

OML-Team/open-metric-learning

762

Tasks

Add Remove

Image Retrieval

Metric Learning

Representation Learning

Re-Ranking

Retrieval

Datasets

DeepFashion

Stanford Online Products

In-Shop

Results from the Paper

Edit

Ranked #2 on Metric Learning on In-Shop

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Metric Learning	In-Shop	STIR	R@1	95	# 2	Compare
Metric Learning	In-Shop	ViT-Triplet	R@1	92.1	# 9	Compare
Metric Learning	Stanford Online Products	STIR	R@1	88.3	# 2	Compare
Metric Learning	Stanford Online Products	ViT-Triplet	R@1	86.5	# 4	Compare

Methods

Add Remove

Absolute Position Encodings • Adam • Dense Connections • Dropout • Label Smoothing • Layer Normalization • Linear Layer • Multi-Head Attention • Position-Wise Feed-Forward Layer • Residual Connection • Scaled Dot-Product Attention • Siamese Network • Transformer • Triplet Loss

Edit Social Preview

STIR: Siamese Transformer for Image Retrieval Postprocessing

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove