TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Image Classification	ImageNet	GLiT-Bases	Top 1 Accuracy	82.3%	# 501
Image Classification	ImageNet	GLiT-Bases	Number of params	96.1M	# 857
Image Classification	ImageNet	GLiT-Bases	GFLOPs	17	# 352
Image Classification	ImageNet	GLiT-Smalls	Top 1 Accuracy	80.5%	# 638
Image Classification	ImageNet	GLiT-Smalls	Number of params	24.6M	# 585
Image Classification	ImageNet	GLiT-Smalls	GFLOPs	4.4	# 208
Image Classification	ImageNet	GLiT-Tinys	Top 1 Accuracy	76.3%	# 847
Image Classification	ImageNet	GLiT-Tinys	Number of params	7.2M	# 454
Image Classification	ImageNet	GLiT-Tinys	GFLOPs	1.4	# 128

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/glit-neural-architecture-search-for-global/image-classification-on-imagenet)](https://paperswithcode.com/sota/image-classification-on-imagenet?p=glit-neural-architecture-search-for-global)`

GLiT: Neural Architecture Search for Global and Local Image Transformer

ICCV 2021 · BoYu Chen, Peixia Li, Chuming Li, Baopu Li, Lei Bai, Chen Lin, Ming Sun, Junjie Yan, Wanli Ouyang ·

We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient transformer variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification.

PDF Abstract ICCV 2021 PDF ICCV 2021 Abstract

Code

Add Remove Mark official

bychen515/glit official

lpxtt/simtrack

Tasks

Add Remove

Image Classification

Neural Architecture Search

Datasets

ImageNet

Results from the Paper

Edit

Ranked #501 on Image Classification on ImageNet

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Image Classification	ImageNet	GLiT-Bases	Top 1 Accuracy	82.3%	# 501	Compare
			Number of params	96.1M	# 857	Compare
			GFLOPs	17	# 352	Compare
Image Classification	ImageNet	GLiT-Smalls	Top 1 Accuracy	80.5%	# 638	Compare
			Number of params	24.6M	# 585	Compare
			GFLOPs	4.4	# 208	Compare
Image Classification	ImageNet	GLiT-Tinys	Top 1 Accuracy	76.3%	# 847	Compare
			Number of params	7.2M	# 454	Compare
			GFLOPs	1.4	# 128	Compare

Methods

Add Remove

1x1 Convolution • Average Pooling • Batch Normalization • Bottleneck Residual Block • Convolution • Dense Connections • Global Average Pooling • Kaiming Initialization • Layer Normalization • Linear Layer • Max Pooling • Multi-Head Attention • ReLU • Residual Block • Residual Connection • ResNet • Scaled Dot-Product Attention • Softmax • Vision Transformer

Edit Social Preview

GLiT: Neural Architecture Search for Global and Local Image Transformer

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove