TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Scene Text Recognition	ICDAR 2003	CSTR	Accuracy	94.8	# 6
Scene Text Recognition	ICDAR2013	CSTR	Accuracy	93.2	# 25
Scene Text Recognition	ICDAR2015	CSTR	Accuracy	81.6	# 15
Scene Text Recognition	SVT	CSTR	Accuracy	90.6	# 22

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/cstr-a-classification-perspective-on-scene/scene-text-recognition-on-icdar-2003)](https://paperswithcode.com/sota/scene-text-recognition-on-icdar-2003?p=cstr-a-classification-perspective-on-scene)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/cstr-a-classification-perspective-on-scene/scene-text-recognition-on-icdar2015)](https://paperswithcode.com/sota/scene-text-recognition-on-icdar2015?p=cstr-a-classification-perspective-on-scene)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/cstr-a-classification-perspective-on-scene/scene-text-recognition-on-svt)](https://paperswithcode.com/sota/scene-text-recognition-on-svt?p=cstr-a-classification-perspective-on-scene)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/cstr-a-classification-perspective-on-scene/scene-text-recognition-on-icdar2013)](https://paperswithcode.com/sota/scene-text-recognition-on-icdar2013?p=cstr-a-classification-perspective-on-scene)`

Revisiting Classification Perspective on Scene Text Recognition

22 Feb 2021 · Hongxiang Cai, Jun Sun, Yichao Xiong ·

The prevalent perspectives of scene text recognition are from sequence to sequence (seq2seq) and segmentation. Nevertheless, the former is composed of many components which makes implementation and deployment complicated, while the latter requires character level annotations that is expensive. In this paper, we revisit classification perspective that models scene text recognition as an image classification problem. Classification perspective has a simple pipeline and only needs word level annotations. We revive classification perspective by devising a scene text recognition model named as CSTR, which performs as well as methods from other perspectives. The CSTR model consists of CPNet (classification perspective network) and SPPN (separated conv with global average pooling prediction network). CSTR is as simple as image classification model like ResNet \cite{he2016deep} which makes it easy to implement and deploy. We demonstrate the effectiveness of the classification perspective on scene text recognition with extensive experiments. Futhermore, CSTR achieves nearly state-of-the-art performance on six public benchmarks including regular text, irregular text. The code will be available at https://github.com/Media-Smart/vedastr.

PDF Abstract

Code

Add Remove Mark official

Media-Smart/vedastr official

531

Tasks

Add Remove

Classification

General Classification

Image Classification

Multi-class Classification

Scene Text Recognition

Datasets

ICDAR 2013

ICDAR 2003

SVT

Results from the Paper

Edit

Ranked #6 on Scene Text Recognition on ICDAR 2003

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Scene Text Recognition	ICDAR 2003	CSTR	Accuracy	94.8	# 6	Compare
Scene Text Recognition	ICDAR2013	CSTR	Accuracy	93.2	# 25	Compare
Scene Text Recognition	ICDAR2015	CSTR	Accuracy	81.6	# 15	Compare
Scene Text Recognition	SVT	CSTR	Accuracy	90.6	# 22	Compare

Methods

Add Remove

1x1 Convolution • Average Pooling • Batch Normalization • Bottleneck Residual Block • Convolution • Global Average Pooling • Kaiming Initialization • Max Pooling • ReLU • Residual Block • Residual Connection • ResNet

Edit Social Preview

Revisiting Classification Perspective on Scene Text Recognition

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove