TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Scene Text Recognition	ICDAR2013	SVTR-L (Large)	Accuracy	97.2	# 16
Scene Text Recognition	ICDAR2013	SVTR-B (Base)	Accuracy	97.1	# 17
Scene Text Recognition	ICDAR2013	SVTR-S (Small)	Accuracy	95.7	# 21
Scene Text Recognition	ICDAR2013	SVTR-T (Tiny)	Accuracy	96.3	# 20

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/svtr-scene-text-recognition-with-a-single/scene-text-recognition-on-icdar2013)](https://paperswithcode.com/sota/scene-text-recognition-on-icdar2013?p=svtr-scene-text-recognition-with-a-single)`

SVTR: Scene Text Recognition with a Single Visual Model

30 Apr 2022 · Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, Yu-Gang Jiang ·

Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text recognition within the patch-wise image tokenization framework, which dispenses with the sequential modeling entirely. The method, termed SVTR, firstly decomposes an image text into small patches named character components. Afterward, hierarchical stages are recurrently carried out by component-level mixing, merging and/or combining. Global and local mixing blocks are devised to perceive the inter-character and intra-character patterns, leading to a multi-grained character component perception. Thus, characters are recognized by a simple linear prediction. Experimental results on both English and Chinese scene text recognition tasks demonstrate the effectiveness of SVTR. SVTR-L (Large) achieves highly competitive accuracy in English and outperforms existing methods by a large margin in Chinese, while running faster. In addition, SVTR-T (Tiny) is an effective and much smaller model, which shows appealing speed at inference. The code is publicly available at https://github.com/PaddlePaddle/PaddleOCR.

PDF Abstract

Code

Add Remove Mark official

PaddlePaddle/PaddleOCR official

38,721

mindspore-lab/mindocr

162

Tasks

Add Remove

Scene Text Recognition

Datasets

ICDAR 2013

Results from the Paper

Edit

Ranked #16 on Scene Text Recognition on ICDAR2013

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Scene Text Recognition	ICDAR2013	SVTR-L (Large)	Accuracy	97.2	# 16	Compare
Scene Text Recognition	ICDAR2013	SVTR-B (Base)	Accuracy	97.1	# 17	Compare
Scene Text Recognition	ICDAR2013	SVTR-S (Small)	Accuracy	95.7	# 21	Compare
Scene Text Recognition	ICDAR2013	SVTR-T (Tiny)	Accuracy	96.3	# 20	Compare

Methods

Add Remove

SPEED

Edit Social Preview

SVTR: Scene Text Recognition with a Single Visual Model

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove