TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK	REMOVE
Document Layout Analysis	RVL-CDIP	VisualWordGrid	FAR	28.7	# 1
Document Layout Analysis	RVL-CDIP	VisualWordGrid	WAR	18.7	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/visualwordgrid-information-extraction-from-1/document-layout-analysis-on-rvl-cdip)](https://paperswithcode.com/sota/document-layout-analysis-on-rvl-cdip?p=visualwordgrid-information-extraction-from-1)`

VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

5 Oct 2020 · Mohamed Kerroumi, Othmane Sayem, Aymen Shabou ·

We introduce a novel approach for scanned document representation to perform field extraction. It allows the simultaneous encoding of the textual, visual and layout information in a 3-axis tensor used as an input to a segmentation model. We improve the recent Chargrid and Wordgrid \cite{chargrid} models in several ways, first by taking into account the visual modality, then by boosting its robustness in regards to small datasets while keeping the inference time low. Our approach is tested on public and private document-image datasets, showing higher performances compared to the recent state-of-the-art methods.

PDF Abstract