TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Document Layout Analysis	PubLayNet val	GLAM	Text	0.878	# 13
Document Layout Analysis	PubLayNet val	GLAM	Title	0.800	# 13
Document Layout Analysis	PubLayNet val	GLAM	List	0.862	# 13
Document Layout Analysis	PubLayNet val	GLAM	Table	0.868	# 14
Document Layout Analysis	PubLayNet val	GLAM	Figure	0.206	# 13
Document Layout Analysis	PubLayNet val	GLAM	Overall	0.722	# 13

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/a-graphical-approach-to-document-layout/document-layout-analysis-on-publaynet-val)](https://paperswithcode.com/sota/document-layout-analysis-on-publaynet-val?p=a-graphical-approach-to-document-layout)`

A Graphical Approach to Document Layout Analysis

3 Aug 2023 · Jilin Wang, Michael Krumdick, Baojia Tong, Hamima Halim, Maxim Sokolov, Vadym Barda, Delphine Vendryes, Chris Tanner ·

Document layout analysis (DLA) is the task of detecting the distinct, semantic content within a document and correctly classifying these items into an appropriate category (e.g., text, title, figure). DLA pipelines enable users to convert documents into structured machine-readable formats that can then be used for many useful downstream tasks. Most existing state-of-the-art (SOTA) DLA models represent documents as images, discarding the rich metadata available in electronically generated PDFs. Directly leveraging this metadata, we represent each PDF page as a structured graph and frame the DLA problem as a graph segmentation and classification problem. We introduce the Graph-based Layout Analysis Model (GLAM), a lightweight graph neural network competitive with SOTA models on two challenging DLA datasets - while being an order of magnitude smaller than existing models. In particular, the 4-million parameter GLAM model outperforms the leading 140M+ parameter computer vision-based model on 5 of the 11 classes on the DocLayNet dataset. A simple ensemble of these two models achieves a new state-of-the-art on DocLayNet, increasing mAP from 76.8 to 80.8. Overall, GLAM is over 5 times more efficient than SOTA models, making GLAM a favorable engineering choice for DLA tasks.

PDF Abstract

Code

Add Remove Mark official

ivanstepanovftw/glam

Tasks

Add Remove

Document Layout Analysis

Datasets

PubLayNet

Results from the Paper

Edit

Ranked #13 on Document Layout Analysis on PubLayNet val

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Document Layout Analysis	PubLayNet val	GLAM	Text	0.878	# 13	Compare
			Title	0.800	# 13	Compare
			List	0.862	# 13	Compare
			Table	0.868	# 14	Compare
			Figure	0.206	# 13	Compare
			Overall	0.722	# 13	Compare

Methods

Add Remove

DLA • Graph Neural Network

Edit Social Preview

A Graphical Approach to Document Layout Analysis

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove