TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Single-step retrosynthesis	USPTO-50k	Augmented Transformer (ATx100)	Top-1 accuracy	53.5	# 6
Single-step retrosynthesis	USPTO-50k	Augmented Transformer (ATx100)	Top-5 accuracy	81.0	# 4
Single-step retrosynthesis	USPTO-50k	Augmented Transformer (ATx100)	Top-10 accuracy	85.7	# 6

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/augmented-transformer-achieves-97-and-85-for/single-step-retrosynthesis-on-uspto-50k)](https://paperswithcode.com/sota/single-step-retrosynthesis-on-uspto-50k?p=augmented-transformer-achieves-97-and-85-for)`

State-of-the-Art Augmented NLP Transformer models for direct and single-step retrosynthesis

5 Mar 2020 · Igor V. Tetko, Pavel Karpov, Ruud Van Deursen, Guillaume Godin ·

We investigated the effect of different training scenarios on predicting the (retro)synthesis of chemical compounds using a text-like representation of chemical reactions (SMILES) and Natural Language Processing neural network Transformer architecture. We showed that data augmentation, which is a powerful method used in image processing, eliminated the effect of data memorization by neural networks, and improved their performance for the prediction of new sequences. This effect was observed when augmentation was used simultaneously for input and the target data simultaneously. The top-5 accuracy was 84.8% for the prediction of the largest fragment (thus identifying principal transformation for classical retro-synthesis) for the USPTO-50k test dataset and was achieved by a combination of SMILES augmentation and a beam search algorithm. The same approach provided significantly better results for the prediction of direct reactions from the single-step USPTO-MIT test set. Our model achieved 90.6% top-1 and 96.1% top-5 accuracy for its challenging mixed set and 97% top-5 accuracy for the USPTO-MIT separated set. It also significantly improved results for USPTO-full set single-step retrosynthesis for both top-1 and top-10 accuracies. The appearance frequency of the most abundantly generated SMILES was well correlated with the prediction outcome and can be used as a measure of the quality of reaction prediction.

PDF Abstract

Code

Add Remove Mark official

bigchem/synthesis official

Tasks

Add Remove

Data Augmentation

Memorization

Retrosynthesis

Single-step retrosynthesis

Datasets

USPTO-50k

Results from the Paper

Edit

Ranked #6 on Single-step retrosynthesis on USPTO-50k

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Single-step retrosynthesis	USPTO-50k	Augmented Transformer (ATx100)	Top-1 accuracy	53.5	# 6	Compare
			Top-5 accuracy	81.0	# 4	Compare
			Top-10 accuracy	85.7	# 6	Compare

Methods

Add Remove

Absolute Position Encodings • Adam • BPE • Dense Connections • Dropout • Label Smoothing • Layer Normalization • Linear Layer • Multi-Head Attention • Position-Wise Feed-Forward Layer • Residual Connection • Scaled Dot-Product Attention • Softmax • Transformer

Edit Social Preview

State-of-the-Art Augmented NLP Transformer models for direct and single-step retrosynthesis

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove