TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Math Word Problem Solving	MATH	GPT-3 2.7B	Accuracy	2.9	# 108
Math Word Problem Solving	MATH	GPT-3 2.7B	Parameters (Billions)	2.7	# 86
Math Word Problem Solving	MATH	GPT-2 (1.5B)	Accuracy	6.9	# 96
Math Word Problem Solving	MATH	GPT-2 (1.5B)	Parameters (Billions)	1.5	# 87
Math Word Problem Solving	MATH	GPT-2 (0.1B)	Accuracy	5.4	# 102
Math Word Problem Solving	MATH	GPT-2 (0.1B)	Parameters (Billions)	0.1	# 90
Math Word Problem Solving	MATH	GPT-2 (0.3B)	Accuracy	6.2	# 99
Math Word Problem Solving	MATH	GPT-2 (0.3B)	Parameters (Billions)	0.3	# 89
Math Word Problem Solving	MATH	GPT-2 (0.7B)	Accuracy	6.4	# 98
Math Word Problem Solving	MATH	GPT-2 (0.7B)	Parameters (Billions)	0.7	# 88
Math Word Problem Solving	MATH	GPT-3-175B (few-shot)	Accuracy	5.2	# 103
Math Word Problem Solving	MATH	GPT-3-175B (few-shot)	Parameters (Billions)	175	# 5
Math Word Problem Solving	MATH	GPT-3 13B	Accuracy	5.6	# 100
Math Word Problem Solving	MATH	GPT-3 13B	Parameters (Billions)	13	# 39
Math Word Problem Solving	MATH	GPT-3-13B (few-shot)	Accuracy	3.0	# 107
Math Word Problem Solving	MATH	GPT-3-13B (few-shot)	Parameters (Billions)	13	# 39

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/measuring-mathematical-problem-solving-with/math-word-problem-solving-on-math)](https://paperswithcode.com/sota/math-word-problem-solving-on-math?p=measuring-mathematical-problem-solving-with)`

Measuring Mathematical Problem Solving With the MATH Dataset

5 Mar 2021 · Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, Jacob Steinhardt ·

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.

PDF Abstract

Code

Add Remove Mark official

hendrycks/math official

724

openai/minif2f

257

facebookresearch/minif2f

rah4927/lean-dojo-mew

Tasks

Add Remove

Math

Mathematical Reasoning

Math Word Problem Solving

Text Generation

Datasets

Introduced in the Paper:

MATH

Used in the Paper:

Mathematics Dataset

HOList

Results from the Paper

Edit

Ranked #96 on Math Word Problem Solving on MATH

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Math Word Problem Solving	MATH	GPT-3 2.7B	Accuracy	2.9	# 108	Compare
Math Word Problem Solving	MATH	GPT-3 2.7B	Parameters (Billions)	2.7	# 86	Compare
Math Word Problem Solving	MATH	GPT-2 (1.5B)	Accuracy	6.9	# 96	Compare
Math Word Problem Solving	MATH	GPT-2 (1.5B)	Parameters (Billions)	1.5	# 87	Compare
Math Word Problem Solving	MATH	GPT-2 (0.1B)	Accuracy	5.4	# 102	Compare
Math Word Problem Solving	MATH	GPT-2 (0.1B)	Parameters (Billions)	0.1	# 90	Compare
Math Word Problem Solving	MATH	GPT-2 (0.3B)	Accuracy	6.2	# 99	Compare
Math Word Problem Solving	MATH	GPT-2 (0.3B)	Parameters (Billions)	0.3	# 89	Compare
Math Word Problem Solving	MATH	GPT-2 (0.7B)	Accuracy	6.4	# 98	Compare
Math Word Problem Solving	MATH	GPT-2 (0.7B)	Parameters (Billions)	0.7	# 88	Compare
Math Word Problem Solving	MATH	GPT-3-175B (few-shot)	Accuracy	5.2	# 103	Compare
Math Word Problem Solving	MATH	GPT-3-175B (few-shot)	Parameters (Billions)	175	# 5	Compare
Math Word Problem Solving	MATH	GPT-3 13B	Accuracy	5.6	# 100	Compare
Math Word Problem Solving	MATH	GPT-3 13B	Parameters (Billions)	13	# 39	Compare
Math Word Problem Solving	MATH	GPT-3-13B (few-shot)	Accuracy	3.0	# 107	Compare
Math Word Problem Solving	MATH	GPT-3-13B (few-shot)	Parameters (Billions)	13	# 39	Compare

Methods

Add Remove

Absolute Position Encodings • Adam • BPE • Dense Connections • Dropout • GELU • Label Smoothing • Layer Normalization • Linear Layer • Multi-Head Attention • Position-Wise Feed-Forward Layer • Residual Connection • Scaled Dot-Product Attention • Softmax • Transformer

Edit Social Preview

Measuring Mathematical Problem Solving With the MATH Dataset

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove