TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Arithmetic Reasoning	GSM8K	RFT 70B	Accuracy	64.8	# 100
Arithmetic Reasoning	GSM8K	RFT 70B	Parameters (Billion)	79	# 102
Arithmetic Reasoning	GSM8K	RFT 7B	Accuracy	51.2	# 122
Arithmetic Reasoning	GSM8K	RFT 7B	Parameters (Billion)	7	# 10
Arithmetic Reasoning	GSM8K	RFT 13B	Accuracy	55.3	# 114
Arithmetic Reasoning	GSM8K	RFT 13B	Parameters (Billion)	13	# 53

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/scaling-relationship-on-learning-mathematical/arithmetic-reasoning-on-gsm8k)](https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k?p=scaling-relationship-on-learning-mathematical)`

Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

3 Aug 2023 · Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan, Chang Zhou, Jingren Zhou ·

Mathematical reasoning is a challenging task for large language models (LLMs), while the scaling relationship of it with respect to LLM capacity is under-explored. In this paper, we investigate how the pre-training loss, supervised data amount, and augmented data amount influence the reasoning performances of a supervised LLM. We find that pre-training loss is a better indicator of the model's performance than the model's parameter count. We apply supervised fine-tuning (SFT) with different amounts of supervised data and empirically find a log-linear relation between data amount and model performance, and we find better models improve less with enlarged supervised datasets. To augment more data samples for improving model performances without any human effort, we propose to apply Rejection sampling Fine-Tuning (RFT). RFT uses supervised models to generate and collect correct reasoning paths as augmented fine-tuning datasets. We find with augmented samples containing more distinct reasoning paths, RFT improves mathematical reasoning performance more for LLMs. We also find RFT brings more improvement for less performant LLMs. Furthermore, we combine rejection samples from multiple models which push LLaMA-7B to an accuracy of 49.3\% on GSM8K which outperforms the supervised fine-tuning (SFT) accuracy of 35.9\% significantly.

PDF Abstract

Code

Add Remove Mark official

ofa-sys/gsm8k-screl official

161

Tasks

Add Remove

Arithmetic Reasoning

GSM8K

Mathematical Reasoning

Datasets

GSM8K

Results from the Paper

Edit

Ranked #100 on Arithmetic Reasoning on GSM8K (using extra training data)

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Arithmetic Reasoning	GSM8K	RFT 70B	Accuracy	64.8	# 100	Compare
Arithmetic Reasoning	GSM8K	RFT 70B	Parameters (Billion)	79	# 102	Compare
Arithmetic Reasoning	GSM8K	RFT 7B	Accuracy	51.2	# 122	Compare
Arithmetic Reasoning	GSM8K	RFT 7B	Parameters (Billion)	7	# 10	Compare
Arithmetic Reasoning	GSM8K	RFT 13B	Accuracy	55.3	# 114	Compare
Arithmetic Reasoning	GSM8K	RFT 13B	Parameters (Billion)	13	# 53	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove