Arvind Rajaraman
0 followers · 236 views
on the atlas — 1
- Let's Verify Step by Step1 savers
highlights — 14
we few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this dataset
Let's Verify Step by Stepwe few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this datase
Let's Verify Step by Stepwe few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this dataset
Let's Verify Step by Stepmath-relevant tokens, which we call MathMix
Let's Verify Step by Stepadditional pretraining step, we finetune all models on a dataset of roughly 1.5B
Let's Verify Step by Stepevaluate a reward model by its ability to perform best-of-N search over uniformly sampled solutions
Let's Verify Step by Stepwe are specifically referring to the supervision given to the reward model.
Let's Verify Step by Steprely on human data-labelers to provide process super- vision
Let's Verify Step by Stepuse a more capable base model, we use significantly more human feedback, and we train and test on the more challenging MATH dataset
Let's Verify Step by Stepadvantages relevant to AI alignment: it is easier for humans to interpret, and it more directly rewards models for following a human-endorsed chain-of- thought.
Let's Verify Step by Stepprovides more precise feedback, since it specifies the exact location of any errors
Let's Verify Step by Stepresulting system is only as reliable as the reward model itself.
Let's Verify Step by StepPRM800K, the complete dataset of 800,000 step-level human feedback labels used to train our best reward model
Let's Verify Step by Stepprocess supervision significantly outper- forms outcome supervision for training models to solve problems from the challenging MATH dataset.
Let's Verify Step by Step