flâneur

Arvind Rajaraman

0 followers · 236 views

on the atlas — 1

highlights — 14

  • we few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this dataset
    Let's Verify Step by Step
  • we few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this datase
    Let's Verify Step by Step
  • we few-shot generate solutions to MATH training problems, filter to those that reach the correct final answer, and finetune the base model on this dataset
    Let's Verify Step by Step
  • math-relevant tokens, which we call MathMix
    Let's Verify Step by Step
  • additional pretraining step, we finetune all models on a dataset of roughly 1.5B
    Let's Verify Step by Step
  • evaluate a reward model by its ability to perform best-of-N search over uniformly sampled solutions
    Let's Verify Step by Step
  • we are specifically referring to the supervision given to the reward model.
    Let's Verify Step by Step
  • rely on human data-labelers to provide process super- vision
    Let's Verify Step by Step
  • use a more capable base model, we use significantly more human feedback, and we train and test on the more challenging MATH dataset
    Let's Verify Step by Step
  • advantages relevant to AI alignment: it is easier for humans to interpret, and it more directly rewards models for following a human-endorsed chain-of- thought.
    Let's Verify Step by Step
  • provides more precise feedback, since it specifies the exact location of any errors
    Let's Verify Step by Step
  • resulting system is only as reliable as the reward model itself.
    Let's Verify Step by Step
  • PRM800K, the complete dataset of 800,000 step-level human feedback labels used to train our best reward model
    Let's Verify Step by Step
  • process supervision significantly outper- forms outcome supervision for training models to solve problems from the challenging MATH dataset.
    Let's Verify Step by Step