The Key Step to Making a Model "Understand Human Speech"
Early large language models such as GPT-3, though already able to generate coherent text, often gave answers off the point and generated harmful or false content. The RLHF training Schulman directed let human evaluators rank the quality of the model’s several candidate answers, then used this ranking data to train the model to answer in a way more in line with human preference, a relatively plain yet exceedingly effective method widely seen as the key technical basis of ChatGPT’s qualitative change over earlier models in the experience of "understanding human speech and giving useful answers."