A Paper That Became the Theoretical Starting Point of ChatGPT’s "Obedience"
The "training a model from human preference" method Christiano proposed in 2017 had at first a relatively simple experimental scene, involving only some basic simulated control tasks, but this core idea of "letting humans rank the model’s outputs by preference and then adjusting the model’s behavior accordingly" was thereafter further developed and extended by researchers such as John Schulman to the alignment training of large language models, becoming one of the key theoretical sources behind ChatGPT’s better ability to "understand human speech" over earlier models.