From "Predicting the Next Word" to a Model That Can Write Essays
The core training goal of the first GPT series was exceedingly plain — merely to predict the most likely next word in a passage of text — but Radford and the team found that as model scale and training-data volume kept growing, this seemingly simple training goal could let the model emerge with complex abilities such as writing coherent essays, answering questions, and even logical reasoning, that previously needed special design to achieve, a finding thereafter one of the most important early empirical supports for the "scaling law" idea.