🤖TechnologyA Brief History of AI
🏠 Home🌐 中文
A Brief History of AITHE MINDS
🎛️

John Schulman

Original nameJohn Schulman

American Computer Scientist, Co-Founder of OpenAI

The Math Behind the Algorithm · Theoretical Founders
The Proximal Policy Optimization Algorithm (PPO)Reinforcement Learning from Human Feedback (RLHF)The Alignment Training of ChatGPT

Who they are

John Schulman is an American computer scientist who in 2015 took part as a co-founder in founding OpenAI, and proposed the reinforcement-learning algorithm "Proximal Policy Optimization" (PPO), which, for its relatively stable training process and relatively simple implementation, was thereafter widely applied in many reinforcement-learning tasks including robot control. Schulman thereafter directed the combination of reinforcement learning with human feedback, developing the training paradigm of "Reinforcement Learning from Human Feedback" (RLHF), letting a large language model fine-tune by human evaluators’ preference rankings of the quality of model outputs, one of the key technical bases of ChatGPT’s great improvement over the earlier GPT versions in following instructions and reducing harmful output. In 2024 Schulman left OpenAI to join Anthropic, and thereafter turned to founding his own research direction.

Primary sourcesSchulman et al., "Proximal Policy Optimization Algorithms" (2017)OpenAI’s official RLHF technical posts

Key stories

The Key Step to Making a Model "Understand Human Speech"

Early large language models such as GPT-3, though already able to generate coherent text, often gave answers off the point and generated harmful or false content. The RLHF training Schulman directed let human evaluators rank the quality of the model’s several candidate answers, then used this ranking data to train the model to answer in a way more in line with human preference, a relatively plain yet exceedingly effective method widely seen as the key technical basis of ChatGPT’s qualitative change over earlier models in the experience of "understanding human speech and giving useful answers."

Relationships

Echoes today

Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.

RLHF becomes a standard training step of large language modelsThe RLHF method Schulman led the development of thereafter became an indispensable standard step in the training process of almost all mainstream conversational large language models.

Appears in

Curiosity mailCurious about world civilization? Leave your email — we’ll tell you when there’s something worth a look.

Free · unsubscribe anytime · Privacy