Letting a Machine Play Itself — an Idea Older Than It Seems
To keep his checkers program improving, Samuel designed a training mechanism letting the program play against itself repeatedly and adjust the weights of its position-evaluation function by the game results, and this core idea of "self-play" is of one lineage in principle with the training method by which DeepMind’s AlphaGo Zero, nearly sixty years later, learned Go from scratch by playing itself. Samuel is thereby often seen as one of the earliest practitioners of the research direction of reinforcement learning, though "reinforcement learning" as a formal disciplinary name was not systematically established until decades later.