⚫David Silver
Original nameDavid Silver
British Computer Scientist, Technical Lead of the AlphaGo System
The Math Behind the Algorithm · Theoretical Founders
AlphaGoAlphaZeroThe Engineering of Deep Reinforcement Learning
Who they are
David Silver (1976– ) is a British computer scientist who studied reinforcement learning under Richard Sutton during his doctorate at the University of Alberta, then joined DeepMind and directed the design of the AlphaGo system, combining a deep neural network with the Monte Carlo tree search algorithm to beat the top Korean Go player Lee Sedol four to one in 2016. Silver thereafter developed the AlphaZero system, abandoning the earlier reliance on human game-record data for training and learning wholly from scratch by self-play, at last surpassing the previous strongest systems in Go, chess, and shogi, and this training paradigm of "from scratch, without human prior knowledge" is thereafter widely seen as a forceful practical confirmation of the theoretical view of Sutton’s "bitter lesson."
Primary sourcesSilver et al., "Mastering the game of Go with deep neural networks and tree search" (2016)Silver et al., the AlphaZero paper (2017)
Key stories
Beating a World Champion Without Looking at a Single Human Game Record
The AlphaZero system Silver designed wholly abandoned the way AlphaGo had earlier pretrained on human professional players’ game records, told only the basic rules of Go, then learned on its own by playing itself repeatedly, and after a mere few dozen hours of self-training its play already surpassed the earlier AlphaGo that had needed to train on much human game-record data, a result that shocked not a few professional players who had long studied Go theory and gave forceful empirical support to the view that "a general self-learning method can at last surpass a method relying on human experience."
Echoes today
Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.
The empirical proof of the "zero human knowledge" training paradigmThe result of AlphaZero surpassing the previous strongest system by self-play alone, wholly without human game records, was thereafter widely cited as powerful empirical proof of the question of "whether a general self-learning method can surpass methods relying on human experience."