🤖TechnologyA Brief History of AI
🏠 Home🌐 中文
A Brief History of AITHE MINDS

David Silver

Original nameDavid Silver

British Computer Scientist, Technical Lead of the AlphaGo System

The Math Behind the Algorithm · Theoretical Founders
AlphaGoAlphaZeroThe Engineering of Deep Reinforcement Learning

Who they are

David Silver (1976– ) is a British computer scientist who studied reinforcement learning under Richard Sutton during his doctorate at the University of Alberta, then joined DeepMind and directed the design of the AlphaGo system, combining a deep neural network with the Monte Carlo tree search algorithm to beat the top Korean Go player Lee Sedol four to one in 2016. Silver thereafter developed the AlphaZero system, abandoning the earlier reliance on human game-record data for training and learning wholly from scratch by self-play, at last surpassing the previous strongest systems in Go, chess, and shogi, and this training paradigm of "from scratch, without human prior knowledge" is thereafter widely seen as a forceful practical confirmation of the theoretical view of Sutton’s "bitter lesson."

Primary sourcesSilver et al., "Mastering the game of Go with deep neural networks and tree search" (2016)Silver et al., the AlphaZero paper (2017)

Key stories

Beating a World Champion Without Looking at a Single Human Game Record

The AlphaZero system Silver designed wholly abandoned the way AlphaGo had earlier pretrained on human professional players’ game records, told only the basic rules of Go, then learned on its own by playing itself repeatedly, and after a mere few dozen hours of self-training its play already surpassed the earlier AlphaGo that had needed to train on much human game-record data, a result that shocked not a few professional players who had long studied Go theory and gave forceful empirical support to the view that "a general self-learning method can at last surpass a method relying on human experience."

Relationships

Echoes today

Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.

The empirical proof of the "zero human knowledge" training paradigmThe result of AlphaZero surpassing the previous strongest system by self-play alone, wholly without human game records, was thereafter widely cited as powerful empirical proof of the question of "whether a general self-learning method can surpass methods relying on human experience."

Appears in

Curiosity mailCurious about world civilization? Leave your email — we’ll tell you when there’s something worth a look.

Free · unsubscribe anytime · Privacy