🤖TechnologyA Brief History of AI
🏠 Home🌐 中文
A Brief History of AITHE MINDS
📖

Jacob Devlin

Original nameJacob Devlin

American Computer Scientist, Main Author of the BERT Model

The Math Behind the Algorithm · Theoretical Founders
The BERT ModelBidirectional Pretrained Language ModelsNatural Language Understanding Benchmarks

Who they are

Jacob Devlin is an American computer scientist who in 2018, while working at Google, directed the design of the BERT (Bidirectional Encoder Representations from Transformers) model, one of the earliest models to prove the overwhelming advantage on natural-language-understanding tasks of the paradigm of "first pretraining on a large scale on vast unannotated text, then a small amount of fine-tuning for a specific task." BERT used a bidirectional attention mechanism, letting the model, in understanding a word, refer at once to the complete context before and after it, a design that greatly refreshed the previous best on many NLP benchmarks such as question-answering and sentiment analysis, and the "pretrain-then-fine-tune" R&D paradigm was thereafter further developed by later models such as the GPT series, becoming the standard workflow of large-language-model R&D.

Primary sourcesDevlin et al., "BERT: Pre-training of Deep Bidirectional Transformers" (2018)

Key stories

"Bidirectional" Reading Let a Model Truly Understand Context for the First Time

The most core innovation of the BERT model lay in its "bidirectional" training — earlier language models could mostly only process text one-directionally from left to right, while BERT, by randomly masking some words in a sentence and requiring the model to predict them combining the complete context before and after, let the model for the first time truly understand context "bidirectionally," a design that let BERT greatly surpass the previous best on many natural-language-understanding benchmarks, and made the "pretrain-then-fine-tune" R&D paradigm the standard workflow of the whole NLP field thereafter.

Relationships

Echoes today

Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.

"Pre-train plus fine-tune" becomes an industry-standard processThe R&D paradigm BERT proved was thereafter further carried forward by subsequent models such as the GPT series, becoming the standard starting point of the whole large-language-model R&D workflow.

Appears in

Curiosity mailCurious about world civilization? Leave your email — we’ll tell you when there’s something worth a look.

Free · unsubscribe anytime · Privacy