🤖TechnologyA Brief History of AI
🏠 Home🌐 中文
A Brief History of AITHE MINDS
🧭

Yoshua Bengio

Original nameYoshua Bengio

Canadian Computer Scientist, Early Founder of Neural Network Language Models and the Attention Mechanism

Those Who Conceived That "a Machine Could Think" · Pioneers
Neural Network Language ModelsThe Early Exploration of the Attention MechanismThe Montreal Institute for Learning Algorithms (Mila)

Who they are

Yoshua Bengio (1964– ) is a Canadian computer scientist long teaching at the University of Montreal, one of the "three giants of deep learning." As early as the early 2000s he began systematically studying neural network language models, and the early prototype of the "attention mechanism" his team thereafter proposed solved the long-standing problem of declining quality in the translation of long sentences by machine translation, an idea that thereafter inspired a Google team to propose in 2017 the Transformer architecture wholly based on self-attention, the common underlying foundation of almost all modern large language models. Bengio founded the Montreal Institute for Learning Algorithms (Mila) at the University of Montreal, thereafter one of the largest academic deep-learning research institutions in the world, training a great body of core researchers of the field including Ian Goodfellow. Compared with contemporary scholars such as Hinton and LeCun, Bengio has in recent years put more energy into AI-safety and governance issues, being one of the main initiators and signatories of the 2023 open letter to "pause giant AI experiments for six months," long calling for an international AI-safety regulatory framework. He won the Turing Award in 2018 with Hinton and LeCun.

Primary sourcesBengio et al., "A Neural Probabilistic Language Model" (2003)The 2023 "pause giant AI experiments" open letter

Key stories

From "Can’t Remember Long Sentences" to "Attention Is All You Need"

When Bengio’s team first studied neural network language models, they found that recurrent neural networks, in translating long sentences, would gradually "forget" the key information at the start of the sentence as it grew longer, and translation quality declined. The attention mechanism they thereafter proposed let the model, in generating each word, dynamically "review" the information at all positions of the input sequence rather than relying on the step-by-step passing of a hidden state, an idea a Google team later carried to its extreme, proposing the Transformer architecture that wholly abandoned recurrent structure and relied only on self-attention, its paper titled the thereafter widely circulated "Attention Is All You Need."

Relationships

Echoes today

Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.

From technical founder to safety advocateBengio has in recent years poured much energy into AI-safety and governance topics, and this shift of path from pure technical researcher to public-policy advocate also reflects the general change of mindset of this generation of deep-learning founders.

Appears in

Curiosity mailCurious about world civilization? Leave your email — we’ll tell you when there’s something worth a look.

Free · unsubscribe anytime · Privacy