🤖TechnologyA Brief History of AI
🏠 Home🌐 中文
A Brief History of AITHE MINDS
🔮

The "Transformer Eight"

Original nameAttention Is All You Need Authors

Google Research Team, Co-Inventors of the Transformer Self-Attention Architecture

The Math Behind the Algorithm · Theoretical Founders
The Transformer ArchitectureThe Self-Attention MechanismThe Paper "Attention Is All You Need"

Who they are

In 2017 eight Google researchers — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin — jointly published the paper "Attention Is All You Need," proposing a wholly new architecture, the "Transformer," that wholly abandoned the previously mainstream recurrent and convolutional neural network structures and relied only on the "self-attention mechanism" to process sequence data. This architecture let a model process in parallel the mutual relations between any two positions in the whole input sequence, rather than having to process step by step in order as a recurrent neural network does, which not only greatly raised training efficiency but, more importantly, let the model more fully capture long-distance context dependencies. The Transformer architecture thereafter became the common underlying architectural basis of almost all large-scale language models — including Google’s own BERT series and OpenAI’s GPT series — profoundly changing the technical path of natural language processing and later of computer vision, protein-structure prediction, and many other fields, one of the most far-reaching single-paper results in deep learning in the 2010s. These eight authors thereafter mostly left Google, in turn joining or founding AI startups including Anthropic, Character.AI, Cohere, and Adept, and this phenomenon of "the paper authors collectively leaving to start companies" also reflects from the side the enormous industrial value this invention itself held.

Primary sourcesVaswani et al., "Attention Is All You Need" (2017, NeurIPS conference paper)

Key stories

A Paper Title That Became a Generation’s Technical Manifesto

The paper title "Attention Is All You Need" directly stated the paper’s most core claim: that processing sequence data no longer needed recurrent or convolutional structure, and attention alone was enough — a claim quite bold at publication, thereafter proven highly farsighted: seven years later, the technical basis of almost all mainstream large language models can be traced to the architecture this paper proposed, and the title "Attention Is All You Need" itself became one of the most cited and joked-about paper titles in the AI field.

Relationships

Echoes today

Below are how modern works borrow or reinterpret this name or story — not the original material. The two differ, so keep them apart.

The phenomenon of "the paper’s authors collectively leaving to start up"Most of these eight authors thereafter left Google for AI startups, a phenomenon that reflects from the side the vast industrial value this invention itself holds, and a topic repeatedly reported by tech media.

Appears in

Curiosity mailCurious about world civilization? Leave your email — we’ll tell you when there’s something worth a look.

Free · unsubscribe anytime · Privacy