Posted in

What Is Multi-Head Attention in Transformers? How It Works

Multi-head attention in Transformers with multiple attention heads and token representations
A visual overview of how multiple attention heads process the same sequence in a Transformer.

Take the sentence “The musician next door practiced all night because she had a concert in the morning.” The word “she” points back to the musician. “Practiced” ties to “night” and “concert” supplies the reason for all of it. A reader follows these links without effort even though each one matters for a different reason. A machine has a harder job. One attention calculation produces a single weighted view of how words relate, and that view has to cover reference, grammar and cause at once. Multi-head attention in Transformers exists because one view tends to blur the others. Instead of asking one question about a sentence, the model asks several at the same time. That idea sits at the center of how Transformer models process text.

What Is Multi-Head Attention?

Attention lets a model decide how much weight every other word deserves when it builds the representation of a given word. Multi-head attention runs several of these attention operations side by side. Each operation is called a head. All heads read the same input, but each views it through its own learned projection and so produces its own account of which words relate to which. Those projections do the real work. Before attention is calculated, the input vectors are multiplied by weight matrices that the model learns during training like the rest of a neural network. Different heads learn different matrices and so operate in different representation spaces. The original paper argues that a single head tends to average its signals together and lose detail, while several heads can keep them apart.

Why Do Transformers Use Multiple Attention Heads?

Consider the sentence “The animal crossed the road because it was tired.” To handle “it” the model has to connect the pronoun to “animal” and not “road”. Swap “tired” for “wide” and the connection flips. Meanwhile “crossed” links the animal to the road and “because” ties the second half of the sentence to the first. A single attention calculation gives each word one set of weights, and that single blend has to carry every one of these relationships. Several heads give the model room to separate them. One head might lean toward the pronoun link while another favors neighboring words and a third follows the verb and its object. This is a tendency and not a rule. Nobody assigns a head its job. Whatever a head ends up doing emerges from training, and researchers who inspect trained models find that some heads follow patterns people can describe while many overlap or resist any tidy label.

How Does Multi-Head Attention Work?

The process begins with input representations. Each token has already been converted into a vector that also carries information about its position in the sequence. In the original Transformer these vectors have 512 dimensions. Every head multiplies them by three learned matrices to produce its own Query, Key and Value versions. The paper used eight heads, and each one projects down to 64 dimensions, so the total cost stays close to that of a single full-size attention operation. Next each head runs its attention calculation on its own projected Queries, Keys and Values. The heads do not wait for one another, which is why the design runs efficiently on modern hardware. Each head returns a new vector for every word. Those vectors are concatenated and passed through a final learned linear projection, and the result moves on to the rest of the Transformer layer.

What Happens Inside Each Attention Head?

Inside a head the three projections play different parts. A Query describes what a word is looking for. A Key describes what each word offers when matched against a Query. A Value holds the information that actually gets passed along once a match is made. Scaled dot-product attention compares every Query with every Key and turns the scores into weights. It then uses those weights to blend the Values.

The original paper writes the calculation as:

Attention(Q,K,V) = softmax(QKᵀ / √dₖ)V

Q, K and V are matrices that stack the Queries, Keys and Values for every word in the sequence. Multiplying Q by the transpose of K produces a table of comparison scores for every pair of words. The term dₖ is the length of the Key vectors, and dividing by its square root keeps the scores from growing too large as those vectors get longer. Without that step the softmax function would saturate and learning would slow down. Softmax then turns the scores into weights that add up to one, and multiplying by V produces the weighted blend.

How Are the Attention Heads Combined?

Each head hands back its own vectors for every word. These are concatenated side by side, so eight heads of 64 dimensions return to the original width of 512. A learned linear projection then mixes the combined vector into a single representation for the next part of the network. The projection matters because the heads were computed separately. It lets the model learn how to weigh and blend what each one found instead of leaving the results as disconnected pieces.

Multi-Head Attention vs Self-Attention

The two terms are easy to confuse because they usually appear together, yet they answer different questions. Where the Queries, Keys and Values come from is the question the self-attention mechanism answers: in self-attention all three are derived from the same sequence. Multi-head attention answers a separate question about how many attention operations run in parallel on that input. Neither concept competes with the other. The encoder of the original Transformer uses multi-head self-attention, which combines both ideas. The decoder also contains a layer where the Queries come from the decoder and the Keys and Values come from the encoder output. That layer is multi-head attention without being self-attention. Running a single head over one sequence would give the reverse case.

Why Does Multi-Head Attention Matter?

Multiple heads allow a Transformer to represent several relationships within the same sequence at once instead of squeezing them into one set of weights. Stack many layers, each with its own heads, and the model can build progressively richer representations of every word in context. This is a large part of why Transformers handle links between distant words so well. The same building block sits inside the language models behind modern chatbots, translation tools and writing assistants. Those systems repeat it across many layers and apply it to text far longer than a single sentence. Recent variants change how heads share projections to save memory, but the basic recipe from the 2017 paper is still recognizable.

Final Thoughts

Multi-head attention works best when pictured as a budget decision. The Transformer splits its representation width across several smaller heads instead of spending it all on one large calculation, and the cost stays about the same. What changes is the variety of patterns the model has room to learn. Nobody tells the heads what to look for, which is both the strength and the puzzle of the design. Training decides how the capacity gets used, and researchers are still studying what individual heads actually do. The safest picture is several independent views of the same sentence that are blended into one result. That blending is what lets a Transformer weigh many relationships at once.

Sources

Krish Shrestha, founder and editor of LegacyVia

Written & Researched by

Krish Shrestha

Krish Shrestha is the founder and editor of LegacyVia. He researches and writes about AI and technology with a focus on understanding how new technologies work and explaining them in a clear, practical way.

Leave a Reply

Your email address will not be published. Required fields are marked *

Krish Shrestha
Founder & Editor Krish Shrestha Founder and editor of LegacyVia, an independent publication covering AI and technology. He researches, writes, and maintains every article on the site.
TRENDING