Archetypal analysis, visually

2026-10-07

Think of a paint shop. It does not sell thousands of colours: it sells a few pure ones, and every other colour is a mix of them. Archetypal analysis (AA) does the same with data. It looks for a handful of extreme, “pure” cases, the archetypes, and describes everything else as a mixture of them.

Instead of asking “which group does this belong to?”, it asks “how much of each extreme does this have?” No maths is needed to follow the ideas below; the formulas are at the end for the curious.

A triangle with a pure blue, a pure green and a pure orange at its corners, whose inside blends smoothly between the three.

Paint mixing. Every colour inside the triangle is a mix of the three pure colours at its corners.

Archetypal analysis looks for the pure “colours” hidden in a dataset. Let us see how it works with a simple example.

A cloud of athletes

Imagine we measured 150 athletes and drew each one as a dot. Athletes who are alike end up close together. (In statistics, each athlete is one observation.)

To be able to draw them, each athlete is described here by just two numbers, say two scores that summarise their abilities. Real data usually have many more, dozens or even thousands of measurements, but the ideas below are exactly the same; we just cannot draw them. (The data here are made up, just to have something to look at.)

A cloud of grey dots with a roughly triangular shape.

150 athletes, one dot each. Similar athletes are close together.

Look at the shape of the cloud: it sticks out in three directions. Near the tip of each one live athletes who are extreme in some way: very fast over a short distance, very enduring, very strong. Let us give those tips names.

The same cloud, with a coloured dot at each tip labelled Sprinter, Marathoner and Weightlifter, joined by a triangle.

The three archetypes: the extreme cases at the tips of the cloud. The triangle they form contains most athletes.

Nobody has to be exactly a sprinter, a marathoner or a weightlifter. A decathlete, for example, is a bit of all three. (Some athletes fall outside the triangle; we will get to them in a moment.) That is the whole idea: the extremes are the building blocks, and every athlete is a mixture of them. This is the first rule of archetypal analysis.

Every athlete is a mixture

To say how much of each archetype an athlete has, we use three percentages that add up to 100%. They are the athlete’s weights. Watch one athlete walk around the triangle: the lines go to the archetypes, the bar chart shows the weights, and the dot takes the blend of the archetypes’ colours that matches them.

At an archetype, the athlete is 100% that archetype. Along an edge, only two archetypes matter. In the middle there is a bit of everything. The closer to an archetype, the larger its share.

An animation of a dot moving inside the triangle of archetypes, with three bars showing how much of each archetype it contains.

One athlete moves through the triangle. The bars show how much of each archetype the athlete has.

Athletes outside the triangle

Look at the cloud again: some athletes lie outside the triangle. A mixture of the archetypes can only land inside it, so these athletes cannot be written exactly as a mixture.

AA deals with them in the simplest possible way: each athlete is described by the closest point of the triangle, the mixture that looks most like them. Watch the athletes outside the triangle slide to their closest mixture. The thin line they leave behind is what the description gets wrong.

An animation in which the dark dots outside the triangle slide to the nearest point of its border, leaving a thin line behind, and then go back.

Athletes outside the triangle (dark dots) are described by the closest mixture of the archetypes, on its edge. The line that joins each athlete to its description is the part that the archetypes fail to explain.

Two things follow from this. First, the leftover gap is exactly the error we will use later to find the archetypes and to decide how many we need. Second, since archetypes are built from real athletes, the triangle can never reach beyond the data: the most extreme athletes will usually be a little outside it. That is the price of keeping the archetypes realistic.

Archetypes come from real data

There is a catch. If we were free to put the three archetypes anywhere, we could invent “athletes” that nobody has ever seen. AA avoids this with a second rule: each archetype must itself be a mixture of real athletes.

So archetypes are extreme, but never imaginary: they sit at the edge of the data, built from observations that actually exist. You can always go and look at the athletes behind an archetype, which makes the result easy to interpret. In the figure below, the zoom visits each archetype in turn.

An animation with two panels. The left one shows the cloud with a small dashed square that moves from one archetype to the next. The right one zooms into that square, where a large coloured dot, the archetype, is joined by lines of different thickness to a few real athletes.

Each archetype is a weighted average of a few real athletes: the thicker the line, the larger the weight. The Sprinter coincides with one real athlete. Right: a zoom that visits each archetype in turn.

Archetypes are not averages

A more familiar way to summarise data is clustering: split the athletes into groups of similar athletes, and describe each group by its average, the centroid. A centroid is the typical member of a group, so it lies in the middle of the cloud. The “average athlete” of a group is average at everything, and may look like nobody in particular. An archetype is the opposite: the extreme member, on the edge.

The closest relative of AA is fuzzy c-means, a “soft” clustering method. Like AA, it gives every athlete percentages that add up to 100%, one for each group. But the percentages mean something different. In c-means they say how close the athlete is to each typical case. In AA they say how much of each extreme the athlete is made of.

Two panels with the same cloud. On the left, three squares inside the cloud, the centres found by fuzzy c-means. On the right, the three archetypes at the tips.

Fuzzy c-means (left) and archetypal analysis (right) on the same athletes. The centres (squares) sit inside the cloud and the archetypes (dots) on its edge. The triangle of the centres leaves most athletes outside; the triangle of the archetypes covers most of them.

Put simply: centres answer “what does a typical athlete look like?”, while archetypes answer “what are the extremes?”

How are archetypes found?

Nobody knows the archetypes in advance, so the computer finds them by trial and improvement:

  1. Guess. Start with three random athletes as archetypes.
  2. Check. Try to rebuild every athlete as a mixture of the archetypes. The gap between each athlete and its rebuilt version is the error.
  3. Improve. Nudge the archetypes, and the weights, in the direction that makes the error smaller.
  4. Repeat until the error stops improving.

It is like walking downhill in the fog: you cannot see the valley, but at every step you can feel which way the ground slopes.

To get a single number for all the gaps together, we compare them with how spread out the athletes are. We call the result the unexplained variation: at 0% the archetypes rebuild everyone perfectly, and at 100% they do no better than describing everybody as “the average athlete”. A bad first guess can be even worse than that, so the curve starts above 100%. Watch the archetypes (left) leave their random starting points and slide towards the tips while the unexplained variation (right) drops.

An animation of three coloured dots moving from the middle of a cloud to its tips, next to a curve of the unexplained variation going down.

Fitting archetypes. Left: they start at three random athletes and move towards the extremes. Right: the unexplained variation drops quickly at first, then flattens out as the archetypes settle.

How many archetypes?

We have used three because that is what this cloud looks like. With real data we have to choose. Too few archetypes cannot describe the data well; too many make the description longer and harder to read, and eventually they stop being “extreme” and start to pick up individual odd athletes.

A common recipe is to try several values and watch how much variation is left unexplained.

Four panels with the same cloud and different numbers of archetypes, each with the unexplained variation below.

The same athletes described with 2, 3, 4 and 5 archetypes. With two, the line cannot reach most of the cloud (a lot of variation is left unexplained). With three, the triangle covers most athletes. Beyond that, the extra archetypes barely help.

Plotting the final unexplained variation against the number of archetypes makes the choice easier. We look for the elbow: the curve looks like a bent arm, and the elbow is where it bends, the point after which adding archetypes stops paying off.

A line chart of the unexplained variation against the number of archetypes, dropping steeply until three and flat afterwards.

Unexplained variation against the number of archetypes. It falls sharply up to three archetypes and hardly changes afterwards: the elbow is at three. In real data the elbow is rarely this sharp, so the choice also depends on what you want to interpret.

Where is it useful?

AA shines whenever the extremes tell you more than the averages. It was introduced by Adele Cutler and Leo Breiman (1994), and it fits questions like these:

Things to keep in mind

In short

The maths, for the curious

Put the nn observations, each described by pp numbers, in a matrix X∈ℝn×pX \in \mathbb{R}^{n \times p}, and the kk archetypes in the rows of Z∈ℝk×pZ \in \mathbb{R}^{k \times p}. The weights of each observation form a matrix A∈ℝn×kA \in \mathbb{R}^{n \times k}, and the data are approximated by

X≈AZ. X \approx A Z .

The second rule, that archetypes are mixtures of real observations, reads Z=BXZ = B X, where B∈ℝk×nB \in \mathbb{R}^{k \times n} holds the weights. Fitting the model means solving

minA,B‖X−ABX‖F2 \min_{A,\,B}\ \lVert X - A B X \rVert_F^2

subject to A𝟏=𝟏A\mathbf{1} = \mathbf{1}, B𝟏=𝟏B\mathbf{1} = \mathbf{1} and A,B≥0A, B \ge 0.

AA says how much of each archetype every observation has; BB says which observations each archetype is made of. The constraints are exactly the two rules above: every row of weights is non-negative and adds up to one, so every mixture is an average. The unexplained variation shown in the figures is ‖X−ABX‖F2\lVert X - ABX\rVert_F^2 divided by the total variation of the data, ‖X−X‾‖F2\lVert X - \bar{X}\rVert_F^2, that is, 1−R21 - R^2. Distances are squared, so a value of 3% corresponds to a typical gap of about 0.03≈17%\sqrt{0.03} \approx 17\% of the spread of the data.

References

Cutler, Adele, and Leo Breiman. 1994. “Archetypal Analysis.” Technometrics 36 (4): 338–47. https://doi.org/10.1080/00401706.1994.10485840.

The animations on this page are looping SVGs made with svganim.