Within and between variation

Longitudinal data is exciting, because it can tell us how individuals change in time and it also lets us understand how much they differ in the first place. That said, it is more complex and harder to understand than data collected once. One of the key concepts you need to learn is the difference between within-variation and between-variation.

Start with a simple graph

Let us look at how individuals change in time based on some simulated data. Here we will focus on two examples: income and satisfaction. Hover or tap any point to follow that individual.

Twelve individuals over four years
Monthly income
Individual trajectories

Calculate each individual’s average

One way to understand this data is to calculate the average for each individual. We have four values for each one, so we can simply take their average. That gives us an idea of what usually happens with this individual. Here we follow just three people: the dashed line is that person’s average, and each orange stem is one year’s distance from it.

Two sources of variation

If we take the difference between the individual average and the observed scores, we can separate two sources of variation. The between variation tells us how different people are from each other. The within variation tells us how different a particular time point is compared to that individual’s own average.

How to calculate their relative importance?

One interesting way to think about this is to ask how much variation there is at each level. What proportion is between, and what proportion is within? That tells us whether something is relatively stable in time or changes a lot, and whether people are very different from each other.

We measure the amount of variation in each, and we represent the between variation as σ²u while the within one is σ²e. For these, a bigger number simply means a wider spread.

Income has more between-person variation, while satisfaction has more within-variation. Drag the slider under the figure and you will see that, depending on where you are, you get different combinations of between and within.

Testing it with a model

We can test this formally by running a statistical model. There are a couple of ways to do it, but if we run something like a multilevel model we can calculate the amount of between- and within-variation. Simulated results based on the variances above lead to these results:

ParameterEstimate

The coefficient that captures the proportion of variation at each level is called the ICC, or intra-class correlation, and it tells us the proportion of variation that is between compared to within.

Strategies for dealing with this complexity

Get rid of the between variation altogether
We can treat the data as a kind of quasi-experimental design, throw away all the between variation and keep only the within variation. This is what the fixed effects model does.
Or model both sources at once
The random effects model, the multilevel model for change and the latent growth model all model both sources of variation, and let us see both how people change over time and how they differ from each other in their rates of change. They can also be added to other models, like the cross-lagged model, to improve interpretation and assumptions.

Want to go further?

These ideas are the foundation of every longitudinal method. To learn how to apply them to your own data, take a course or read the book.

Browse the courses Get the book