When Compound Symmetry Is a Lie: Unequal Variances, Unstructured Correlation, and the Identifiability Ceiling

On this page
Abstract
Nearly every reader arrives with one sentence installed: a random intercept means every pair of repeated measurements from the same person is equally correlated. The sentence is truncated rather than false. A random intercept guarantees only that it contributes the same between-person variance to every pair's covariance; equal correlation needs a second, usually unstated assumption, namely residuals that are uncorrelated and share one variance. This post relaxes that half. With site-specific residual variances the covariance stays 6.25 in every off-diagonal cell while the six correlations run from 0.44 to 0.74. It then counts what richer structures cost: a symmetric four-by-four matrix holds ten distinct quantities, unstructured correlation plus unequal variances spends all ten, and a person variance on top makes eleven parameters for ten quantities.

One Number Per Person: Linear Mixed Models from the Ground Up — Post 6 of 8.
Arriving here without the earlier posts? Here is where we are. When the same patient is measured several times, those measurements are not independent of one another, and a mixed model deals with that by handing every patient one private number — an offset that lifts or lowers that patient's whole profile. Everything the model then assumes about the repeated measurements sits in a single line, $V_i = Z_i G Z_i' + R_i$, whose only two changeable objects are $Z$, what each person is allowed to carry, and $R$, what a single measurement carries on its own. This post is about $R$, and about a sentence most of us were taught in a form that is only half true. A gentler, non-technical tour of all eight posts is in the series guide.
In Post 5 we enriched the left-hand term of $V_i = Z_i G Z_i' + R_i$ by giving $Z$ a second column, and each person stopped carrying an offset and started carrying a whole trajectory. Post 6 leaves $Z$ alone and goes after the right-hand term, $R_i$ — the residual covariance that almost every analysis silently fixes at $\sigma^2 I$ without ever writing the assumption down. We follow that term up the ladder of covariance structures until it hits a hard ceiling, and Post 7 then walks through the only door left in the room.
Every number in this series comes from simulated teaching data; later posts introduce clearly labelled variant simulations. It is not an empirical finding, it describes no real cohort, and it must not be cited as evidence about skin physiology.
The claim under examination
Before the hard part: Two words are about to do most of the work, and confusing them is the commonest error in this whole area. Covariance is how much two measurements move together in their own raw units. Correlation is that same quantity after it has been rescaled by how variable each of the two measurements is on its own. Giving every person a shared offset fixes the first for every pair of measurements from that person. It does not fix the second, and the gap between those two facts is what this post is about.
Nearly every reader arrives here with one sentence already installed: a random intercept means that every pair of repeated measurements from the same person is equally correlated. The sentence is not false so much as truncated, and the missing half is the part that does all the work. What a random intercept guarantees is that it contributes exactly the same quantity, $\tau^2$, to the covariance of every pair of measurements from the same person. Whether that contribution is the whole covariance is a separate matter, and it belongs to the residuals. A pair covaries by $\tau^2 + R_{jk}$, which collapses to $\tau^2$ only when the two residuals are uncorrelated. Equal correlation is a further claim again, and it holds only if a second assumption — one that nobody in the room has stated out loud — happens to be true as well.
The precise statement is this: a random intercept plus mutually independent residuals of equal variance produces compound symmetry. Drop the second half and the first half survives untouched, while compound symmetry quietly dies. In this post we take a magnifier to that second half, count what it costs to relax it, and then find the point at which the covariance matrix has no room left for one more parameter.
Spec B: the same shared offset, four different residual variances
Before the hard part: Some measurements are simply noisier than others. A blood pressure taken in a busy corridor scatters more than one taken after five quiet minutes; an assay run on an ageing analyser scatters more than the same assay on a new one; a measurement taken where the patient cannot hold still scatters more than one taken where they can. The standard model denies all of this by decree, because it hands every measurement the same residual variance. The simulation below relaxes that one assumption and changes nothing else.
Spec A, which carried Posts 1 through 4, gave every one of the four measurement sites the same residual variance: $\sigma^2 = 3.75$ everywhere, sitting under a between-person variance of $\tau^2 = 6.25$, for a total marginal variance of 10.00 at every site. Spec B is a separate simulation, not a re-analysis of the Spec A data — same 80 simulated participants' worth of design, same four sites in the same order, same population means, same $\tau^2 = 6.25$, but each site is now given its own residual variance.
| site | residual variance $\sigma_j^2$ (g²·m⁻⁴·h⁻²) | marginal variance $\tau^2+\sigma_j^2$ (g²·m⁻⁴·h⁻²) |
|---|---|---|
| forearm | 1.75 | 8.00 |
| hand | 2.75 | 9.00 |
| shin | 13.75 | 20.00 |
| back | 3.75 | 10.00 |
Of note, the back was assigned a residual variance of exactly 3.75 — the same number that Spec A used for all four sites — so one and only one site would return the familiar total of 10.00. The shin was assigned the largest residual variance by construction. That is a property of the simulation and nothing else; it is not a claim that anterior shin skin behaves in any particular way. Now suppose a pooled-variance model were fitted to Spec B data. The single $\widehat{\sigma}^{2}$ it reported would land near the average of the four site variances, which in a balanced design is their simple mean of $22/4 = 5.5$, and not at 3.75. The pooled estimate would not be wrong arithmetic; it would simply be answering a question about an average that no site actually has.
The reveal: identical covariance, six different correlations
Before the hard part: Keep those two words apart and the next page is easy. Every pair of measurements from the same person shares the same person-level quantity, so the top of every correlation is the same number. Each correlation is then divided by how variable its own two measurements are, and those bottoms now differ from pair to pair. Same top, different bottoms, different answers.
Now take the magnifier to the off-diagonal cells. The derivation from Post 1 has not changed at all:
\[ \operatorname{Cov}(Y_{ij}, Y_{ik}) = \operatorname{Cov}(b_i + e_{ij},\ b_i + e_{ik}) = \operatorname{Var}(b_i) = \tau^2 = 6.25 \]
Read the right-hand side carefully: there is no $\sigma_j^2$ anywhere in it, and no site index either. The residual variances cancelled out of the covariance before the arithmetic even began, because $e_{ij}$ and $e_{ik}$ are independent of each other and both are independent of $b_i$. Every off-diagonal cell of the Spec B covariance matrix therefore reads 6.25, exactly as it did under Spec A. The shared offset is still shared, and it is still shared by the same amount.
The correlation is a different quantity, because it divides that covariance by two site-specific standard deviations:
\[ \operatorname{Corr}(Y_{ij}, Y_{ik}) = \frac{\tau^2}{\sqrt{(\tau^2+\sigma_j^2)(\tau^2+\sigma_k^2)}} \]
The numerator is a constant. The denominator is not. Work two of the six cells by hand and the mechanism is unmistakable:
\[ \operatorname{Corr}(\text{forearm},\text{hand}) = \frac{6.25}{\sqrt{8.00\times 9.00}} = \frac{6.25}{\sqrt{72}} = \frac{6.25}{8.485} = 0.74 \]
\[ \operatorname{Corr}(\text{forearm},\text{shin}) = \frac{6.25}{\sqrt{8.00\times 20.00}} = \frac{6.25}{\sqrt{160}} = \frac{6.25}{12.649} = 0.49 \]
The same 6.25 enters both calculations and two very different numbers come out. Repeat the division for all six pairs:
| pair | denominator $\sqrt{(\tau^2+\sigma_j^2)(\tau^2+\sigma_k^2)}$ | correlation |
|---|---|---|
| forearm–hand | $\sqrt{8\times9}$ | 0.74 |
| forearm–shin | $\sqrt{8\times20}$ | 0.49 |
| forearm–back | $\sqrt{8\times10}$ | 0.70 |
| hand–shin | $\sqrt{9\times20}$ | 0.47 |
| hand–back | $\sqrt{9\times10}$ | 0.66 |
| shin–back | $\sqrt{20\times10}$ | 0.44 |
One shared offset, one identical covariance in every cell, and six different correlations ranging from 0.44 to 0.74. Compound symmetry is formally false in this simulated dataset, and the random intercept had nothing to do with its death — the culprit is entirely in the denominators.

What varIdent() actually does
Before the hard part: This section is bookkeeping rather than theory, but the bookkeeping is what trips people up when they read the output.
nlmedoes not print four residual standard deviations. It prints one — for whichever level of the factor happens to come first — and then a multiplier for each of the others. The residual standard deviation of any other unit is the printed one multiplied by that unit's own multiplier.
The model that generated Spec B is fitted by attaching one residual variance to each level of site:
# Spec B — site-specific residual variance (Post 6)
fitB <- lme(tewl ~ site * group + age + sex + phototype,
random = ~ 1 | id, weights = varIdent(form = ~ 1 | site), data = skin)
In nlme, varIdent() does not estimate four free variances directly. It estimates one residual standard deviation for the reference stratum — forearm, the first factor level — and then a multiplier for each remaining stratum, with the reference multiplier fixed at 1 for identification. Read the output accordingly. The printed residual standard error belongs to the reference site, the variance-function table holds the three remaining multipliers, and a site's residual standard deviation is the reference value times its own multiplier. Counted as parameters, that is $1 + (k-1) = 4$ variance parameters where the pooled model spent 1. The random-intercept variance $\tau^2$ is still estimated on top and still reported in the VarCorr panel; the value that generated Spec B is 6.25, and any single fit will return an estimate that lands near that value rather than exactly on it.
The person-level covariates in that call — age, sex and phototype — enter the mean structure only. Each of them is constant within a participant, so they may absorb part of the between-person spread that $\tau^2$ would otherwise carry, but they leave the covariance algebra of the previous section untouched.
Three extra parameters is a modest price. The next section explains why the price is not the interesting part.
The other direction: the correlation ladder
Before the hard part: So far we have relaxed how noisy each measurement is allowed to be. The other thing you can relax is how strongly the measurements are allowed to agree with one another. The structures below form a short ladder, from the stingiest assumption to the most generous, and the only currency is how many parameters each rung spends. Higher up the ladder is not better; it is merely more expensive.
Relaxing the variance half of compound symmetry is one move. Relaxing the correlation half is the other, and the structures form a short ladder ordered by how many parameters they spend on a person's $k \times k$ marginal block $V_i$. The count below is a count of that assembled block — everything in it, whether it arrives as $\tau^2$ or as part of $R_i$ — because that is the object the ten slots belong to. For $k = 4$:
| Structure | What it says about a person's four measurements | Parameters of the marginal $V_i$, $k=4$ |
|---|---|---|
| Independence, $V_i = \sigma^2 I_4$ | No within-person dependence at all; the repeated measurements are treated as unrelated rows. | 1 |
| Compound symmetry (exchangeable) | Equal variances and one common correlation for every pair; what Spec A produces. | 2 |
| First-order autoregressive, AR(1) | One common variance and a correlation that decays as a power of the distance between occasions. | 2 |
| Heteroscedastic compound symmetry (Spec B) | One variance per unit, one shared covariance $\tau^2$ for every pair. | 5 |
Unstructured (corSymm + varIdent) | Every variance and every pairwise correlation free to take its own value. | 10 |
AR(1) deserves one caution of its own. It is only meaningful when the repeated dimension is genuinely ordered and the spacing between measurements carries meaning, which is why it belongs to the visit sub-study of Post 5 and not here. The four sites of the running example are a fixed, unordered set of measurement locations — nominal levels rather than positions on a scale. The order forearm, hand, shin, back is a labelling convention. Fitting a structure that assumes hand is "closer" to shin than to back would impose a metric that the design does not contain. Compound symmetry, by contrast, is invariant to how the four units are listed.
If you want to watch the ladder move, the ICC explorer built in Post 2 has a $k$ selector: compound symmetry stays at two parameters however far you push it, whereas the unstructured count is $k(k+1)/2$, which is 3 at $k=2$ and 21 at $k=6$.
The identifiability ceiling, counted in coins
Before the hard part: A dataset can only answer as many questions as it actually contains answers to. That sounds too obvious to need saying, yet it is entirely possible to write down a model that quietly asks one question more than the data holds — and recruiting more patients will never fix it, because the missing answer was never there to be collected. The next few lines count the questions, then count the askers. When the second number exceeds the first, the model is over-parameterised.
Here is the whole argument, and it is arithmetic rather than opinion. A person's marginal covariance matrix is symmetric and $4 \times 4$, so it contains
\[ \frac{k(k+1)}{2} = \frac{4 \times 5}{2} = 10 \]
distinct quantities: four variances on the diagonal and six covariances above it. Ten slots, and not one more. Now count what an unstructured specification spends. varIdent buys four variances. corSymm buys six correlations. That is ten parameters describing ten quantities — the budget is exactly and completely spent.
\[ \underbrace{4}_{\texttt{varIdent}} \;+\; \underbrace{6}_{\texttt{corSymm}} \;+\; \underbrace{1}_{\tau^2} \;=\; 11 \quad \text{parameters for} \quad \frac{k(k+1)}{2} = 10 \quad \text{quantities} \]
Adding a random intercept to that model asks eleven coins to fill ten slots. The eleventh coin is not greedy; it is homeless. This is redundancy in the strict sense, not merely a tight fit, and no amount of data collection will resolve it.

Three splits, one matrix
Before the hard part: Here is the same idea told as a clinical story. Suppose you know that two readings from one patient agree by a certain amount, but you have no way of telling how much of that agreement is "the same person" and how much is "the same day, the same machine, the same observer". You may attribute the agreement whichever way you like, and the data will look identical either way. The table below does precisely that, three times over, and all three attributions reproduce the same fitted matrix.
The intuition behind the ceiling is worth stating slowly, because it explains why the failure is not a numerical accident. Suppose the residual covariance is free to be any number $c$ rather than forced to zero. Then the two things a person's marginal matrix reports are
\[ \operatorname{Cov}(Y_{ij}, Y_{ik}) = \tau^2 + c, \qquad \operatorname{Var}(Y_{ij}) = \tau^2 + \sigma_j^2 \]
Fix the target: the Spec B matrix, with 6.25 in every off-diagonal cell and 8.00, 9.00, 20.00, 10.00 down the diagonal. Any pair $(\tau^2, c)$ satisfying $\tau^2 + c = 6.25$, combined with residual variances $\sigma_j^2 = V_{jj} - \tau^2$, reproduces that matrix cell for cell. The only restriction is that the residual matrix it implies must itself be a legitimate covariance matrix, which confines $\tau^2$ to the interval from 0 to about 7.03. Three examples, all sitting comfortably inside that interval, make the point:
| Split | $\tau^2$ | residual covariance $c$ | residual variances (forearm, hand, shin, back) | reproduces $V$? |
|---|---|---|---|---|
| A | 6.25 | 0.00 | 1.75, 2.75, 13.75, 3.75 | yes |
| B | 3.25 | 3.00 | 4.75, 5.75, 16.75, 6.75 | yes |
| C | 0.00 | 6.25 | 8.00, 9.00, 20.00, 10.00 | yes |
Check split B by hand: $3.25 + 3.00 = 6.25$ in every off-diagonal cell, and $3.25 + 4.75 = 8.00$, $3.25 + 5.75 = 9.00$, $3.25 + 16.75 = 20.00$, $3.25 + 6.75 = 10.00$ down the diagonal. The fitted matrix is identical to split A's. So is split C's, in which there is no person effect at all and the entire dependence has been absorbed into the residual covariance.
Because the three specifications produce the same $V$, they produce the same multivariate normal density, and therefore the same likelihood value at every data point. The likelihood is flat along that direction. The data cannot prefer one split over another, not because the sample is small but because the question has no answer in the data. What the data do fix, and fix with arbitrary precision as the sample grows, is the sum $\tau^2 + c = 6.25$. Where inside the admissible interval that sum is divided is not a question more observations can be asked.
What over-parameterisation looks like in output
Software does not print "this model is not identified". It prints symptoms instead. Most of them are shared by any over-parameterised covariance specification, whichever package produced it, and they are worth recognising on sight. First, the fit may fail to converge, or converge only after the iteration limit is raised, with a warning about the covariance matrix or the optimiser's step. Second, a variance component may settle exactly on its boundary — $\widehat{\tau}^{2}$ collapsing to zero, or the random-effects correlation being estimated at exactly $\pm 1$ — which lme4 reports as a singular fit. Third, refitting with different starting values may move the variance components substantially while leaving the fixed-effect estimates and the fitted covariance matrix essentially unchanged, which is the flat likelihood revealing itself directly. Finally, standard errors for the variance components may be enormous, or may not be computed at all.
None of these should be read as evidence that between-person variability is genuinely absent. A boundary estimate of $\widehat{\tau}^{2} = 0$ in a model that also contains a free residual covariance is the arithmetic of split C, not a finding about persons. The appropriate response is to remove the redundancy from the specification, and not to conclude that the person effect was never there.
If gls has no person effect, does it ignore the repeated measurements?
No. This is the most common misreading of the marginal specification, and it is worth answering flatly. A gls fit with corSymm and varIdent states the within-person covariance matrix directly instead of deriving it from a latent quantity:
# unstructured marginal covariance, no latent person effect (Posts 6, 7)
fitU <- gls(tewl ~ site * group + age + sex + phototype,
correlation = corSymm(form = ~ siteN | id),
weights = varIdent(form = ~ 1 | site), data = skin)
The | id in that formula is doing exactly the job the reader is worried about: it tells the model which rows belong to the same person, and the estimated correlation matrix is applied within those blocks. The dependence is fully accounted for; only its explanation has been declined.
What you lose is real and should be listed honestly. You lose $\tau^2$ as an interpretable estimand, and with it the variance-partition intraclass correlation that Post 2 was built on. The six marginal correlations are still estimated, and estimated freely, but a correlation read straight off the matrix describes the dependence rather than apportioning it between persons and occasions. You also lose the BLUPs, so there is no person-specific prediction and no way to ask which participants sit above their group's profile. What you keep is also real: consistent fixed-effect estimates, standard errors that respect the within-person dependence, and contrasts that are correct under the stated covariance. Whether that trade is acceptable depends entirely on whether your research question is about persons or about populations, which is the whole subject of the next post.
Choosing a covariance model with evidence, not aesthetics
Before the hard part: The rest of this post has shown what is possible. This section is about what is defensible. In one breath: look at the residual spread before you choose anything, compare only models that share the same fixed effects, and settle the comparison before you have seen the fixed-effect results. The paragraph below sets that out as five steps in order.
The ladder invites a bad habit — climbing it until the output looks impressive — so the selection procedure deserves to be stated as a short protocol rather than a preference. First, plot residuals against fitted values and inspect the spread separately for each unit; structured widening, or an obviously fatter band at one site, is the diagnostic that motivates a heteroscedastic residual. Second, compare the residual spread by unit numerically, for instance as standardised-residual summaries per site, so that the impression from the plot is quantified rather than eyeballed. Third, compare nested variance structures by a likelihood-ratio test, with the fixed-effects specification held identical and both models fitted by REML, because REML likelihoods are not comparable across different mean structures. Fourth, treat AIC and BIC as supporting evidence rather than as the decision. Keep the boundary caveat attached to the comparison it actually belongs to: a test of whether a random-effect variance is zero sits on the edge of the parameter space and is conservative, whereas a pooled-versus-varIdent comparison tests variance ratios whose null value lies in the interior, so the ordinary reference distribution applies. Finally, pre-specify the comparison rather than running it after the fixed-effect results are known.
Two things do not belong in this protocol. Selecting a covariance structure by looking at the p-values of the fixed effects it produces is not model checking; it is choosing the answer first. And the tidiness of the output is not evidence of anything at all. The rule that Post 8 leans on is worth adopting now: change the covariance model because the assumptions are contradicted by diagnostics, never because the output looks too tidy.
What this post does not license
The argument above is easy to over-read in three directions, and each of them is worth blocking explicitly.
First, it is not an argument for fitting the richest available structure by default. Ten covariance parameters may well be affordable with 80 simulated participants and four measurements each. The same unstructured specification at $k = 10$ occasions, however, asks for 55, and with a modest number of clusters the estimates could become unstable and the standard errors optimistic. Parsimony is not timidity; it may be the difference between an estimable model and a decorative one.
Second, nothing here says compound symmetry is generally wrong. Spec B was built to be heteroscedastic, so the demonstration is a construction rather than a discovery, and it carries no implication whatsoever about which measurement sites are more variable in real skin. In a design where the units are unordered and their residual spread really does look similar across them, compound symmetry may be both defensible and preferable, precisely because it spends two parameters instead of ten.
Third, the reverse inference is not available either. A fitted model that reports equal correlations across all pairs has not demonstrated that the correlations are equal — it has imposed it. Equality in the output of a compound-symmetry fit is a restatement of the assumption that was handed to the software, and it can only be examined by relaxing the assumption and looking at what changes.
What to do in your own analysis
- State the residual covariance assumption explicitly in the methods, in the same sentence as the random-effects structure. "A random intercept for participant with independent, homoscedastic residuals" is a complete specification; "a mixed model was used to account for repeated measurements" is not.
- Look at residual spread per unit before choosing a structure, and let the diagnostic drive the decision rather than the model-fit statistic.
- If the diagnostics contradict pooling, fit the heteroscedastic version and compare it on identical fixed effects, and consider reporting that comparison as a pre-specified sensitivity analysis rather than as the primary model.
- Do not stack a random intercept on an unstructured residual covariance. Choose one of the two — either explain the dependence with a latent person effect, or state it directly — because the combination asks for a parameter that the covariance matrix has no room to hold.
- Let the estimand decide when the two specifications disagree in convenience. If the question requires $\tau^2$, the ICC, variance partitioning or person-level prediction, the random-effects formulation might be the only one that can answer it, however flexible the marginal alternative looks.
- Report what changed, not that nothing changed. If a sensitivity analysis alters an interval or a secondary contrast, that could be worth a sentence of its own in the results.
Key takeaways
- A random intercept contributes the same $\tau^2 = 6.25$ to the covariance of every pair of repeated measurements from the same person, and when the residuals are uncorrelated that contribution is the entire covariance; what it never forces is the same correlation.
- Compound symmetry requires two assumptions — one shared offset and independent residuals of equal variance — and the second one is the half that is usually left unstated.
- In Spec B the covariance is 6.25 in every off-diagonal cell while the six correlations run from 0.44 to 0.74, because each correlation is divided by a different pair of marginal standard deviations.
- A symmetric $4 \times 4$ covariance matrix contains $k(k+1)/2 = 10$ distinct quantities,
corSymmplusvarIdentspends exactly 10, and adding $\tau^2$ makes 11 parameters for 10 quantities — a redundancy that no sample size can resolve. - Because several $(\tau^2, c)$ splits reproduce the identical fitted matrix — here, every $\tau^2$ between 0 and about 7.03 — the likelihood is flat along that direction, and convergence warnings, boundary estimates and starting-value sensitivity are what that flatness looks like from the outside.
- A
glsfit with a stated covariance matrix accounts for the repeated measurements fully; what it gives up is the interpretation of where the dependence came from, together with the variance-partition ICC and the person-level predictions.
Next
Post 7 takes the marginal road all the way — GLS and GEE, which decline to explain the correlation at all — and shows that for a continuous outcome this costs almost nothing, while for a binary outcome it changes the odds ratio itself.
References
- Verbeke G, Molenberghs G. Linear Mixed Models for Longitudinal Data. New York: Springer; 2000.
- Diggle PJ, Heagerty P, Liang KY, Zeger SL. Analysis of Longitudinal Data. 2nd ed. Oxford: Oxford University Press; 2002.
- Pinheiro JC, Bates DM. Mixed-Effects Models in S and S-PLUS. New York: Springer; 2000.