A group can improve after an intervention without showing that the intervention caused the improvement. The comparator may change too. In a controlled experiment, the relevant causal comparison usually concerns how the groups differ, not whether each group separately crossed a significance threshold against its own baseline.
Identify what each calculation is comparing
Bland and Altman’s methodological article demonstrates why separate within-group significance tests can mislead readers of randomised trials. The intervention question requires comparing groups directly; comparing their p-value labels does not perform that test.Bland and Altman — Comparisons within randomised groups can be very misleading (opens in a new tab)
A within-group change compares a group with itself at an earlier time. A between-group comparison contrasts outcomes or changes between conditions. Both calculations can describe data, but they answer different questions.
| Calculation | Question answered |
|---|---|
| A after minus A before | How did group A change? |
| B after minus B before | How did group B change? |
| Change in A minus change in B | How did the changes differ? |
| Adjusted follow-up comparison | How did outcomes differ under the stated model? |
The appropriate analysis depends on the design, including randomisation, pairing and baseline information. The table identifies questions rather than prescribing one universal analysis method.
See why improvement in one group is incomplete evidence
In an original hypothetical experiment, group A changes from a mean of 40 to 50 signal units. Group B changes from 42 to 51. The observed changes are +10 and +9 units respectively.
The difference between changes is therefore +1 unit. Reporting only A’s ten-unit increase would omit the substantial change also seen in the comparator. The one-unit contrast is the descriptive difference between those changes, not a significance result.
Suppose instead that both groups change by ten units. Each group might show clear evidence of change over time, while the between-group change contrast is zero. A common time-related influence can affect both groups.
Do not compare significance labels as if they were effects
Imagine that the +10 change in A has a p-value below the chosen threshold, while the +9 change in B does not. That difference in labels can arise from different precision, even when the estimated changes are close.
To assess whether the changes differ, the analysis needs the between-group contrast and its uncertainty. One small p-value and one larger p-value are not enough to reconstruct it.
The same issue can arise when a paper claims an effect only in one subgroup because that subgroup’s test was significant. Evidence that two effects differ requires a comparison of the effects, not merely separate tests of each against zero.
For a reader, the practical check is to find the exact test named in the caption or methods. Does it compare each condition with baseline, compare conditions with one another, or test a group-by-time relationship?
Read baseline adjustment and design together
A paper may analyse follow-up values with baseline adjustment rather than subtracting raw group changes. That can be a valid way to address the intervention comparison, depending on the model and design. Do not assume a change-score calculation is always required.
Similarly, a raw difference between changes does not automatically establish causality in a non-randomised study. The groups may differ in other time-varying influences or in the process that determined their allocation.
A clear summary distinguishes the observed trajectories from the estimated intervention contrast. It gives the latter’s uncertainty and identifies whether the analysis accounted for relevant baseline and repeated-measure structure.
This preserves a useful observation that a group changed while avoiding a stronger claim that the intervention was responsible. The comparator and design provide the additional evidence needed to interpret that responsibility.
Sources and further detail
- Bland and Altman — Comparisons within randomised groups can be very misleading (opens in a new tab)
BMJ 2011;342:d561, DOI 10.1136/bmj.d561. Publisher-indexed methodological text and simulation description read; direct full-page retrieval returned 403. No published simulation numbers copied. The 40/50 and 42/51 example is original.
Sources checked 19 September 2026. Worked examples are illustrative unless a supplied report is explicitly identified. This article has not undergone independent scientific peer review.