Novum Peptides

Research use only

Before you enter

Please confirm the following before browsing Novum Peptides.

Adults onlyYou must be 18 years or older to enter.

Laboratory research onlyOur products are not for human or veterinary use.

We’ll remember your confirmation on this browser where storage is available.

Novum Peptides · For laboratory research only

Multiple comparisons in multi-endpoint experiments

Recognise how testing many outcomes changes the interpretation of a selected significant result.

A peptide experiment may measure several markers, compare several conditions and examine several time points. Each comparison can answer a useful question. The difficulty arises when one favourable result is interpreted as if it were the only test that had an opportunity to produce an unusual finding.

Count the opportunities relevant to the claim

NIST warns that repeatedly applying pairwise comparison procedures does not generally preserve the intended overall error control. Multiple-comparison methods address a set of comparisons rather than treating each selected result in isolation.NIST/SEMATECH — How can we make multiple comparisons? (opens in a new tab)

For a reader, the relevant set might include all comparisons supporting one main conclusion. It can involve several outcomes, several groups or several analyses of the same outcome. The number of panels in the final figure may be smaller than the number of analyses considered.

Where additional comparisons can arise
ChoiceWhat to inspect
Several markersAll measured outcomes, including unreported ones
Several conditionsWhich groups were compared with which
Several timesWhether one time point was selected afterwards
Several analysesAlternative transformations, subgroups or models

This is not a reason to avoid rich measurements. It is a reason to make the analysis plan and the role of each result visible, especially when a headline depends on a selected finding.

See the difference between per-test and overall error

Consider an original idealised example with twenty independent tests, all of whose null hypotheses are true. Suppose each valid test has exactly a 5% false-positive probability. The chance that none is a false positive is 0.95 raised to the twentieth power.

The chance of at least one false positive is therefore 1 − 0.95²⁰, approximately 64.2%. It is not 5%, and it is not a claim that any particular reported finding has a 64.2% probability of being false.

The independence assumption matters. Biological endpoints often correlate, so this calculation is an illustration rather than a universal correction for every twenty-endpoint study. Different dependencies change the joint probability.

Read what the reported adjustment controls

NIST’s Bonferroni guidance uses the sum of individual error probabilities to bound the overall error for a specified finite set. For testing, allocating an overall 0.05 across twenty valid tests gives a per-test threshold of 0.0025; the bound does not require the tests to be independent.NIST/SEMATECH — Bonferroni’s method (opens in a new tab)

Under that illustrative rule, an unadjusted p-value of 0.01 does not meet the family-level criterion, despite being below 0.05. The effect estimate itself has not become zero; the decision criterion for the set has changed.

Other methods have different purposes and assumptions. A paper should name its procedure and explain the group of comparisons to which it was applied. An adjusted label without that information leaves the reader unsure what protection was intended.

Avoid choosing a method solely because it retains a preferred result after inspection. A defensible analysis connects the procedure to the planned claims, design and acceptable error behaviour, while reporting exploratory work transparently.

Preserve the complete pattern of results

Read the main outcome alongside the other relevant measurements. A result that appears alone in an abstract may have emerged from a broader set whose overall pattern is mixed or uncertain.

Multiplicity also affects how simultaneous intervals should be interpreted. A collection of individually calculated 95% intervals does not automatically have 95% simultaneous coverage for all the parameters together.

Separate a planned primary claim from secondary and exploratory observations, and record any independent follow-up. A discovered signal can be worth investigating even when the original experiment does not establish it under the stated family-level criterion.

A useful evidence summary describes the selected effect, the wider set examined and the multiplicity approach. That is more informative than simply removing or retaining a significance star.

Sources and further detail

  1. NIST/SEMATECH — How can we make multiple comparisons? (opens in a new tab)

    Warning about repeated pairwise testing read. The twenty-independent-test calculation is original and explicitly conditional.

  2. NIST/SEMATECH — Bonferroni’s method (opens in a new tab)

    General inequality and simultaneous coverage read. The 0.05/20 testing example applies the same error bound; no claim that this is optimal for every experiment.

Sources checked 19 September 2026. Worked examples are illustrative unless a supplied report is explicitly identified. This article has not undergone independent scientific peer review.