MedDeviceGuideMedDeviceGuide
Back

Multiplicity and Multiple-Endpoint Testing in Medical Device Clinical Trials

Master multiplicity control in device trials: learn FDA 2022 & CDRH rules, co-primary & composite exemptions, Holm/gatekeeping methods, and IVD endpoint strategies.

Ran Chen
Ran Chen
Global MedTech Expert | 10× MedTech Global Access
Published 2026-08-11Last reviewed 2026-08-1133 min read

When designing a pivotal Investigational Device Exemption (IDE) study or EU MDR clinical investigation, device sponsors frequently evaluate multiple clinical outcomes—such as pain reduction, functional recovery, adverse event rates, and radiographic healing—to capture the full clinical benefit of a novel medical technology. However, conducting multiple statistical hypothesis tests across several primary endpoints, key secondary claims, patient subgroups, or repeated interim looks creates a well-documented statistical vulnerability known as multiplicity.

Without formal statistical adjustments, every additional hypothesis test inflates the probability of committing a false-positive (Type I) error. Testing three independent primary endpoints at a standard two-sided significance level of $\alpha = 0.05$ increases the overall trial-wide false-positive rate from 5.0% to over 14.2%. If a sponsor claims pivotal trial success based on unadjusted significance across multiple endpoints, regulatory agencies such as the US FDA and EU Notified Bodies will reject the efficacy claims or require additional clinical data.

This comprehensive guide clarifies when multiplicity adjustment is legally and statistically required, details the explicit exemptions recognized by regulatory authorities, evaluates FDA-accepted statistical methods (Bonferroni, Holm, Hochberg, fixed-sequence, gatekeeping, and graphical procedures), and maps device-native endpoint patterns—including In Vitro Diagnostic (IVD) co-primary sensitivity and specificity, cardiovascular composite MACE endpoints, and single-arm performance goal benchmarks.


Direct Answer: When Is Multiplicity Adjustment Required vs. Exempt?

For medical device sponsors and biostatisticians asking whether a trial requires formal multiplicity adjustment:

Core Rule: Multiplicity adjustment is mandatory whenever success on any one of multiple statistical tests could be used independently to claim overall trial success, expand labeling claims, or support secondary efficacy indications.

Key Exemptions: Formal alpha-splitting is not required when:

  1. All co-primary endpoints must succeed simultaneously to declare trial success (e.g., IVD diagnostic sensitivity and specificity, or orthopedics safety and efficacy).
  2. A single composite endpoint is evaluated as the primary outcome (e.g., Major Adverse Cardiac Events, or MACE).
  3. Hierarchical ordered testing is pre-specified for a single endpoint (e.g., testing for non-inferiority first, and only if non-inferiority is established, proceeding to test superiority on the exact same endpoint).

While the FDA's guidance document Multiple Endpoints in Clinical Trials (published by CDER and CBER in October 2022) provides the foundational taxonomy of statistical methods, it is explicitly written for drugs and biologics. Medical device trials operate under CDRH's 2013 guidance Design Considerations for Pivotal Clinical Investigations for Medical Devices (78 FR 66943), which enforces the same overarching mandate for Type I error control but requires device-specific adaptations for diagnostic accuracy, operator learning curves, and performance goal benchmarks.


What Is Multiplicity, and Why Does It Inflate False Positives in Device Trials?

Multiplicity occurs whenever a clinical investigation performs more than one statistical hypothesis test. In classical frequentist hypothesis testing, the probability of falsely rejecting a true null hypothesis ($\mathrm{H}_0$) for a single test is constrained by the significance level $\alpha$ (typically $0.05$ two-sided or $0.025$ one-sided).

However, when $k$ independent hypothesis tests are conducted in a single trial, each at significance level $\alpha$, the cumulative probability of making at least one false-positive error—known as the Family-Wise Error Rate (FWER) or overall Type I error rate—rises sharply according to the probability formula:

$$\mathrm{FWER} = 1 - (1 - \alpha)^k$$

Mathematical Inflation of Family-Wise Error Rate

The table below illustrates how the overall false-positive rate escalates as the number of independent statistical hypothesis tests ($k$) increases:

Number of Independent Tests ($k$) One-Sided Nominal $\alpha = 0.025$ Two-Sided Nominal $\alpha = 0.050$ FWER Inflation Factor (vs. Target $\alpha=0.05$)
1 Test (Standard Baseline) 2.50% 5.00% $1.0\times$ (No Inflation)
2 Independent Tests 4.94% 9.75% $1.95\times$
3 Independent Tests 7.31% 14.26% $2.85\times$
5 Independent Tests 11.89% 22.62% $4.52\times$
10 Independent Tests 22.37% 40.13% $8.03\times$

Note: FWER calculations assume independent test statistics. Positive correlation between endpoints slightly mitigates, but does not eliminate, Type I error inflation.

Four Primary Operational Sources of Multiplicity

Multiplicity in device investigations stems from four primary operational sources:

  1. Multiple Primary Endpoints: Evaluating several distinct clinical measures (e.g., pain score, functional disability index, and motion range) where a positive result on any single endpoint is interpreted as trial success.
  2. Multiple Secondary Claims: Testing secondary endpoints intended for promotional or labeling claims after evaluating the primary endpoint.
  3. Multiple Patient Subgroups: Analyzing efficacy across pre-specified subgroups (e.g., mild vs. severe disease, male vs. female, or specific anatomical dimensions).
  4. Multiple Treatment Arms or Device Configurations: Comparing two or more device iterations, dose levels, or settings against a control group.

(Note: Multiplicity arising from repeated interim data looks is governed by alpha-spending functions and interim stopping rules pre-specified for data monitoring committees, as detailed in our guide on data monitoring committee and interim stopping boundaries).


Detailed Case Study: The Cost of Unadjusted Multiplicity in a Device PMA Efficacy Claim

To understand the severe regulatory consequences of unadjusted multiplicity, consider an illustrative composite scenario assembled from documented FDA interaction patterns in artificial spinal disc replacement PMA submissions. The statistical mechanics below are exact; the commercial-impact figures are representative industry ranges for a confirmatory-study extension rather than a single reported case.

Trial Design and Protocol Deficit

The trial sponsor designed a 300-patient randomized controlled trial comparing the new artificial disc against traditional spinal fusion. The protocol pre-specified three primary effectiveness endpoints:

  • Endpoint A: Visual Analog Scale (VAS) pain reduction $\ge 20\text{ mm}$ at 12 months.
  • Endpoint B: Oswestry Disability Index (ODI) score improvement $\ge 15$ points at 12 months.
  • Endpoint C: Radiographic motion preservation confirmed by dynamic flexion/extension X-rays.

The trial protocol stated that the device would be deemed successful if it demonstrated statistical superiority over fusion on any of the three primary endpoints at a nominal two-sided $\alpha = 0.05$. Crucially, the protocol and initial draft of the statistical analysis plan contained no multiplicity adjustment strategy.

Trial Readout and Statistical Evaluation

At trial completion, the statistical analysis yielded the following results across the intent-to-treat (ITT) population:

  • Endpoint A (Pain VAS): $p = 0.038$ (Statistically significant at nominal $0.05$).
  • Endpoint B (ODI Disability): $p = 0.120$ (Not significant).
  • Endpoint C (Radiographic Motion): $p = 0.082$ (Not significant).

The sponsor submitted the PMA claiming pivotal trial success based on Endpoint A. During CDRH panel review, FDA biostatisticians calculated the trial's overall FWER. Because trial success required rejecting only one of three independent null hypotheses, the effective Type I error rate was 14.26%.

Adjusting for multiplicity using the conservative Bonferroni correction required each endpoint to achieve $p \lt 0.0167$ ($0.05 / 3$). Using the Holm step-down procedure, the smallest $p$-value ($0.038$) failed to beat the adjusted threshold of $0.0167$.

Regulatory Outcome and Commercial Impact

FDA would issue a Major Deficiency determination, ruling that trial success had not been demonstrated at an overall $\alpha = 0.05$. Because the statistical analysis plan had not pre-specified a hierarchical testing sequence (which could have placed Endpoint A first and preserved it at full alpha), the sponsor would typically be required to run a confirmatory extension study of roughly 150–200 patients. In practice, such a multiplicity omission commonly produces a two- to three-year delay in market launch and on the order of $10–15 million in additional clinical-trial cost for a spinal-implant pivotal program.


Recommended Reading
Non-Inferiority Clinical Trials for Medical Devices: Margin & Design Guide
Clinical Evidence Regulatory2026-08-08 · 27 min read

When Is Multiplicity Adjustment NOT Required? (Regulatory Exemptions)

Understanding when multiplicity adjustments are unnecessary is just as critical as knowing how to adjust. Over-adjusting alpha unnecessarily penalizes statistical power, driving up sample sizes and trial costs. Regulatory authorities (FDA CDRH, EMA, and MHRA) explicitly recognize three structural exemptions.

Exemption Decision Rule Summary

  • Scenario A — Any Endpoint Can Declare Success: Multiplicity adjustment is mandatory. Examples include testing multiple primary effectiveness outcomes independently, claiming secondary efficacy features without a hierarchy, or analyzing multiple subgroups for independent indications.
  • Scenario B — All Endpoints Must Succeed Simultaneously: No multiplicity adjustment is required. Examples include IVD sensitivity and specificity co-primary goals, or orthopedic safety and efficacy co-primary criteria.
  • Scenario C — Single Composite Primary Endpoint: No multiplicity adjustment is required. Examples include 30-day composite MACE (mortality, stroke, revascularization) in cardiovascular trials.
  • Scenario D — Sequential Non-Inferiority then Superiority: No multiplicity adjustment is required. Non-inferiority is tested first; if passed, superiority is tested on the exact same endpoint without alpha penalty.

1. Co-Primary Endpoints ("All-or-None" Rule)

When a trial protocol defines multiple primary endpoints and requires that every single endpoint must achieve statistical significance for the trial to be declared successful, no multiplicity adjustment is required. This is known as an intersection-union test.

  • Statistical Logic: Because trial success requires rejecting the null hypotheses for Endpoint 1 AND Endpoint 2, the overall Type I error rate is constrained by the maximum Type I error of any individual test ($\text{FWER} \le \alpha$).
  • Device Example: An IVD assay evaluating diagnostic performance for Hepatitis C must demonstrate both Sensitivity $\ge 95%$ ($\alpha = 0.025$ one-sided) AND Specificity $\ge 98%$ ($\alpha = 0.025$ one-sided).
  • The Joint-Power Trade-off: While co-primary endpoints avoid alpha splitting, they severely penalize statistical power. As highlighted in the FDA 2022 guidance, if Endpoint 1 has 80% power and independent Endpoint 2 has 80% power, the overall joint power of the trial drops to:

$$\text{Joint Power} = \text{Power}_1 \times \text{Power}_2 = 0.80 \times 0.80 = 64%$$

To maintain overall trial power at 80% with two independent co-primary endpoints, each individual endpoint must be individually powered to approximately 89.5% ($\sqrt{0.80} \approx 0.8944$), expanding the required sample size.

2. Single Composite Endpoints

Combining multiple clinical events into a single composite binary or time-to-event outcome does not incur a multiplicity penalty because only one statistical test is performed on the composite variable.

  • Device Example: A transcatheter aortic valve replacement (TAVR) trial uses a primary composite endpoint of Major Adverse Cardiac and Cerebrovascular Events (MACCE) at 30 days, defined as all-cause mortality, disabling stroke, or acute kidney injury stage 3.
  • Regulatory Caveat: FDA guidance dictates that while the composite endpoint itself requires no alpha adjustment, testing individual components of the composite for standalone labeling claims does introduce multiplicity and requires a pre-specified hierarchical strategy.

3. Ordered Non-Inferiority to Superiority Testing

Testing a single primary endpoint sequentially for non-inferiority and then for superiority requires no alpha adjustment, provided the testing sequence is strictly ordered.

  • Sequential Procedure:
    1. Test $\mathrm{H}_{01}$: Investigational device is inferior to active control by margin $\delta$ at $\alpha = 0.025$ (one-sided).
    2. If $\mathrm{H}{01}$ is rejected (establishing non-inferiority), proceed to test $\mathrm{H}{02}$: Investigational device is superior to active control ($\text{difference} = 0$) at $\alpha = 0.025$ (one-sided) using the exact same patient data.
  • Why Alpha Is Preserved: Because $\mathrm{H}{02}$ is a subset of the parameter space ruled out by rejecting $\mathrm{H}{01}$, the two hypotheses are nested and closed. If non-inferiority fails, testing stops immediately, preserving overall Type I error at $\alpha = 0.025$.
  • Reference: See our detailed guide on non-inferiority trial design and margin for margin derivation rules ($M_1$ and $M_2$).

Taxonomy of FDA-Accepted Multiplicity Adjustment Methods

When multiplicity adjustment is necessary, biostatisticians select from a hierarchy of statistical procedures ranging from simple single-step corrections to flexible graphical approaches.

Classification of Adjustment Methods

  1. Single-Step Methods (Fixed Equal or Custom Alpha Allocation):
    • Bonferroni Method: Divides overall alpha equally among $k$ hypotheses ($\alpha_i = \alpha / k$).
    • Sidak Method: Adjusts for independence among tests ($\alpha_i = 1 - (1 - \alpha)^{1/k}$).
    • Prospective Alpha Allocation Scheme (PAAS): Custom unequal alpha allocation across endpoints ($\sum \alpha_i \le \alpha$).
  2. Stepwise Data-Driven Methods (Sequential P-Value Testing):
    • Holm Step-Down Procedure (1979): Tests smallest $p$-value first against $\alpha / k$, stepping down dynamically.
    • Hochberg Step-Up Procedure (1988): Tests largest $p$-value first against $\alpha$, stepping up under Simes property.
    • Hommel Procedure (1988): Closed testing procedure based on Simes test, offering slightly higher power than Hochberg.
  3. Hierarchical & Graphical Methods (Power-Preserving Procedures):
    • Fixed-Sequence Testing: Strict prioritized sequence ($\mathrm{E}_1 \rightarrow \mathrm{E}_2 \rightarrow \mathrm{E}_3$) protecting 100% of primary alpha.
    • Serial Gatekeeping: Primary endpoint family must succeed fully before alpha passes to secondary families.
    • Parallel Gatekeeping: Alpha passes to secondary families if at least one primary endpoint succeeds.
    • Graphical Approach (Bretz et al., 2009): Directed graph with node alpha allocations and weighted alpha transfer rules.

Detailed Examination of Key Procedures

Holm Step-Down Procedure (Holm, 1979)

The Holm step-down method is one of the most widely accepted sequential procedures. The observed $p$-values for $k$ hypotheses are ordered from smallest to largest:

$$p_{(1)} \le p_{(2)} \le \dots \le p_{(k)}$$

The algorithm proceeds sequentially as follows:

  • Step 1: Compare $p_{(1)}$ to $\alpha / k$. If $p_{(1)} \lt \alpha / k$, reject $\mathrm{H}_{(1)}$ and proceed to Step 2. If not, stop testing and reject no hypotheses.
  • Step 2: Compare $p_{(2)}$ to $\alpha / (k - 1)$. If $p_{(2)} \lt \alpha / (k - 1)$, reject $\mathrm{H}_{(2)}$ and proceed to Step 3. If not, stop.
  • Step $i$: Compare $p_{(i)}$ to $\alpha / (k - i + 1)$. Continue as long as $p_{(i)} \lt \alpha / (k - i + 1)$.

Because Holm step-down evaluates hypotheses sequentially, it is strictly more powerful than Bonferroni while making zero assumptions about the joint correlation structure of the test statistics.

Hochberg Step-Up Procedure (Hochberg, 1988)

The Hochberg step-up procedure begins at the largest observed $p$-value $p_{(k)}$:

  • Step 1: Compare $p_{(k)}$ to $\alpha$. If $p_{(k)} \lt \alpha$, reject all $k$ hypotheses and stop. If not, proceed to Step 2.
  • Step 2: Compare $p_{(k-1)}$ to $\alpha / 2$. If $p_{(k-1)} \lt \alpha / 2$, reject $\mathrm{H}{(k-1)}$ and all hypotheses with smaller $p$-values ($p{(1)} \dots p_{(k-2)}$). If not, continue stepping up.

Hochberg step-up is more powerful than Holm step-down, but it requires that test statistics satisfy the Simes inequality (which holds under independence or positive orthant dependence). FDA biostatisticians accept Hochberg for most standard clinical trial endpoints.


Comparison Matrix: Multiplicity Control Procedures for Medical Device Pivotal Trials

The following matrix compares statistical procedures across key operational dimensions relevant to device sponsors:

Method FWER Control Power Preservation for Primary Pre-Specification Complexity FDA / Notified Body Acceptance Best Device Use Case
Bonferroni Strong (Any correlation) Very Low Minimal Universal 2 simple, independent secondary endpoints.
Holm (Step-Down) Strong (Any correlation) Moderate Low Universal Multiple secondary endpoints of equal priority.
Hochberg (Step-Up) Strong (Simes property) High Low High (Requires Simes verification) Multiple secondary claims where endpoints are positively correlated.
Fixed-Sequence Strong (Strict order) Maximum (100% $\alpha$) Low Gold Standard Pivotal trial with 1 clear primary endpoint and 2 prioritized secondary claims.
Serial Gatekeeping Strong (Family-wise) High for Primary Family Moderate High Primary endpoint family (e.g., Safety & Efficacy) leading to Secondary claims.
Graphical (Bretz) Strong (Flexible graph) High High (Requires specialized biostat package) High (Requires FDA Pre-Sub alignment) Complex trials with multiple primary endpoints and interconnected secondary claims.

Recommended Reading
Should You Trust AI in Regulatory Work? A Vendor-Neutral Checklist After DJ Fang's Pilot
Digital Health & AI Regulatory2026-08-07 · 21 min read

Five Device-Native Multiplicity Patterns

While drug trials primarily evaluate multiple dose groups or multiple symptom scales, medical device investigations present five distinct clinical trial patterns.

Pattern Clinical Device Sector Primary Statistical Challenge Recommended Strategy
1. IVD Diagnostic Accuracy Molecular Diagnostics, IVDs Sensitivity & Specificity co-primary Co-primary exemption; power each endpoint $\ge 89.5%$
2. Composite MACE Gatekeeping Structural Heart, TAVR, Stents Composite primary vs. component claims Primary composite tested at full alpha; serial gatekeeping for components
3. OPC / Performance Goals Orthopedics, Ophthalmic, Vascular Multiple single-arm historical goals Co-primary if all required; fixed-sequence if prioritized
4. Primary Efficacy to PRO Spine implants, Joint Replacement Clinical success to Quality-of-Life claims Fixed-sequence hierarchy putting clinical endpoints first
5. Multi-Arm Device Iterations Electrophysiology, Energy devices Generation 2.0 vs 2.1 vs Control Dunnett's procedure or Hochberg step-up vs. common control

Pattern 1: IVD Diagnostic Accuracy (Sensitivity and Specificity Co-Primary)

In Vitro Diagnostic (IVD) performance evaluations under FDA 510(k), De Novo, or PMA pathways require evaluating both Sensitivity ($TP / [TP + FN]$) and Specificity ($TN / [TN + FP]$) against a clinical reference standard.

  • Multiplicity Rule: Co-primary endpoint exemption applies. Both Sensitivity $\ge P_1$ and Specificity $\ge P_2$ must achieve statistical significance at one-sided $\alpha = 0.025$.
  • Statistical Hazard: Do not split alpha (e.g., do not test at $0.0125$). However, sample size calculations must account for joint power. If disease prevalence in the study population is low (e.g., 5%), enrolling enough true positive subjects to achieve 90% power on sensitivity will dictate the total sample size.

Pattern 2: Cardiovascular Composite MACE and Component Gatekeeping

Interventional cardiology devices (stents, heart valves, vascular grafts) routinely use a composite primary endpoint of Major Adverse Cardiac Events (MACE) at 12 months.

  • Multiplicity Rule: The primary MACE composite is tested at full $\alpha = 0.05$.
  • Gatekeeping Structure: If MACE superiority is achieved, a serial gatekeeping sequence unlocks testing for individual component claims:
    • Step 1: Test Primary Composite MACE ($p \lt 0.05$).
    • Step 2: Test Target Lesion Revascularization (TLR) ($p \lt 0.05$).
    • Step 3: Test Stent Thrombosis ($p \lt 0.05$).
  • Reference: Contrast this with superiority and equivalence trial design when establishing superiority margins.

Pattern 3: Single-Arm Performance Goals vs. Multiple Historic Benchmarks

Orthopedic, ophthalmic, and vascular devices frequently utilize single-arm pivotal trials compared against pre-specified Performance Goals (PG) or Objective Performance Criteria (OPC) derived from historical registry data.

  • Multiplicity Rule: If a device is evaluated against two independent OPC benchmarks (e.g., 12-month primary patency $\ge 75%$ AND 12-month freedom from major adverse events $\ge 88%$), the protocol must define whether these are co-primary (all-or-none, no alpha split) or multiple primary endpoints (either suffices, requiring Bonferroni or fixed-sequence adjustment).

Pattern 4: Primary Efficacy to Patient-Reported Outcome (PRO) Claims

Sponsors seeking to include Patient-Reported Outcome (PRO) claims (e.g., health-related quality of life, pain interference) in device labeling must integrate PRO endpoints into a formal gatekeeping hierarchy.

  • Multiplicity Rule: PRO endpoints must occupy secondary or tertiary positions in a fixed-sequence chain. Testing a PRO endpoint at $\alpha = 0.05$ is only permitted after all primary clinical endpoints have achieved statistical significance.

Pattern 5: Multi-Arm Device Iterations or Multi-Size Families

When a trial compares two novel device iterations (e.g., Generation 2.0 and Generation 2.1) against a common control, or compares two energy delivery settings against control:

  • Multiplicity Rule: This creates a multi-arm comparison structure. Dunnett's procedure or a Hochberg step-up adjustment is required to control FWER across the two active arms versus control.

Worked Example: Fixed-Sequence Gatekeeping in an Orthopedic Total Joint Trial

To illustrate the exact operational mechanics of fixed-sequence gatekeeping, consider a 400-patient randomized controlled trial evaluating a novel ceramic-on-ceramic total hip replacement system versus a standard metal-on-polyethylene control.

Hypothesis Hierarchy Pre-Specified in the SAP

The sponsor pre-specified a four-step fixed-sequence testing hierarchy in the SAP prior to unblinding:

  1. Hypothesis 1 ($\mathrm{H}_1$ - Primary Non-Inferiority): 24-month Harris Hip Score (HHS) non-inferiority margin $\delta = 5.0$ points at one-sided $\alpha = 0.025$.
  2. Hypothesis 2 ($\mathrm{H}_2$ - Primary Superiority): 24-month Harris Hip Score superiority over control ($\Delta > 0$) at two-sided $\alpha = 0.05$.
  3. Hypothesis 3 ($\mathrm{H}_3$ - Key Secondary Wear Rate): 24-month linear femoral head wear rate superiority ($\text{wear}{\text{test}} \lt \text{wear}{\text{control}}$) at two-sided $\alpha = 0.05$.
  4. Hypothesis 4 ($\mathrm{H}_4$ - Secondary Pain VAS): 24-month visual analog scale pain reduction superiority at two-sided $\alpha = 0.05$.

Trial Analysis Readout & Stepwise Evaluation

At the 24-month database lock, biostatisticians executed the sequential analysis:

  • Step 1 ($\mathrm{H}_1$ Test): Non-inferiority margin $p \lt 0.001$. $\mathrm{H}_1$ is rejected; non-inferiority is established. Proceed to Step 2.
  • Step 2 ($\mathrm{H}_2$ Test): HHS superiority test yields $p = 0.018$. $\mathrm{H}_2$ is rejected at two-sided $\alpha = 0.05$; superiority on HHS is established. Proceed to Step 3.
  • Step 3 ($\mathrm{H}_3$ Test): Wear rate reduction test yields $p = 0.004$. $\mathrm{H}_3$ is rejected at two-sided $\alpha = 0.05$; superiority on wear rate is established. Proceed to Step 4.
  • Step 4 ($\mathrm{H}_4$ Test): Pain VAS reduction test yields $p = 0.082$. $\mathrm{H}_4$ fails to reach significance ($p \ge 0.05$). Testing halts immediately.

Regulatory Claim Outcome

Because the fixed-sequence gatekeeping rules were strictly followed:

  • The sponsor successfully claimed Non-Inferiority, Superiority on HHS, and Superiority on Linear Wear Rate in FDA PMA Section 14 labeling and the EU MDR SSCP.
  • The Pain VAS result ($p = 0.082$) was designated as exploratory without promotional claims.
  • Because testing halted at Step 4, overall trial FWER remained strictly controlled at $\alpha = 0.05$ with zero sample size inflation penalty.

Real-World Multiplicity Strategies Across Additional Medical Device Specialties

To further demonstrate how multiplicity strategies adapt to diverse clinical technology sectors, we examine three specialized device categories.

1. Continuous Glucose Monitoring (CGM) & Software as a Medical Device (SaMD)

Continuous Glucose Monitors evaluated under FDA De Novo or 510(k) (iCGM classification) require evaluating accuracy across glycemic ranges:

  • Co-Primary Accuracy Endpoints: Percentage of CGM readings within $\pm 15% / \pm 15\text{ mg/dL}$ of reference YSI values (target $\ge 85%$) AND percentage within $\pm 20% / \pm 20\text{ mm/dL}$ (target $\ge 98%$).
  • Exemption Rule: Co-primary intersection-union rule applies. Both accuracy criteria must be satisfied simultaneously at $\alpha = 0.025$ one-sided.
  • Secondary Serial Gatekeeping: Once accuracy co-primary criteria are met, testing proceeds to secondary hypoglycemia detection alert sensitivity and time-in-range (TIR 70-180 mg/dL) superiority claims.

2. Neurovascular Thrombectomy Catheters

Mechanical thrombectomy catheters for acute ischemic stroke evaluation in single-arm pivotal investigations:

  • Primary Reperfusion Goal: Modified Treatment in Cerebral Infarction (mTICI) $2\text{b}/3$ success rate at immediate post-procedure angiogram compared against a pre-specified OPC benchmark of $70%$ ($\alpha = 0.025$ one-sided).
  • Key Secondary Hierarchy: 90-day functional independence (modified Rankin Scale $\text{mRS} \le 2$) tested at $\alpha = 0.05$ only if primary mTICI reperfusion meets its OPC goal.
  • Safety Endpoint: 24-hour symptomatic intracranial hemorrhage (sICH) rate evaluated against a historical safety threshold as a co-primary safety criterion.

3. Ophthalmic Premium Intraocular Lenses (IOLs)

Extended Depth of Focus (EDOF) or multifocal intraocular lenses evaluated for PMAs:

  • Primary Visual Acuity Co-Primary: Monocular distance uncorrected visual acuity (UCVA) AND intermediate uncorrected visual acuity (DCIVA) both meeting non-inferiority vs. monofocal control lenses at 6 months.
  • Secondary Binocular Gatekeeping: Binocular near uncorrected visual acuity (DNCVA) superiority tested in a fixed-sequence step following primary monocular success.

Recommended Reading
FDA Medical Device Development Tools (MDDT): Program & Qualified Tools Guide
Clinical Evidence Regulatory2026-07-29 · 19 min read

Interplay: Multiplicity, Sample Size, Subgroups, and the SAP

Sample Size Inflation Under Alpha Splitting

When an alpha-splitting method (such as Bonferroni) is selected, reducing the nominal significance level from $\alpha = 0.05$ to $\alpha / k$ directly increases the critical z-score threshold ($z_{1 - \alpha/2}$), inflating the required sample size ($N$).

For a two-sample parallel group comparison of means, sample size per arm is derived as:

$$N_{\text{arm}} = \frac{2 \left( z_{1 - \alpha/2} + z_{1 - \beta} \right)^2 \sigma^2}{\Delta^2}$$

The table below demonstrates the sample size penalty associated with alpha splitting for a trial powered at 80% ($\beta = 0.20$, $z_{0.80} = 0.842$) with a standardized effect size of $\Delta / \sigma = 0.35$:

Adjustment Strategy Effective Alpha ($\alpha$) Critical Value ($z_{1 - \alpha/2}$) Required $N$ Per Arm Sample Size Increase vs. Baseline Total Patients (2 Arms)
Unadjusted Baseline / Fixed-Sequence 0.0500 1.960 129 0% (Baseline) 258
2-Endpoint Bonferroni Split 0.0250 2.241 156 +20.9% 312
3-Endpoint Bonferroni Split 0.0167 2.394 171 +32.6% 342
5-Endpoint Bonferroni Split 0.0100 2.576 191 +48.1% 382

Lesson for Sponsors: Using fixed-sequence gatekeeping instead of a 3-endpoint Bonferroni split saves 84 total patients (342 vs. 258) while fully preserving primary endpoint power. For detailed guidance on sample size calculations across device trial types, review our framework for sample size calculation for device investigations.

Pre-Specified Subgroups vs. Post-Hoc Explorations

Analyzing treatment efficacy across demographic or anatomical subgroups introduces severe multiplicity hazards:

  1. Pre-Specified Subgroups: If a sponsor intends to seek a targeted indication claim for a specific subgroup (e.g., patients with vessel diameter under 2.5 mm), the subgroup hypothesis must be integrated into the formal multiplicity gatekeeping plan.
  2. Exploratory Subgroups: Subgroup analyses not included in the pre-specified gatekeeping hierarchy are strictly exploratory. They cannot support formal efficacy claims in FDA PMA summaries or EU MDR Summary of Safety and Clinical Performance (SSCP) documents.

Interaction with Interim Analyses and Data Monitoring Committees

Conducting interim analyses for early stopping (for efficacy or futility) introduces an additional dimension of multiplicity over time (interim looks). This temporal multiplicity must be controlled using alpha-spending functions (such as O'Brien-Fleming or Pocock boundaries). When a trial incorporates both multiple endpoints and multiple interim looks, the SAP must specify how alpha spending at interim looks interacts with the endpoint gatekeeping hierarchy. Detailed interim stopping governance rules are available in our guide to data monitoring committee and interim stopping boundaries.

Strict Pre-Specification in the Statistical Analysis Plan (SAP)

Regulatory agencies strictly enforce the requirement that the multiplicity control strategy must be fully detailed in the SAP prior to unblinding or initiating definitive primary analysis.

Key SAP elements required by FDA biostatistical reviewers include:

  • Exact mathematical definition of the FWER to be controlled (typically two-sided $0.05$ or one-sided $0.025$).
  • Complete list of all primary and secondary hypotheses, categorized into explicit endpoint families.
  • Complete specification of the adjustment algorithm (including graph weights, gatekeeping rules, or ordering sequences).
  • Pre-planned handling of missing data and its interaction with endpoint ordering.
  • For adaptive trial designs, explicit integration of interim alpha spending with endpoint multiplicity, matching principles established for adaptive designs for medical device clinical studies and Bayesian methods for device clinical trials.

ISO 14155:2020 & EU MDR Annex XV Regulatory Alignment

In the European Union, clinical investigations for medical devices are governed by Regulation (EU) 2017/745 (EU MDR) Annex XV and standard ISO 14155:2020 (Clinical investigation of medical devices for human subjects — Good clinical practice).

ISO 14155:2020 Section 6.5 Requirements

ISO 14155:2020 (Clause 6.5, the clinical investigation plan's statistical design and analysis section) requires that the Clinical Investigation Plan (CIP) and Statistical Analysis Plan explicitly describe:

  • The statistical rationale for the choice of primary performance and safety endpoints.
  • The explicit handling of multiple comparisons and procedures for controlling Type I error.
  • The definition of analysis populations (Full Analysis Set, Per-Protocol Set) and how missing data handling impacts endpoint testing order.

EU Notified Body Evaluation Criteria

EU Notified Bodies (such as TÜV SÜD, BSI, and DEKRA) review clinical investigation data submitted in the Technical Documentation for CE marking. Notified Body clinical experts apply MDCG guidance (including MDCG 2020-6 on adequate clinical evidence for legacy devices) to verify that claims made in the Summary of Safety and Clinical Performance (SSCP) are supported by statistically valid analyses.

If an EU clinical trial evaluates multiple secondary endpoints without pre-specified multiplicity controls, Notified Bodies will reject secondary clinical claims from the SSCP and Instructions for Use (IFU), restricting the device's intended purpose statement.


Advanced Multiplicity Models: Adaptive Multi-Arm & Platform Designs

Medical device clinical development increasingly employs adaptive multi-arm multi-stage (MAMS) designs, platform trials, and seamless Phase II/III device studies.

Multi-Arm Adaptive Adaptations

In a multi-arm trial evaluating three novel device iterations ($\text{Dev}_A, \text{Dev}_B, \text{Dev}_C$) against a common control ($\text{Control}$):

  • Arm Selection Adaptation: At an interim look, the unblinded Data Monitoring Committee (DMC) selects the best-performing arm ($\text{Dev}_A$) and drops $\text{Dev}_B$ and $\text{Dev}_C$.
  • Multiplicity Risk: Selecting the best arm based on interim data inflates Type I error because the final p-value is evaluated on a selected hypothesis.
  • Control Procedure: The trial must use a Combination Test (Bauer & Köhne, 1994) or a Closed Testing Procedure combined with Dunnett's test to ensure that final pairwise comparisons control FWER at $\alpha = 0.025$ one-sided.

CDRH 2016 Adaptive Study Designs Guidance

FDA CDRH's guidance Adaptive Designs for Medical Device Clinical Studies (finalized July 2016) explicitly requires sponsors of adaptive device trials to perform prospective simulation studies demonstrating that overall Type I error remains $\le 0.05$ across all potential adaptation paths.


Recommended Reading
Bayesian Statistics for Medical Device Clinical Trials: FDA Guidance & Approvals
Clinical Evidence Regulatory2026-08-07 · 17 min read

Global Regulatory Comparison: FDA, EU MDR, PMDA, and NMPA

Managing multiplicity across international device registration filings requires understanding subtle regulatory differences across key global authorities:

Regulatory Jurisdiction Governing Law / Guidance Document Multiplicity Enforcement Strictness Key Regional Requirement
United States (FDA CDRH) CDRH 2013 Pivotal Investigations Guidance; CDER/CBER 2022 Multiple Endpoints Guidance Very High Formal SAP lock prior to unblinding; Pre-Sub alignment strongly recommended.
European Union (MDR / Notified Bodies) EU MDR 2017/745 Annex XV; ISO 14155:2020 Clause 6.5 High Unadjusted secondary endpoints excluded from SSCP and IFU claims.
Japan (PMDA) PMDA Technical Guidance on Medical Device Clinical Trials High Aligns with ICH E9; requires strict justification for multi-arm trial adaptations.
China (NMPA) Medical Device Clinical Trial Design Guideline (CFDA Announcement No. 6, 2018) High Requires multiplicity and Type I error control when multiple primary endpoints could each support conclusions.

Mathematical Foundations: Closed Testing Principle & Graphical Procedures

To implement advanced multiplicity adjustments such as serial gatekeeping or graphical methods, biostatisticians rely on the Closed Testing Principle (Marcus et al., 1976).

The Closed Testing Principle Explained

Given $k$ elementary hypotheses $\mathrm{H}_1, \mathrm{H}_2, \dots, \mathrm{H}_k$, the closed testing principle constructs an intersection hypothesis $\mathrm{H}I = \bigcap{j \in I} \mathrm{H}_j$ for every non-empty subset $I \subseteq {1, \dots, k}$, yielding $2^k - 1$ intersection hypotheses.

An elementary hypothesis $\mathrm{H}_i$ is rejected at local significance level $\alpha$ if and only if every intersection hypothesis $\mathrm{H}_I$ containing $\mathrm{H}_i$ is rejected at local level $\alpha$ by an alpha-consistent local test (such as Simes or Bonferroni).

Graphical Alpha Transfer (Bretz et al., 2009)

Bretz's graphical approach translates closed testing into an intuitive framework:

  1. Assign initial significance levels $\alpha_i$ to each node $\mathrm{H}_i$ such that $\sum \alpha_i \le \alpha$.
  2. Assign directed edge weights $w_{i,j} \in [0, 1]$ representing the fraction of $\mathrm{H}_i$'s alpha passed to $\mathrm{H}_j$ if $\mathrm{H}_i$ is rejected.
  3. Upon rejection of $\mathrm{H}_i$, node $\mathrm{H}_i$ is removed, its remaining alpha is distributed to adjacent nodes according to edge weights, and edge weights between remaining nodes are updated via explicit algorithm rules:

$$w_{j,k}^{(i)} = \frac{w_{j,k} + w_{j,i} \cdot w_{i,k}}{1 - w_{j,i} \cdot w_{i,j}}$$

This graphical framework allows biostatisticians to model complex clinical trial structures—such as multiple primary endpoints linked to multiple secondary endpoints—while guaranteeing strict FWER control.


Practical Step-by-Step Multiplicity Implementation Framework

To assist device biostatisticians and clinical operations managers in designing robust, regulatory-compliant pivotal trial protocols, we present a five-step practical implementation framework.

Step 1: Endpoint Classification & Family Definition

During protocol development, map all trial objectives into a strict hierarchy of endpoint families:

  • Primary Endpoint Family: The core efficacy and safety outcomes required for primary marketing authorization. Limit to 1 composite or co-primary pair whenever feasible.
  • Key Secondary Endpoint Family: High-priority clinical claims intended for inclusion in section 14 (Clinical Studies) of FDA labeling or EU MDR SSCP.
  • Exploratory Endpoint Family: Mechanistic, health economic, or exploratory biomarker outcomes intended solely for internal R&D or hypothesis generation.

Step 2: Determine Exemption Eligibility

Evaluate each primary endpoint family against the three regulatory exemption rules:

  • Are all endpoints required to be significant simultaneously? (If yes $\rightarrow$ Co-primary, no alpha split).
  • Is the primary outcome evaluated as a single composite event? (If yes $\rightarrow$ Single test, no alpha split).
  • Are hypotheses ordered sequentially (e.g., Non-inferiority $\rightarrow$ Superiority)? (If yes $\rightarrow$ Nested tests, no alpha split).

Step 3: Select the Optimal Adjustment Procedure

If multiplicity adjustment is required across $k$ endpoints in a family:

  • Default to fixed-sequence hierarchical testing if endpoints can be logically prioritized ($\mathrm{E}_1 \rightarrow \mathrm{E}_2 \rightarrow \mathrm{E}_3$).
  • Select Holm step-down or Hochberg step-up if endpoints carry equal clinical weight and no natural ordering exists.
  • Utilize a graphical approach (Bretz) if complex weight-sharing between primary and secondary families is required.

Step 4: Conduct Power Simulations & Sample Size Adjustments

Calculate trial sample size incorporating the selected multiplicity procedure:

  • For co-primary endpoints, adjust per-endpoint power to maintain overall joint power $\ge 80%$.
  • For alpha-splitting methods (Bonferroni, Holm), calculate sample size using adjusted alpha $\alpha / k$.
  • Quantify the sample size savings of fixed-sequence testing vs. alpha splitting to present to executive leadership.

Step 5: Document Strategy in Protocol & SAP and Request FDA Alignment

Detail the complete multiplicity control algorithm in the clinical protocol and standalone Statistical Analysis Plan:

  • Write explicit pseudo-code or step-by-step decision rules in the SAP.
  • Submit the draft SAP to FDA CDRH via a Pre-Submission (Q-Sub) meeting to obtain formal agreement on the multiplicity strategy prior to enrolling the first patient.

Biostatistician Checklist for Protocol & SAP Submissions

Before locking the Statistical Analysis Plan or unblinding pivotal trial data, verify that the submission package satisfies every item on this regulatory checklist:

  • Pre-Specification: Is the complete multiplicity strategy specified in writing prior to unblinding or initiating the primary analysis?
  • Target FWER Defined: Is the target Family-Wise Error Rate explicitly stated (e.g., two-sided $\alpha = 0.05$ or one-sided $\alpha = 0.025$)?
  • Hypothesis Categorization: Are all primary, secondary, and exploratory hypotheses explicitly listed and grouped into endpoint families?
  • Exemption Rationale: If co-primary or composite exemptions are claimed, is the mathematical justification documented?
  • Testing Sequence: If fixed-sequence or gatekeeping procedures are used, is the exact order of testing unambiguously defined?
  • Statistical Software Code: Is macro code or SAS/R syntax (e.g., SAS PROC MULTTEST, R gMCP package) included in the SAP appendix?
  • Missing Data Interaction: Does the SAP define how missing data imputations (e.g., jump-to-reference, tipping point sensitivity analysis) interact with endpoint ordering?
  • Interim Analysis Integration: Are endpoint gatekeeping rules integrated with interim alpha-spending functions?
  • Subgroup Scope: Are pre-specified subgroups for labeling claims separated from exploratory subgroup analyses?
  • FDA Q-Sub Consensus: Has the multiplicity control strategy been reviewed and agreed upon during an FDA Q-Submission meeting?

Strategic Recommendations for Medical Device Biostatisticians

To optimize pivotal trial design and ensure seamless regulatory approval, sponsors should follow a four-step strategic workflow:

  1. Minimize Primary Endpoints: Whenever possible, restrict the primary endpoint family to a single clinical outcome or a validated composite score.
  2. Prioritize Fixed-Sequence Testing: For secondary endpoints intended for labeling claims, establish a hierarchical fixed-sequence testing order. Avoid Bonferroni alpha-splitting unless endpoints carry equal clinical weight and cannot be ordered logically.
  3. Account for Joint Power in Co-Primary Designs: When co-primary endpoints are mandatory (such as IVD diagnostic sensitivity and specificity), power each endpoint individually to $\ge 89%$ to prevent joint-power collapse.
  4. Lock the Multiplicity Plan in the Q-Sub Process: Submit the detailed multiplicity section of the draft SAP to FDA CDRH via a formal Q-Submission (Pre-Sub) meeting prior to protocol freeze and IDE approval.

Frequently Asked Questions (FAQ)

Does a single-arm device trial with one primary performance goal require a multiplicity adjustment?

No. A single-arm trial evaluating one primary endpoint against a single pre-specified Performance Goal (PG) or Objective Performance Criteria (OPC) involves only one statistical hypothesis test ($H_0: \mu \le \text{PG}$ vs. $H_1: \mu > \text{PG}$) and carries no multiplicity penalty. Multiplicity only arises if the single arm is evaluated against multiple independent PGs where meeting any one PG is claimed as trial success.

For an IVD study, do I need to adjust alpha when both sensitivity and specificity must meet target goals?

No. Because diagnostic validation requires demonstrating both Sensitivity $\ge \text{Goal}_1$ AND Specificity $\ge \text{Goal}_2$ simultaneously, this functions as a co-primary intersection-union test. Overall Type I error is bounded by $\alpha = 0.025$ one-sided. However, you must power each endpoint higher (typically $\ge 90%$) to maintain acceptable joint statistical power.

What is the most multiplicity-efficient strategy for a device pivotal trial with one primary and three key secondary endpoints?

Hierarchical fixed-sequence testing is the most efficient strategy. By testing the primary endpoint first at full $\alpha = 0.05$, you preserve 100% of statistical power without sample size inflation. Secondary endpoints are tested sequentially ($\mathrm{E}_1 \rightarrow \mathrm{E}_2 \rightarrow \mathrm{E}_3$) at $\alpha = 0.05$, with testing stopping as soon as any endpoint in the chain fails to reach statistical significance.

If my superiority test is not significant, can I still claim equivalence or non-inferiority on the same endpoint without adjustment?

Only if the protocol and SAP pre-specified a non-inferiority hypothesis before testing superiority. The standard, accepted ordered-hypothesis sequence is to test Non-Inferiority first; if non-inferiority is established ($p \lt 0.025$ one-sided), you may test Superiority at $\alpha = 0.025$ without penalty. Reversing the order or adding a post-hoc non-inferiority margin after a failed superiority test is rejected by regulatory reviewers.

Does the FDA Multiple Endpoints guidance apply to medical devices given it was issued by CDER and CBER?

The 2022 FDA Multiple Endpoints in Clinical Trials guidance was formally published by CDER and CBER for drugs and biologics. However, CDRH biostatisticians reference its underlying mathematical framework and taxonomy of statistical methods (Holm, Hochberg, gatekeeping, graphical models) during device PMA and De Novo reviews. Device sponsors must apply these statistical principles while adhering to CDRH-specific rules defined in the 2013 CDRH Pivotal Investigations guidance.


Regulatory Guidance & Primary Reference Sources

  • U.S. FDA CDER/CBER Guidance (October 2022): Multiple Endpoints in Clinical Trials Guidance for Industry (Docket FDA-2016-D-4460; FR Doc 2022-22882).
  • U.S. FDA CDRH Guidance (November 2013): Design Considerations for Pivotal Clinical Investigations for Medical Devices (78 FR 66943).
  • European Medicines Agency (EMA): Guideline on Multiplicity Issues in Clinical Trials (Draft / CHMP/EWP/908/99).
  • International Council for Harmonisation (ICH): ICH E9 Statistical Principles for Clinical Trials (Step 4, September 1998).
  • ISO Standard (2020): ISO 14155:2020 Clinical investigation of medical devices for human subjects — Good clinical practice.
  • Bretz et al. (2009): A graphical approach to sequentially rejective multiple test procedures. Statistics in Medicine, 28(4), 586–604.
  • Dmitrienko et al. (2008): Gatekeeping procedures in clinical trials. Journal of Biopharmaceutical Statistics, 18(6), 1147–1165.