MedDeviceGuideMedDeviceGuide
Back

Superiority and Equivalence Clinical Trial Designs for Medical Devices

Master superiority and equivalence trial designs for medical devices: navigate TOST mechanics, OPC benchmark derivation, CLSI EP09 agreement, and FDA standards.

Ran Chen
Ran Chen
Global MedTech Expert | 10× MedTech Global Access
Published 2026-08-10Last reviewed 2026-08-1027 min read

When medical device clinical leaders, biostatisticians, and regulatory affairs executives design a pivotal investigation under an Investigational Device Exemption (IDE) or Pre-Market Approval (PMA) submission, selecting the correct primary hypothesis structure is one of the most critical statistical and strategic decisions in the product development lifecycle. While non-inferiority trials are frequently evaluated for incremental or iterative device modifications, sponsors aiming to establish clear therapeutic superiority over existing clinical comparators—or proving sameness for a novel device, diagnostic assay, or sensor—must construct formal superiority or equivalence trial architectures.

In pharmaceutical development, superiority and equivalence trial designs typically assume prospective randomization against a concurrent placebo or active drug control. Medical device clinical investigations, however, operate under distinct clinical, ethical, and practical constraints. High-risk Class III devices, surgical implants, and diagnostic instruments often cannot be ethically evaluated against sham interventions or outdated active controls. Consequently, device-native superiority frequently manifests as single-arm comparison against a fixed historical benchmark—known as an Objective Performance Criterion (OPC) or Performance Goal (PG)—derived from multi-center registry data and clinical literature.

Furthermore, medical device equivalence spans two entirely distinct domains: therapeutic clinical equivalence evaluated via formal statistical hypothesis testing (such as Two One-Sided Tests, or TOST), and analytical measurement equivalence evaluated in In Vitro Diagnostic (IVD) devices and continuous sensors through consensus standards like CLSI EP09-A3 and Bland-Altman agreement analysis.

This guide provides an authoritative decision and implementation framework for superiority and equivalence clinical trial designs in medical devices. It details the statistical hypothesis structures, analyzes regulatory guidance from FDA CDRH and ICH, evaluates recent empirical evidence on nonconcurrent control usage in FDA device approvals, provides step-by-step methods for IVD measurement agreement, and clarifies essential boundary conditions with related regulatory concepts.


Executive Summary & Core Decision Framework

Scenario & Direct Answer

Sponsor Scenario: We are designing a pivotal clinical investigation under an IDE/PMA framework. Our new device is either expected to outperform the active control (or an established historical benchmark) on the primary clinical endpoint, or is intended to perform identically to a comparator where our regulatory claim is sameness (such as an IVD assay matching a reference measurement procedure). How do we structure and analyze superiority or equivalence hypotheses, when is a single-arm OPC or performance-goal benchmark acceptable to FDA CDRH instead of a concurrent control, and how do we avoid critical statistical pitfalls like claiming equivalence from a failed superiority trial?

Direct Answer: A superiority trial tests $H_0: \mu_T - \mu_C \le 0$ against $H_1: \mu_T - \mu_C > 0$ using a two-sided $\alpha = 0.05$ (or one-sided $\alpha = 0.025$). Superiority is established when the $95%$ confidence interval for the treatment effect lies entirely above zero (in the favorable direction). An equivalence trial tests whether the device is neither better nor worse than a comparator within a pre-specified symmetric margin $\delta$, using the Two One-Sided Tests (TOST) procedure. Equivalence is declared at $\alpha = 0.05$ when the entire $(1 - 2\alpha) = 90%$ confidence interval for the treatment difference falls strictly within $[-\delta, +\delta]$. Failing to reject the null hypothesis in a superiority test ($p > 0.05$) is never evidence of equivalence, as the burden of proof and boundary conditions differ fundamentally (Walker & Nowacki, 2010).

For medical devices, the dominant real-world pattern for superiority is single-arm comparison against an Objective Performance Criterion (OPC) or Performance Goal (PG) defined under FDA CDRH's Design Considerations for Pivotal Clinical Investigations for Medical Devices (2013 guidance, Section 7.6.1). Empirical analysis by Mooghali et al. (2025, JAMA Network Open) revealed that $59.1%$ of original high-risk therapeutic device PMA approvals between 2019 and 2023 relied on nonconcurrent controls ($39$ performance goals, $6$ historical controls, $6$ OPCs, and $1$ mixed design), with $80.8%$ conducted as single-group studies. However, only $3.8%$ of these analyses provided formal statistical justification for the comparator. For IVD and sensor applications, equivalence refers to analytical measurement agreement against a reference method governed by CLSI EP09-A3 and Bland-Altman $95%$ limits of agreement ($\text{mean difference} \pm 1.96 \text{ SD}$) evaluated against a Total Allowable Error ($\text{TAE}$), rather than clinical TOST. Finally, statistical equivalence trial designs must be crisply distinguished from EU MDR Article 61 clinical-equivalence data-reuse pathways.

                      PRIMARY HYPOTHESIS & BENCHMARK SELECTION
                                         |
     +-----------------------------------+-----------------------------------+
     |                                                                       |
[Superiority Claim]                                                [Equivalence Claim]
     |                                                                       |
     +-----------------------+                                               +-----------------------+
     |                       |                                               |                       |
(Concurrent Control)   (Historical Benchmark)                           (Therapeutic Outcome)     (IVD / Measurement)
  RCT Design              Single-Arm Study                                 TOST Procedure           CLSI EP09 / Bland-Altman
  H0: mu_T - mu_C <= 0    H0: P_test <= OPC                              H01: diff <= -delta      Mean Bias +- 1.96 SD
  95% CI > 0              95% LCB > OPC                                  H02: diff >= +delta      Limits within TAE
  Two-sided alpha 0.05    FDA 2013 Guidance Sec 7.6.1                    90% CI within [-d, +d]   Assay Migration / EP09-A3

What is a Superiority Trial Design for a Medical Device, and When Should You Use It?

Formal Hypothesis Structure and Alpha Allocation

A superiority clinical trial is designed to demonstrate that an investigational medical device provides a clinically meaningful and statistically significant treatment advantage over a control intervention (active comparator, sham control, or historical benchmark).

Formally, for a continuous primary endpoint where a higher value represents a favorable clinical outcome, let $\mu_T$ represent the mean response in the test device group and $\mu_C$ represent the mean response in the control group. The superiority hypothesis is structured as:

$$H_0: \mu_T - \mu_C \le 0 \quad \text{versus} \quad H_1: \mu_T - \mu_C > 0$$

For binary endpoints (such as treatment success rate $P_T$ versus $P_C$), the null hypothesis states that $P_T - P_C \le 0$.

Under regulatory standards established by ICH E9 (Statistical Principles for Clinical Trials) and FDA CDRH guidelines, superiority testing requires strict control of the Type I error rate ($\alpha$). By standard convention:

  • Two-Sided Alpha Standard: Superiority is evaluated at a two-sided significance level of $\alpha = 0.05$.
  • One-Sided Alpha Equivalent: Equivalently, a one-sided hypothesis test at $\alpha = 0.025$ may be specified in the protocol and Statistical Analysis Plan (SAP).
  • Confidence Interval Decision Rule: Superiority is demonstrated if and only if the lower bound of the two-sided $95%$ confidence interval for the treatment difference ($\mu_T - \mu_C$) is strictly greater than zero (or greater than a pre-specified clinical superiority margin $\delta_{\text{sup}} > 0$).

Superiority Against Concurrent Active Controls vs. Historical Benchmarks

In randomized controlled clinical trials (RCTs), the investigational device is compared directly to a concurrent active control (e.g., an established catheter ablation system or surgical prosthetic). When a concurrent control group is randomized 1:1 or 2:1 with the test device, trial execution aligns with conventional statistical methods.

However, in medical device development, concurrent randomized controls are frequently unfeasible due to rapid technological iteration, clinical equipoise challenges, or surgical ethics. In these circumstances, FDA CDRH permits sponsors to execute single-arm superiority trials comparing device performance against a fixed historical benchmark:

$$\text{Superiority vs. Historical Benchmark: } H_0: P_{\text{device}} \le \text{Benchmark} \quad \text{vs.} \quad H_1: P_{\text{device}} > \text{Benchmark}$$

The sponsor establishes superiority by demonstrating that the lower bound of the one-sided $97.5%$ (or two-sided $95%$) confidence interval for the device success rate exceeds the historical benchmark.

Superiority Margins vs. Super-Superiority

While standard superiority tests against a null difference of zero ($H_0: \mu_T - \mu_C \le 0$), certain regulatory scenarios require demonstrating that the device exceeds the control by a minimum clinically relevant threshold $\delta_{\text{sup}} > 0$. This design—sometimes termed "super-superiority"—requires:

$$H_0: \mu_T - \mu_C \le \delta_{\text{sup}} \quad \text{versus} \quad H_1: \mu_T - \mu_C > \delta_{\text{sup}}$$

Super-superiority is typically employed when an investigational device incurs higher procedural risk, elevated toxicity, or substantially greater cost than existing clinical therapies, requiring proof of incremental clinical benefit to justify regulatory authorization or reimbursement coverage. Sponsors establishing sample sizes for superiority trials should cross-reference formulas and power curves detailed in our guide on sample size calculation for medical device clinical investigations.


How Do You Design and Analyze an Equivalence Trial Using TOST and the Confidence-Interval Rule?

The Two One-Sided Tests (TOST) Procedure

An equivalence trial aims to demonstrate that an investigational device is clinically and analytically indistinguishable from a reference comparator. Unlike non-inferiority trials—which enforce a one-sided boundary to prove the test device is "not unacceptably worse"—equivalence trials enforce symmetric upper and lower boundaries ($-\delta$ and $+\delta$) to prove the device is neither worse nor better than the control by an amount exceeding the equivalence margin $\delta$.

The formal statistical framework for equivalence relies on the Two One-Sided Tests (TOST) formulation established by Schuirmann (1987) and expanded by Walker & Nowacki (2010, Journal of General Internal Medicine, PMC3019319). The null hypothesis of non-equivalence ($H_0$) is expressed as a composite hypothesis comprising two distinct one-sided null hypotheses:

$$H_0: H_{01} \cup H_{02}$$

Where:

  • $H_{01}: \mu_T - \mu_C \le -\delta$ (The test device is inferior to the control by at least $\delta$)
  • $H_{02}: \mu_T - \mu_C \ge +\delta$ (The test device is superior to the control by at least $\delta$)

To claim equivalence, the analyst must reject both one-sided null hypotheses simultaneously in favor of the alternative hypothesis $H_1$:

$$H_1: -\delta < \mu_T - \mu_C < +\delta$$

                           TOST EQUIVALENCE MARGIN & CONFIDENCE INTERVAL
                                    
        Reject H01 (Inferiority)                  Reject H02 (Superiority)
                 =======>                                <=======
   +-----------------------------------------------------------------------------------+
   |  Non-Equivalent (Inferior) |      EQUIVALENT REGION       | Non-Equivalent (Superior) |
   +-----------------------------------------------------------------------------------+
   ^                            ^                               ^                      ^
 $-\delta$                      0                            $+\delta$               Difference
 
                       |----------- 90% CI -----------|
                       (Entire CI must fall inside [-d, +d])

The $90%$ Confidence Interval Decision Rule

Because TOST requires rejecting two separate one-sided hypotheses each at significance level $\alpha$, the overall Type I error rate for the equivalence procedure is maintained strictly at $\alpha$ by operationalizing the test through a single confidence interval.

Under the TOST procedure at nominal significance level $\alpha = 0.05$:

  1. Calculate the two-sided $(1 - 2\alpha) = 90%$ confidence interval for the treatment difference ($\mu_T - \mu_C$).
  2. Compare the $90%$ confidence interval endpoints $[L_{90}, U_{90}]$ to the pre-specified equivalence boundaries $[-\delta, +\delta]$.
  3. Declare equivalence if and only if:

$$-\delta < L_{90} \quad \text{and} \quad U_{90} < +\delta$$

If either boundary of the $90%$ confidence interval touches or crosses $-\delta$ or $+\delta$, the trial fails to establish equivalence.

Methodological Note on Alpha Conventions: Although a $90%$ confidence interval corresponds to two one-sided tests at $\alpha = 0.05$, some conservative regulatory bodies or clinical journals request a $95%$ confidence interval for equivalence reporting. Utilizing a $95%$ confidence interval for TOST corresponds to one-sided tests at $\alpha = 0.025$, which slightly increases sample size requirements while offering additional regulatory conservatism.

Sample Size Sensitivity to the Equivalence Margin

Equivalence trials require substantially larger sample sizes than superiority or non-inferiority trials because sample size $N$ scales inversely with the square of the equivalence margin $\delta$. Furthermore, sample size calculations must account for any true expected delta ($\Delta = \mu_T - \mu_C$) between treatment arms. If the device and control differ slightly in reality ($\Delta \ne 0$), the effective margin narrows to $\delta - |\Delta|$, driving sample size requirements exponentially higher.

Walker & Nowacki (2010, Table 2) provided a classic empirical demonstration of equivalence sample size scaling for two independent proportions (baseline success rate $\approx 28%$ vs $33%$, $\alpha = 0.05$, $80%$ power):

  • Setting an equivalence margin $\delta = 0.12$ ($12%$ margin) requires $N = 535$ patients per group.
  • Halving the equivalence margin to $\delta = 0.06$ ($6%$ margin) increases the required sample size to $N = 26,185$ patients per group—nearly a 50-fold increase.

Sponsors must carefully balance clinical justification with sample size feasibility when negotiating equivalence margins with FDA CDRH during Pre-Submission (Q-Submission) meetings.


Recommended Reading
FDA Medical Device Development Tools (MDDT): Program & Qualified Tools Guide
Clinical Evidence Regulatory2026-07-29 · 19 min read

What is an Objective Performance Criterion (OPC) and Performance Goal, and Why is Benchmark 'Superiority' the Dominant Device Pattern?

FDA Guidance Definitions: OPC vs. Performance Goal (PG)

In 2013, FDA CDRH published its landmark final guidance, Design Considerations for Pivotal Clinical Investigations for Medical Devices (Section 7.6.1). This guidance formalized historical benchmark comparisons as an accepted alternative to concurrent randomized controls in pivotal device trials:

Objective Performance Criterion (OPC): "An Objective Performance Criterion (OPC) is a numerical target value derived from historical data from clinical studies and/or registries that is used for comparison with the outcome of a single-arm study... OPCs are typically used in a dichotomous (pass/fail) manner and are generally established when the clinical progression of a disease and treatment outcomes are well documented."

While the terms OPC and Performance Goal (PG) are frequently used interchangeably in industry discussions, regulatory science maintains an important hierarchy:

Feature Objective Performance Criterion (OPC) Performance Goal (PG)
Evidence Maturity High maturity derived from established registry networks, standardized historical clinical trials, or consensus guidelines. Emerging or moderate maturity; derived from limited clinical literature or exploratory feasibility studies.
Regulatory Status Formally established or recognized by FDA CDRH panel guidance for a specific device category. Defined by the trial sponsor for a specific study protocol; requires protocol-by-protocol regulatory justification.
Endpoint Focus Predominantly effectiveness endpoints and major primary composite safety endpoints (e.g., valve safety). Frequently safety endpoints, adverse event rate upper bounds, or secondary performance targets.
Statistical Test Single-arm superiority test against fixed numerical target: $H_0: P \le \text{OPC}$. Single-arm test against performance target: $H_0: P \le \text{PG}$ or $H_0: \text{Safety Rate} \ge \text{PG}$.

Empirical Reality: Mooghali et al. (2025, JAMA Network Open) Approval Analysis

Industry literature often portrays pivotal device trials as randomized controlled trials comparing an investigational device to an active comparator. However, recent empirical research reveals that single-arm historical benchmark superiority is the actual dominant paradigm in high-risk US medical device authorization.

In a comprehensive empirical study published in JAMA Network Open, Mooghali et al. (2025; doi:10.1001/jamanetworkopen.2025.6230) analyzed the pivotal-study evidence behind original high-risk therapeutic medical device Pre-Market Approvals (PMAs) granted by FDA CDRH between 2019 and 2023. Of the $101$ such approvals, $13$ ($12.9%$) were granted without any pivotal study, leaving $88$ approvals with analyzable pivotal evidence. Their analysis revealed a striking reliance on nonconcurrent historical benchmarks:

          FDA HIGH-RISK THERAPEUTIC DEVICE PMAs (2019-2023, N = 88)
          
   +-------------------------------------------------------------------+
   |  Nonconcurrent Controls Used: 52 PMAs (59.1%)                      |
   |    By approval: 39 PG | 6 Historical | 6 OPC | 1 Mixed             |
   |    Across 79 primary analyses: 63 PG | 9 Historical | 7 OPC         |
   +-------------------------------------------------------------------+
   |  No Nonconcurrent Controls: 36 PMAs (40.9%)                        |
   |    (33 concurrent-control + 3 single-group without nonconcurrent)  |
   +-------------------------------------------------------------------+

   Key Methodological Findings from Mooghali et al. (2025):
   * Single-Group Study Design: 42 of 52 (80.8%) nonconcurrent control PMAs were single-arm trials.
   * Justification Deficit: Only 3 of 79 (3.8%) primary analyses provided formal statistical justification
     for utilizing a nonconcurrent control instead of a concurrent randomized control.
   * Regulatory Precedent Gap: 0 of 63 (0%) Performance Goals had been previously established by FDA;
     4 of 7 (57.1%) Objective Performance Criteria had prior formal FDA establishment.

The Mooghali 2025 aggregate demonstrates that while nonconcurrent benchmarks power nearly $60%$ of high-risk device approvals, FDA review divisions are increasingly scrutinizing the statistical justification and historical constancy of these benchmarks. Sponsors submitting single-arm OPC or PG superiority designs must proactively address benchmark derivation, historical drift, and patient selection bias in their IDE protocols.

Classic Device OPC Case Study: Mechanical & Bioprosthetic Heart Valves

The historical foundation of the OPC benchmark paradigm originated in interventional cardiology and cardiothoracic surgery. In 2006, Grunkemeier et al. published their classic analysis in The Annals of Thoracic Surgery (82(3):776-780, PMID 16928482), titled Prosthetic Heart Valves: Objective Performance Criteria Versus Randomized Clinical Trial.

The authors documented how FDA CDRH and international standards committees established standardized OPC event rates (expressed as percent per patient-year for complications such as thromboembolism, valve thrombosis, major hemorrhage, structural valve deterioration, and endocarditis) based on pooled historical registry data from thousands of valve-years. Rather than requiring sponsors to conduct 2,000-patient RCTs against older surgical valves, FDA permitted single-arm PMA pivotal trials evaluated against established heart-valve OPCs. This framework was reaffirmed by Head et al. (2016, Circulation), illustrating how OPC benchmarks accelerate clinical access to life-saving technology while maintaining rigorous quantitative safety thresholds.

For novel cardiovascular platforms such as intracranial thrombus aspiration catheters (e.g., ClinicalTrials.gov NCT05119647), single-arm OPC designs continue to serve as the definitive pivotal trial pathway under FDA IDE authorization.


Why a Non-Significant Superiority Test Does NOT Prove Equivalence (and CI Interpretation Rules)

The Fallacy of Declaring Equivalence from $p > 0.05$

One of the most persistent errors in medical device literature and regulatory submissions is concluding that two treatments are "equivalent" or "equally effective" simply because a superiority trial failed to achieve statistical significance ($p > 0.05$).

This logical fallacy—often summarized as "absence of evidence is not evidence of absence"—violates fundamental statistical principles:

  1. Different Null Hypotheses: Superiority tests $H_0: \mu_T - \mu_C \le 0$ (assuming no difference until proven otherwise). Equivalence tests $H_0: |\mu_T - \mu_C| \ge \delta$ (assuming non-equivalence until proven otherwise).
  2. Burden of Proof: In a superiority trial, an underpowered study, small sample size, high dropout rate, or excessive outcome variance drives $p > 0.05$. Failing to reject $H_0$ may reflect poor study execution or inadequate sample size, rather than therapeutic sameness.
  3. Contradictory Empirical Results: Walker & Nowacki (2010) re-analyzed published trial datasets comparing naive superiority tests to formal TOST. In their evaluation of published comparisons (citing Barker et al.), 9 out of 21 trial outcomes ($42.9%$) generated contradictory conclusions between non-significant superiority $p$-values and formal equivalence tests.
                  CONFIDENCE INTERVAL INTERPRETATION TAILORED TO HYPOTHESES
                  
                             Inferiority       Equivalence      Superiority
                               Margin            Margin           Margin
                              $-\delta$             0            $+\delta$
                                  |                 |                 |
  (1) Superiority Demonstrated    |                 |        |--- 95% CI ---|
                                  |                 |                 |
  (2) Equivalence Demonstrated    |     |--- 90% CI ---|              |
                                  |                 |                 |
  (3) Non-Inferiority Only        |    |----- 95% CI -----|           |
                                  |                 |                 |
  (4) Inconclusive / Failed       |  |--------- 95% CI ---------|       |
                                  |                 |                 |

Comprehensive Confidence Interval Decision Rules

To ensure clarity during regulatory review and SAP specification, clinical teams must interpret confidence intervals relative to zero and the pre-specified margins ($-\delta$ and $+\delta$):

  1. Superiority Established: The two-sided $95%$ confidence interval lies entirely above zero ($L_{95} > 0$).
  2. Equivalence Established: The two-sided $90%$ confidence interval lies entirely within the equivalence bounds ($-\delta < L_{90}$ and $U_{90} < +\delta$).
  3. Non-Inferiority Established (Without Equivalence): The two-sided $95%$ confidence interval lies entirely above the non-inferiority margin ($L_{95} > -\delta$), but the upper bound $U_{95}$ extends past $+\delta$ or the lower bound $L_{95}$ falls below zero.
  4. Inconclusive / Failed Trial: The confidence interval spans across $-\delta$ or zero without meeting pre-specified boundary criteria.

For a detailed analysis of non-inferiority margin derivation ($M_1$ and $M_2$ preservation logic) and the formal switch between non-inferiority and superiority testing, refer to our companion guide on non-inferiority clinical trials for medical devices.


How is IVD and Sensor Measurement Equivalence Different from Therapeutic Equivalence?

Analytical vs. Clinical Equivalence: CLSI EP09-A3 and Bland-Altman

For In Vitro Diagnostics (IVDs), continuous glucose monitors (CGMs), wearable bio-sensors, and laboratory diagnostic instrumentation, "equivalence" does not refer to clinical trial outcome rates or TOST hypotheses. Instead, it refers to analytical measurement equivalence—demonstrating that a candidate measurement procedure (new assay or instrument) agrees with a reference measurement procedure or established predicate device within clinically acceptable limits.

Diagnostic measurement equivalence is governed by consensus standards established by the Clinical and Laboratory Standards Institute (CLSI), specifically CLSI EP09-A3: Measurement Procedure Comparison and Bias Estimation Using Patient Samples.

                      IVD MEASUREMENT EQUIVALENCE WORKFLOW (CLSI EP09-A3)
                                               |
     +-----------------------------------------+-----------------------------------------+
     |                                                                                   |
[Sample Collection]                                                      [Scatter & Bias Regression]
  >= 40-100 Patient Samples                                                Passing-Bablok / Deming
  Covering Full Measuring Range                                            Calculate Mean Bias & 95% CI
     |                                                                                   |
     +-----------------------------------------+-----------------------------------------+
                                               v
                                   [Bland-Altman Agreement]
                                   Plot Difference vs. Average
                                   Calculate 95% Limits of Agreement:
                                   Bias +- 1.96 * SD_diff
                                               |
                                               v
                                [Total Allowable Error Check]
                                Is |Bias| + 1.96 * SD_diff <= TAE?

Statistical Methods for Measurement Agreement

When conducting an IVD measurement equivalence study under CLSI EP09-A3:

  1. Sample Scope: Test a minimum of $40$ to $100$ individual patient samples (preferably run in duplicate over multiple days) spanning the full medically relevant measurement interval.
  2. Regression Analysis: Ordinary Least Squares (OLS) regression is inappropriate because both the candidate and reference methods contain measurement error. Sponsors must use Deming Regression (when error variance ratio is known) or Passing-Bablok Regression (non-parametric regression robust to outliers).
  3. Bland-Altman Analysis: Published originally by Bland & Altman (1986, The Lancet; 1999, Statistical Methods in Medical Research, PMID 10501650; $>11,500$ citations), this method plots the difference between paired measurements ($Y_i - X_i$) against their mean ($(X_i + Y_i)/2$).
  4. 95% Limits of Agreement: Calculate mean bias ($\bar{d}$) and the standard deviation of differences ($SD_d$). The $95%$ limits of agreement are defined as:

$$\text{Limits of Agreement} = \bar{d} \pm 1.96 \times SD_d$$

  1. Total Allowable Error ($\text{TAE}$): Measurement equivalence is established if the $95%$ limits of agreement fall entirely within the predefined Total Allowable Error ($\text{TAE}$) specification established by clinical practice guidelines or CLIA criteria.

Regulatory Application: FDA Assay Migration Guidance

FDA CDRH explicitly incorporates CLSI EP09 and Bland-Altman agreement in its guidance document Assay Migration Studies for In Vitro Diagnostic Devices. When a manufacturer transfers an approved diagnostic assay to a new automated analyzer or updates reagent formulations, FDA allows the sponsor to submit an assay migration study validating measurement equivalence rather than conducting new clinical performance studies.

Real-world clinical trial registry examples illustrate how device equivalence splits across these two domains:

  • Masimo Rainbow SpHb sensor (NCT03134326 / NCT03157232) — measurement equivalence: These are non-invasive hemoglobin sensor sub-range performance equivalence studies in which a candidate Rainbow SpHb sensor (R2-25; reusable DCI) is compared against a predicate/reference SpHb sensor (R1-25) using method-comparison and agreement analysis — the IVD/measurement-equivalence paradigm described above, rather than a clinical-endpoint TOST.
  • DERMABOND PROTAPE (NCT00547638) — clinical-performance equivalence, NOT measurement equivalence: A randomized, multicenter study demonstrating equivalence of two topical skin adhesives (DERMABOND PROTAPE vs. DERMABOND HVD) for wound closure. This is a therapeutic / clinical-performance equivalence claim on a clinical endpoint, distinct from the CLSI EP09 / Bland-Altman measurement-agreement paradigm — included here precisely to mark that boundary.

Recommended Reading
Bayesian Statistics for Medical Device Clinical Trials: FDA Guidance & Approvals
Clinical Evidence Regulatory2026-08-07 · 17 min read

Assay Sensitivity, Constancy Assumption, and Control Selection (ICH E10 & FDA 2013)

The Critical Role of Assay Sensitivity

Under ICH E10 (Choice of Control Group and Related Issues in Clinical Trials) and FDA CDRH's 2013 pivotal investigation guidance, assay sensitivity is defined as the property of a clinical trial that permits it to distinguish between an effective treatment and a less effective or ineffective treatment.

                           ASSAY SENSITIVITY & HYPOTHESIS TYPES
                           
   +-----------------------------------------------------------------------------------+
   |  CONCURRENT SUPERIORITY TRIAL                                                     |
   |  Assay Sensitivity Tested Directly: Rejecting H0 proves trial had sensitivity.    |
   +-----------------------------------------------------------------------------------+
   |  EQUIVALENCE / NON-INFERIORITY / SINGLE-ARM OPC TRIAL                             |
   |  Assay Sensitivity MUST Be Assumed: Requires Constancy Assumption verification.    |
   +-----------------------------------------------------------------------------------+

In a concurrent randomized superiority trial, assay sensitivity is demonstrated internally: if the test device statistically beats the control ($p < 0.05$), the trial inherently possessed the sensitivity required to detect a difference.

Conversely, in an equivalence trial, a non-inferiority trial, or a single-arm OPC benchmark trial, assay sensitivity cannot be proven internally. If the test device and control show equal outcomes, it may mean both treatments were highly effective—or it may mean the trial was poorly conducted, patient compliance was dismal, endpoint measurement was noise-dominated, or the historical control effect has eroded over time.

The Constancy Assumption and Benchmark Validity

When executing single-arm OPC superiority or active-control equivalence designs, sponsors implicitly rely on the constancy assumption: the premise that the historical effect of the benchmark or active control remains unchanged in the current trial setting.

Historical constancy can break down due to:

  • Evolution in Background Medical Therapy: Concurrent background medical therapy (e.g., modern antiplatelet regimens in stent trials) improves patient baseline outcomes, making historical control event rates outdated.
  • Diagnostic Shift: Improved imaging techniques detect milder disease stages, altering the baseline risk profile of enrolled cohorts.
  • Investigator Learning Curves: Surgical technique improvements alter complication rates over time.

To maintain benchmark validity, sponsors utilizing OPCs or performance goals must systematically audit historical dataset freshness and document constancy in their IDE submission.


Disambiguation: Statistical Equivalence Trials vs. EU MDR Article 61 Clinical-Equivalence Data Reuse

A major point of confusion for regulatory professionals working across US and European markets is the term "equivalence." It is essential to disambiguate statistical equivalence clinical trials from EU MDR clinical equivalence data reuse:

                              THE EQUIVALENCE DISAMBIGUATION MATRIX
                              
   Feature              Statistical Equivalence Trial              EU MDR Article 61 Data Reuse
   -------------------  -----------------------------------------  -----------------------------------------
   Primary Domain       Clinical Trial Design & Biostatistics      European Regulatory Access (CE Marking)
   Governing Framework  ICH E9 / TOST / CLSI EP09                  EU MDR 2017/745 / MDCG 2020-5 Guidance
   Core Objective       Prove 2 treatments have equal clinical     Claim technical, biological, and clinical
                        or measurement performance via statistics  sameness to reuse predicate clinical data
   Mechanism            Prospective Clinical Investigation         Technical Documentation & Evaluation
   Output               $90\%$ Confidence Interval within $[-\delta, +\delta]$  MDCG 2020-5 Demonstration Table

Under Regulation (EU) 2017/745 (EU MDR) Article 61(4) and Annex XIV, a manufacturer may claim "equivalence" to an existing CE-marked device to utilize the predicate's clinical data for its own Clinical Evaluation Report (CER). MDCG 2020-5 enforces strict technical, biological, and clinical identity criteria. This is a regulatory data-use pathway, completely distinct from conducting a prospective statistical equivalence clinical trial.

Sponsors seeking guidance on European data-reuse requirements should consult our dedicated guide on clinical equivalence assessment under EU MDR (MDCG 2020-5).


Regulatory Justification Playbook: Securing IDE/PMA Approval for Single-Arm and Equivalence Designs

To ensure your superiority or equivalence trial protocol survives FDA CDRH Pre-Submission review and IDE assessment, execute the following three-step regulatory playbook:

Step 1: Justify Control Selection & Benchmark Provenance

If proposing a single-arm OPC or Performance Goal superiority design:

  • Document why a concurrent randomized control is ethically unfeasible or clinically impractical.
  • Provide a systematic literature review and registry audit demonstrating the stability and constancy of the historical benchmark.
  • Reference established FDA CDRH panel guidance or recognized consensus standards (e.g., ISO 14155 clinical investigation principles) supporting OPC usage for your device code. Cross-reference regulatory requirements in our guide on FDA IDE requirements for medical device clinical trials and ISO 14155 GCP standards.

Step 2: Formally Pre-Specify Margins and Alpha Allocation in the SAP

  • For equivalence trials, pre-specify the symmetric margin $\delta$ and justify its clinical acceptability based on clinical consensus and historical variability.
  • Define the exact confidence interval calculation ($(1 - 2\alpha) = 90%$) and explicit TOST software procedures in the SAP prior to protocol freeze.
  • Incorporate sensitivity analyses and estimand definitions aligning with ICH E9(R1) to address missing data and intercurrent events.

Step 3: Address Feasibility and Early Signals


Recommended Reading
Auto-Injector Critical-Task Matrix for Human Factors Validation
Design Controls Clinical Evidence2026-05-05 · 21 min read

Frequently Asked Questions (FAQ)

Can I claim equivalence if my superiority test came back non-significant ($p > 0.05$)?

No. A non-significant superiority result indicates only that the trial failed to demonstrate a statistically significant difference between treatments under $H_0: \mu_T - \mu_C \le 0$. It does not prove sameness. To claim equivalence, you must pre-specify an equivalence margin $\delta$ in the SAP, execute the TOST procedure, and demonstrate that the entire $90%$ confidence interval for the treatment difference falls strictly within $[-\delta, +\delta]$. Walker & Nowacki (2010) demonstrated that naive superiority $p$-values and formal TOST produce contradictory conclusions in over $40%$ of trial comparisons.

What confidence interval level is required for a TOST equivalence trial?

Standard regulatory TOST procedures at significance level $\alpha = 0.05$ utilize a two-sided $90%$ confidence interval ($(1 - 2\alpha)$). If the entire $90%$ confidence interval lies within $[-\delta, +\delta]$, equivalence is declared at $\alpha = 0.05$. However, if regulatory reviewers request a $95%$ confidence interval, the test operationalizes TOST at $\alpha = 0.025$, requiring a slightly larger sample size.

What is the exact difference between an OPC and a Performance Goal (PG)?

An Objective Performance Criterion (OPC) is a numerical target value derived from mature historical clinical trial or registry data, typically recognized in FDA guidance or panel standards for a well-characterized device class (e.g., heart valves). A Performance Goal (PG) is a performance target defined by a sponsor for a specific study where historical data is less mature or where safety upper bounds are evaluated. Mooghali et al. (2025) found that while $59.1%$ of high-risk device PMAs used nonconcurrent controls, $0%$ of Performance Goals had prior FDA establishment, whereas $57.1%$ of OPCs had formal FDA establishment.

How does IVD measurement equivalence differ from clinical outcome equivalence?

IVD measurement equivalence evaluates analytical agreement between a new assay/instrument and a reference measurement procedure using clinical specimens across the measuring range. It is governed by CLSI EP09-A3 and Bland-Altman $95%$ limits of agreement ($\bar{d} \pm 1.96 \text{ SD}$) compared against a Total Allowable Error ($\text{TAE}$), rather than clinical trial TOST hypothesis testing.

Do I need to demonstrate assay sensitivity for a superiority trial?

In a concurrent randomized superiority trial, assay sensitivity is demonstrated internally if the test device beats the control ($p < 0.05$). However, for equivalence trials, non-inferiority trials, or single-arm OPC benchmark studies, assay sensitivity cannot be proven internally and relies entirely on the constancy assumption—requiring sponsors to prove that historical control effect sizes remain valid in the current clinical environment (ICH E10).


Conclusion & Implementation Checklist

Designing a successful superiority or equivalence clinical investigation for a medical device requires aligning clinical claims, statistical hypothesis structures, regulatory guidance, and empirical benchmark data. By distinguishing therapeutic clinical hypotheses from IVD analytical measurement agreement—and proactively justifying nonconcurrent control benchmarks—sponsors can construct robust pivotal trial protocols that navigate FDA CDRH review.

Implementation Checklist for Clinical & Regulatory Teams

  • Define Primary Claim: Clarify whether the regulatory claim is superiority over control, super-superiority ($>\delta_{\text{sup}}$), therapeutic equivalence (TOST), or IVD measurement agreement (CLSI EP09).
  • Select Control Architecture: Determine whether a concurrent randomized control is feasible, or justify a single-arm OPC/Performance Goal design under FDA CDRH 2013 Pivotal Investigation guidance.
  • Perform Benchmark Provenance Audit: If using an OPC or PG, document historical dataset freshness, address historical drift, and provide statistical justification to withstand scrutiny under Mooghali 2025 standards.
  • Specify Alpha & Confidence Intervals: In the SAP, pre-specify two-sided $\alpha = 0.05$ ($95%$ CI) for superiority, or TOST at $\alpha = 0.05$ ($90%$ CI within $[-\delta, +\delta]$) for equivalence.
  • Calculate Sample Size: Account for margin scaling ($\delta^2$) and non-zero true delta ($\Delta$), cross-referencing power calculations in our sample size calculation guide.
  • Execute IVD Agreement Methods: For diagnostic assays and sensors, execute Deming/Passing-Bablok regression and Bland-Altman $95%$ limits of agreement under CLSI EP09-A3 against Total Allowable Error.
  • Disambiguate Regulatory Concepts: Ensure clinical protocols crisply separate statistical equivalence testing from EU MDR Article 61 clinical equivalence data reuse.