MedDeviceGuideMedDeviceGuide
Back

External and Historical Control Arms in Medical Device Clinical Trials

Guide to external and historical control arms in medical device trials: FDA acceptance, OPC vs external controls, MAIC, and exchangeability rules.

Ran Chen
Ran Chen
Global MedTech Expert | 10× MedTech Global Access
Published 2026-08-12Last reviewed 2026-08-1225 min read

When designing a pivotal clinical trial for a novel medical device, establishing an appropriate control group is a primary determinant of study feasibility, scientific validity, and regulatory approval. While a concurrent randomized controlled trial (RCT) with active or sham control remains the gold standard for clinical evidence, running a randomized study is frequently impractical, ethically unviable, or methodologically flawed in medical device evaluation. In rare disease indications, pediatric populations, humanitarian device exemptions, or situations where surgical sham controls pose unacceptable risk of harm without therapeutic benefit, sponsors must look to nonconcurrent comparative data.

In these settings, sponsors increasingly propose external control arms (including historical controls, registry-derived comparator cohorts, and synthetic control arms) or Objective Performance Criteria (OPC). However, designing an externally controlled device trial requires navigating a nuanced regulatory landscape marked by structural guidance gaps, stringent statistical justification requirements, and subtle methodological distinctions between fixed benchmark numbers and constructed patient cohorts.

                    ┌──────────────────────────────────────────────────────────┐
                    │      Is a Concurrent Randomized or Sham Control          │
                    │                  Feasible & Ethical?                     │
                    └────────────────────────────┬─────────────────────────────┘
                                                 │
                        ┌────────────────────────┴────────────────────────┐
                        │                                                 │
                        YES                                               NO
                        ▼                                                 ▼
          ┌───────────────────────────┐                     ┌───────────────────────────┐
          │  Concurrent RCT / Sham    │                     │   Does Patient-Level Data     │
          │   Standard Design         │                     │  Exist for Historical Arm?│
          └───────────────────────────┘                     └─────────────┬─────────────┘
                                                                          │
                                                ┌─────────────────────────┴────────────────────────┐
                                                │                                                  │
                                                YES                                                NO
                                                ▼                                                  ▼
                                  ┌───────────────────────────┐                      ┌───────────────────────────┐
                                  │ Constructed External Arm  │                      │   Fixed Benchmark (OPC)   │
                                  │ (Propensity / Registry /  │                      │  or Performance Goal (PG) │
                                  │   Synthetic Control Arm)  │                      │  (CDRH 2013 Guidance)     │
                                  └───────────────────────────┘                      └───────────────────────────┘

This guide details when the FDA and international authorities accept external controls in medical device pivotal trials, analyzes the structural differences between an OPC benchmark and a constructed external cohort, evaluates six primary construction methodologies, and outlines protocol-level strategies for demonstrating exchangeability and satisfying the constancy assumption.


Executive Summary & Direct Answer

Scenario Question: We are designing a single-arm pivotal study for a first-in-human or high-risk Class III device in a small patient population where a concurrent randomized or sham control is not ethical or feasible. Under what conditions will the FDA accept an external or historical control, how do we construct the control arm, and how does an external control arm differ from an Objective Performance Criterion benchmark?

Direct Answer: The FDA accepts an external or historical control for a medical device trial when a concurrent randomized or sham control is unviable or unethical (e.g., severe ethical barriers to sham intervention, rare disease populations, high unmet medical need, or well-characterized device categories with mature historical performance data), provided the sponsor pre-specifies the analysis, establishes baseline exchangeability between external and investigational cohorts, and demonstrates that the constancy assumption holds (proving the historical effect has not drifted over time).

Crucially, sponsors must distinguish between an Objective Performance Criterion (OPC)—a single, fixed numerical benchmark target derived from pooled historical data as defined in Section 7.6.1 of the FDA CDRH 2013 guidance (Design Considerations for Pivotal Clinical Investigations for Medical Devices)—and a constructed external control arm, which is a dynamic cohort of real-world or historical patients adjusted via propensity scoring or matching-adjusted indirect comparison (MAIC).

Empirical evidence shows that nonconcurrent controls are widely used in device approvals but poorly justified statistically. A 2025 cross-sectional study of original high-risk therapeutic device PMA approvals between 2019 and 2023 published in JAMA Network Open by Mooghali et al. revealed that 59.1% (52 of 88) relied on nonconcurrent controls (39 performance goals, 6 historical controls, 6 OPCs, and 1 mixed design), yet only 3.8% (3 of 79 analyses) provided formal statistical justification for the selected comparator. Furthermore, the official FDA draft guidance issued on February 1, 2023 (Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products, FR Doc 2023-02094, Docket FDA-2022-D-2983) covers human drugs and biological products under CDER and CBER scope; there is no CDRH device-specific equivalent guidance. Medical device sponsors must therefore construct external control arms by synthesizing CDRH's 2013 pivotal investigation guidance, ICH E10 control-group principles, and device-specific empirical precedents.


What is an external or historical control arm, and how does it differ from an OPC benchmark?

Understanding control arm terminology is essential to avoiding critical protocol classification errors during FDA pre-submission (Q-Submission) interactions.

Control Group Terminology

  1. Historical Control Group: A comparator cohort composed of patients treated in an earlier time period, drawn from prior clinical trials, published literature, or retrospective hospital records. The control patients were managed under standard-of-care practices existing at that past time.
  2. Concurrent External Control Group: A comparator cohort treated during the same timeframe as the pivotal device study, but outside the trial protocol (e.g., patients enrolled in a parallel prospective real-world registry or receiving standard medical management at non-investigational centers).
  3. Synthetic Control Arm (SCA): A constructed cohort created by matching real-world patient-level data (from electronic health records, claims databases, or disease registries) or historical trial data to investigational trial subjects using advanced causal inference algorithms (such as propensity score matching, entropy balancing, or machine learning target trial emulation).
  4. Objective Performance Criterion (OPC): A fixed, numerical rate or threshold parameter (e.g., "a 1-year major adverse event rate of less than 12.5%") established by FDA guidance or broad consensus from extensive historical trial data across a well-characterized device class. An OPC is evaluated as a single-arm pass/fail hypothesis test.
  5. Performance Goal (PG): A numerical benchmark defined by an individual sponsor when formal FDA-established OPCs do not exist, derived from literature or preliminary data but lacking prior formal regulatory establishment.

OPC Benchmark vs. Constructed External Control Arm

Although both eliminate the need to enroll concurrent control subjects within the pivotal trial, an OPC and an external control arm differ fundamentally in structure, statistical handling, and data requirements:

Dimension Objective Performance Criterion (OPC) / Performance Goal Constructed External Control Arm
Data Format Single fixed numerical benchmark (rate or mean threshold). Patient-level dataset (individual row-per-patient data).
Statistical Testing Single-arm test against fixed constant ($H_0: p \ge p_0$ vs $H_1: p < p_0$). Two-sample statistical comparison ($H_0: \mu_{\text{dev}} = \mu_{\text{ctrl}}$).
Covariate Adjustment Fixed rate applies to a broad target population; no individual matching. Individual-level propensity matching, weighting (IPTW), or MAIC.
FDA Precedent Well-established for specific matured classes (e.g., heart valves, coronary stents). Evaluated case-by-case; common in rare disease and high unmet need.
Sample Size Impact Minimizes total sample size; requires zero control subjects. Requires external dataset with sufficient sample size for effective matching.
Primary Regulatory Risk Selection of an obsolete or non-conservative benchmark number. Residual confounding, unmeasured covariates, and temporal drift.

Sponsors seeking to establish a single-arm superiority or equivalence hypothesis test using fixed rates should consult our dedicated guide on superiority and equivalence trial design with OPC and performance-goal benchmarks.


When will the FDA accept an external or historical control instead of a concurrent randomized control?

Under FDA regulations (21 CFR 812 for Investigational Device Exemptions and 21 CFR 814 for Premarket Approvals), clinical investigations must provide "valid scientific evidence" to determine device safety and effectiveness. While randomized controls provide maximum protection against bias, the FDA recognizes that randomized or sham controls are not always viable.

4 Pillars of Regulatory Acceptability

To gain FDA agreement during an IDE submission or Q-Submission, an external control arm design must satisfy four core conditions:

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                        4 PILLARS OF FDA ACCEPTABILITY FOR EXTERNALLY                   │
│                                  CONTROLLED DEVICE TRIALS                              │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                            │
    ┌───────────────────────┬───────────────┴───────────────┬───────────────────────┐
    ▼                       ▼                               ▼                       ▼
┌──────────────────┐    ┌──────────────────┐    ┌──────────────────┐    ┌──────────────────┐
│   1. ETHICAL &   │    │ 2. PROTOCOL PRE- │    │3. DEMONSTRABLE   │    │  4. CONSTANCY    │
│  PRACTICAL NEED  │    │  SPECIFICATION   │    │ EXCHANGEABILITY  │    │   ASSUMPTION     │
├──────────────────┤    ├──────────────────┤    ├──────────────────┤    ├──────────────────┤
│• Unacceptable    │    │• Control source, │    │• Overlapping     │    │• Standard of care│
│  sham risk       │    │  eligibility, &  │    │  baseline        │    │  is unchanged    │
│• Ultra-rare      │    │  analytic SAP    │    │  covariates      │    │• Diagnostic      │
│  population      │    │  finalized prior │    │• Measured &      │    │  criteria match  │
│• No standard     │    │  to inspecting   │    │  unmeasured      │    │• No version      │
│  comparator      │    │  device outcomes │    │  confounding     │    │  drift           │
└──────────────────┘    └──────────────────┘    └──────────────────┘    └──────────────────┘
  1. Demonstrated Unfeasibility of Concurrent Control:

    • Ethical Barrier: Performing a sham surgical or invasive intervention (e.g., burr-hole craniotomy, intra-cardiac catheterization, or spinal implantation) exposes subjects to surgical risk without prospective therapeutic benefit.
    • Feasibility Barrier: The target condition is an ultra-rare disease (e.g., Humanitarian Device Exemption candidate populations) where enrolling a 1:1 or 2:1 randomized control arm would require multi-decade recruitment timelines.
    • Lack of Standard Comparator: No established medical device or pharmaceutical standard of care exists for the disease entity.
  2. Protocol Pre-Specification:

    • The FDA mandates that the external control source, patient selection criteria, matching methodology, statistical analysis plan (SAP), and primary endpoints must be finalized before analyzing pivotal device trial outcomes. Post-hoc construction of historical control arms is viewed by CDRH reviewers as cherry-picking and routinely triggers deficiency decisions.
  3. Demonstrable Baseline Exchangeability:

    • The investigational device cohort and external control subjects must be comparable across all key prognostic factors (age, disease severity, baseline functional score, comorbidities, prior therapies, and anatomical sub-types).
  4. Satisfaction of the Constancy Assumption:

    • Derived from ICH E10 principles (and closely related to the non-inferiority design and the constancy assumption), the constancy assumption requires that historical control performance remains valid in the present study environment. If background medical care, diagnostic sensitivity, or adjunctive pharmacological therapy has evolved, historical control outcomes cannot serve as a reliable counterfactual.

Empirical FDA Precedent: Acceptance vs. Justification Deficits

Empirical evidence demonstrates both the necessity and the regulatory vulnerabilities of external control arms in medical device approvals:

  • Jahanshahi et al. (2021): In an empirical review published in Therapeutic Innovation and Regulatory Science, Jahanshahi and colleagues identified 45 FDA product approvals where regulatory decision-making accepted external control data in benefit-risk assessments. FDA acceptance was overwhelmingly clustered in severe, life-threatening, or rare conditions lacking satisfactory alternative therapies.
  • Mooghali et al. (2025): A comprehensive cross-sectional study in JAMA Network Open analyzed original high-risk therapeutic device Premarket Approvals (PMAs) granted by FDA CDRH between 2019 and 2023. Of 88 analyzable original PMAs:
    • 59.1% (52 of 88) relied on nonconcurrent controls.
    • Nonconcurrent study designs comprised 39 performance goals, 6 historical controls, 6 objective performance criteria, and 1 mixed design.
    • 80.8% (42 of 52) were conducted as single-group studies.
    • The Justification Gap: Crucially, only 3.8% (3 of 79 analyses) provided formal statistical justification for the comparator chosen. Furthermore, while 4 of 7 OPC comparators possessed formal prior FDA establishment, 0 of 63 performance goals had prior regulatory establishment.

This 3.8% statistical justification rate highlights why CDRH panel reviews and IDE deficiency letters increasingly target nonconcurrent control selection. Sponsors must move beyond narrative claims of "historical precedent" to rigorous, pre-specified statistical frameworks.


Recommended Reading
Bayesian Statistics for Medical Device Clinical Trials: FDA Guidance & Approvals
Clinical Evidence Regulatory2026-08-07 · 17 min read

Which construction method should a device sponsor choose?

When designing an externally controlled device study, sponsors must select a construction methodology aligned with data availability, patient-level access, and regulatory expectations.

                                  ┌──────────────────────────────────────────────────────────┐
                                  │           EXTERNAL CONTROL CONSTRUCTION DECISION         │
                                  └────────────────────────────┬─────────────────────────────┘
                                                               │
                                    ┌──────────────────────────┴──────────────────────────┐
                                    │                                                     │
                        Patient-Level Data Accessible?                           Aggregate Data Only?
                                    │                                                     │
                        ┌───────────┴───────────┐                             ┌───────────┴───────────┐
                        │                       │                             │                       │
                Trial-to-Trial Data      Real-World Data               Published Trial Rates     Historical Literature
                        │                       │                             │                       │
                        ▼                       ▼                             ▼                       ▼
              ┌──────────────────┐    ┌──────────────────┐          ┌──────────────────┐    ┌──────────────────┐
              │ Propensity-Score │    │ Registry-Based / │          │ Matching-Adjusted│    │ Fixed Benchmark  │
              │  Matched Cohort  │    │ Synthetic Control│          │ Indirect Compar. │    │  (OPC / PG)      │
              │ (IPTW / Matching)│    │    Arm (SCA)     │          │   (MAIC / STC)   │    │ (CDRH 2013 Guid.)│
              └──────────────────┘    └──────────────────┘          └──────────────────┘    └──────────────────┘

6 Construction Methodologies Compared

The following decision matrix outlines the six primary methodologies for constructing or establishing external comparative data in medical device trials:

Methodology Data Requirements Statistical Mechanism Key Advantage Primary Regulatory Risk Ideal Use Case
1. Objective Performance Criterion (OPC) Pooled historical trial data across mature device class. Single-arm test against fixed rate threshold ($p_0$). Minimizes sample size; zero control subjects needed. Obsolete rate due to improving standard of care. Class III devices with mature regulatory precedent (e.g., heart valves).
2. Sponsor Performance Goal (PG) Published literature or preliminary pilot study data. Single-arm test against sponsor-defined threshold. High flexibility when no formal OPC exists. High FDA scrutiny (0% had prior FDA establishment in Mooghali 2025). Early feasibility or De Novo pathways for novel device concepts.
3. Registry-Based External Arm Prospective/retrospective disease or device registry. Direct cohort comparison with propensity matching. High external validity; real-world clinical setting. Missing data, variable diagnostic coding, protocol gaps. Post-market commitment studies or expansion of indication PMA supplements.
4. Propensity-Score Matched External Cohort Patient-level data from prior trial or curated database. 1:1 or 1:k matching / IPTW weighting on propensity scores. Balances measured baseline confounders effectively. Unmeasured confounding; loss of sample size from non-matches. Single-arm IDE pivotal studies with access to completed competitor/prior trials.
5. Matching-Adjusted Indirect Comparison (MAIC) Patient-level data for device; aggregate data for control. Re-weighting device subjects to match aggregate control vector. Enables comparison when patient-level control data is unavailable. Reduced effective sample size (ESS); extreme weights. Health Technology Assessment (HTA), EU MDR CE-mark, or global submissions.
6. Synthetic Control Arm (SCA) Deep EHR/claims databases with machine-learning curation. Target trial emulation and causal inference algorithms. Leverages large-scale real-world datasets. Data quality, variable outcome definitions, FDA skepticism of EHR. Rare disease device indications where prior trial data does not exist.

How does a sponsor justify exchangeability and the constancy assumption?

When submitting an externally controlled device trial protocol to FDA CDRH, demonstrating exchangeability—the principle that external control subjects would have experienced the same outcome distribution as investigational subjects had they received the same treatment—is the single most critical methodological hurdle.

Device-Specific Sources of Non-Exchangeability

Unlike pharmaceutical trials where drug molecules act consistently regardless of the administering clinician, medical device outcomes are inherently bound to human operator skills, surgical technique, and mechanical engineering iterations. Sponsors must systematically evaluate five device-specific confounding mechanisms:

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                   5 DEVICE-SPECIFIC CONFOUNDING MECHANISMS IN EXTERNAL CONTROLS        │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                            │
    ┌───────────────────────┬───────────────┼───────────────┬───────────────────────┐
    ▼                       ▼               ▼               ▼                       ▼
┌──────────────────┐    ┌──────────────┐┌──────────────┐┌──────────────────┐    ┌──────────────────┐
│  1. OPERATOR     │    │  2. DEVICE   ││ 3. EVOLVING  ││  4. DIAGNOSTIC   │    │  5. ANATOMICAL   │
│  LEARNING CURVE  │    │  ITERATION   ││ MEDICAL CARE ││   & SELECTION    │    │  & PROCEDURAL    │
├──────────────────┤    ├──────────────┤├──────────────┤├──────────────────┤    ├──────────────────┤
│Early procedural  │    │Hardware or   ││Background    ││Higher resolution ││Center volume,    │
│complications in  │    │software      ││medication    ││imaging selects   ││concomitant       │
│historical trials │    │upgrades alter││changes alter ││different baseline││devices, or dual  │
│skew baseline     │    │device risk   ││event rates   ││risk tiers        ││therapy variations│
│safety profile    │    │profile       ││over time     ││over time         ││skew outcomes     │
└──────────────────┘    └──────────────┘└──────────────┘└──────────────────┘    └──────────────────┘
  1. Operator Learning Curve Confounding: Historical control data from initial feasibility trials may incorporate operator learning curve complications (e.g., vascular access site bleeding or implant malpositioning). Comparing a refined second-generation device implanted by experienced operators against an unadjusted early historical control artificially inflates relative device efficacy.
  2. Device Version Drift: Engineering modifications between the historical device generation and the investigational device (e.g., changes in stent strut thickness, balloon material, or software delivery algorithms) break temporal comparability.
  3. Evolution of Background Standard of Care: Changes in adjunctive pharmacological therapy (e.g., transition from dual antiplatelet therapy to novel oral anticoagulants) alter baseline event rates independently of device performance.
  4. Diagnostic and Selection Drift: Improvements in imaging resolution (e.g., high-definition OCT vs. historical intravascular ultrasound) or biomarker sensitivity result in stage migration, where historical patients classified as "moderate disease" would today be categorized as "severe."
  5. Anatomical and Procedural Selection Criteria: Discrepancies in inclusion criteria regarding anatomical complexity (e.g., calcification scoring or vessel tortuosity) between external registries and trial protocols destroy comparability.

Quantitative Protocol Checklist for FDA Pre-Submissions

To prevent the 3.8% statistical justification deficit identified by Mooghali et al. (2025), sponsors should incorporate the following quantitative diagnostics into their Statistical Analysis Plan:

  • Pre-specified Causal Diagram (DAG): Include a Directed Acyclic Graph identifying all known prognostic variables, treatment effect modifiers, and potential confounders.
  • Standardized Mean Difference (SMD) Thresholds: Specify that baseline covariate balance between matched external controls and device subjects will achieve an SMD $< 0.10$ across all key demographic and anatomical variables.
  • Positivity Diagnostics: Evaluate propensity score distributions to verify sufficient overlap across treatment arms (ensuring every device subject has a non-zero probability of matching an external control subject).
  • Effective Sample Size (ESS) Reporting: For MAIC or weighted comparisons, pre-specify a minimum ESS floor to prevent high-variance estimates driven by extreme weights.
  • E-Value / Quantitative Sensitivity Analysis: Calculate E-values to quantify the minimal strength of unmeasured confounding that would be required to explain away the observed device-control treatment effect.

Sponsors using registry data to build external control cohorts should review our comprehensive guide on real-world evidence and registry-based controls for devices.


How do Bayesian dynamic borrowing and MAIC apply to device external controls?

When individual-level historical data is accessible but heterogeneous, or when only aggregate published data exists, advanced biostatistical techniques allow sponsors to construct valid external comparators while controlling bias.

Bayesian Dynamic Borrowing

Rather than completely pooling external data (which assumes perfect exchangeability) or completely ignoring historical controls, Bayesian dynamic borrowing adaptively weights historical information based on how closely the historical data aligns with incoming trial data.

                    ┌──────────────────────────────────────────────────────────┐
                    │               HISTORICAL DATA ALIGNMENT CHECK            │
                    └────────────────────────────┬─────────────────────────────┘
                                                 │
                        ┌────────────────────────┴────────────────────────┐
                        │                                                 │
            High Data Consistency                            High Data Heterogeneity
            (Historical & Trial Align)                       (Conflict Detected)
                        │                                                 │
                        ▼                                                 ▼
          ┌───────────────────────────┐                     ┌───────────────────────────┐
          │   Borrowing Weight (α → 1) │                     │   Borrowing Weight (α → 0) │
          │ Full Historical Discounting│                     │ Down-weighting / Prior    │
          │   Sample Size Reduced     │                     │    Inflation Protection   │
          └───────────────────────────┘                     └───────────────────────────┘

Methods such as Power Priors, Commensurate Priors, and Robust Meta-Analytic Predictive (MAP) Priors dynamically tune the borrowing parameter ($\alpha_0 \in [0, 1]$):

  • If the current trial outcomes align with historical control data, $\alpha_0$ approaches 1, allowing maximum borrowing and reducing the required investigational sample size.
  • If a data conflict emerges (e.g., historical control mortality is 5% but preliminary trial control data shows 12%), the Bayesian model automatically down-weights the historical data ($\alpha_0 \to 0$), protecting the trial against Type I error inflation.

Sponsors interested in the mathematical formulation of power priors and commensurate priors should refer to our detailed post on Bayesian historical-data borrowing for device trials, as well as our framework for calculating the resulting sample size savings in sample size calculation and historical-data borrowing.

Matching-Adjusted Indirect Comparison (MAIC)

When a sponsor possesses individual patient data (IPD) for their proprietary device trial but can only access published aggregate data (e.g., mean baseline vectors and event rates) for a competitor historical control arm, traditional propensity score matching is impossible.

Originally proposed by Signorovitch et al. (2012) and formalized by Phillippo et al. in NICE Decision Support Unit Technical Support Document 18 (TSD 18), MAIC solves this by re-weighting individual device patients so that their weighted baseline covariate means match the aggregate covariate vector reported in the historical publication.

The effective sample size ($ESS$) of the re-weighted device cohort is calculated as:

$$ESS = \frac{\left( \sum_{i=1}^n w_i \right)^2}{\sum_{i=1}^n w_i^2}$$

where $w_i$ represents the individual propensity weight assigned to subject $i$. If the historical population differs drastically from the device trial population, individual weights become highly skewed, reducing the $ESS$ and flagging loss of statistical power.

When an externally controlled study incorporates multiple endpoints or serial statistical comparisons against benchmark targets, sponsors must also implement alpha-control adjustments detailed in our guide on multiplicity and multiple-endpoint testing in device trials.


Recommended Reading
Superiority and Equivalence Clinical Trial Designs for Medical Devices
Clinical Evidence Regulatory2026-08-10 · 27 min read

What does the FDA 2023 draft guidance say, and why is there a device guidance gap?

On February 1, 2023, the FDA published a draft guidance document titled Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products (Federal Register Doc 2023-02094, Docket FDA-2022-D-2983).

The CDER/CBER Scope vs. CDRH Reality

While this guidance represents a major milestone in regulatory science, its title and administrative scope explicitly restrict its application to human drugs, biological products, and oncology therapeutics:

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                        REGULATORY SCOPE OF FDA EXTERNAL CONTROL GUIDANCE               │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                            │
    ┌───────────────────────────────────────┴───────────────────────────────────────┐
    ▼                                                                               ▼
┌───────────────────────────────────────────┐   ┌───────────────────────────────────────────┐
│     FDA 2023 DRAFT GUIDANCE               │   │      FDA CDRH DEVICE FRAMEWORK            │
│  (FR Doc 2023-02094; FDA-2022-D-2983)      │   │     (FR 2013-26690 / ICH E10)            │
├───────────────────────────────────────────┤   ├───────────────────────────────────────────┤
│• Issuing Bodies: CDER, CBER, OCE          │   │• Issuing Body: CDRH (Center for Devices)  │
│• Scope: Drugs & Biological Products       │   │• Core Guidance: Nov 2013 Design           │
│• Device Status: EXCLUDED FROM SCOPE       │   │  Considerations for Pivotal Investigations│
│• Focus: EHR data, causal inference for    │   │• Device Mechanisms: Objective Performance │
│  small-molecule/biologic single-arm trials│   │  Criteria (OPC) & Performance Goals (PG)  │
└───────────────────────────────────────────┘   └───────────────────────────────────────────┘

Because CDRH did not co-author or adopt FR Doc 2023-02094, medical device sponsors operate in a guidance gap. Device sponsors cannot directly cite the 2023 draft guidance as binding CDRH policy. Instead, they must synthesize a compliant regulatory strategy from three distinct sources:

  1. CDRH 2013 Pivotal Investigations Guidance (FR 2013-26690): Establishes the formal definitions of Objective Performance Criteria (Section 7.6.1) and performance goals, serving as the foundational device-native benchmark framework.
  2. ICH E10 Choice of Control Group: Supplies the international standard for control group classification, defining historical/external controls and establishing exchangeability principles.
  3. FDA IDE Regulatory Framework (21 CFR 812): Outlines the submission process for obtaining agency agreement on nonconcurrent control designs prior to trial initiation. Sponsors preparing IDE applications should review our guide on IDE study designs and control types for device investigations.

Global Regulatory Comparison: FDA vs. EMA vs. MHRA

International regulatory authorities exhibit varying degrees of flexibility regarding external controls for medical devices and IVDs:

Regulatory Body Primary Guidance Document External Control Stance for Devices Key Requirements
US FDA (CDRH) 2013 Pivotal Investigations Guidance; 21 CFR 812 / 814. Acceptable when RCT is unfeasible; strong preference for OPC/PG in mature classes. Pre-specification in IDE; exchangeability proof; constancy justification.
EMA (European Union) Clinical evaluation under EU MDR (2017/745) Annex XIV; MDCG 2020-6. Acceptable for CE-mark clinical evaluation if state-of-the-art benchmark is established. Rigorous systematic literature review; demonstration of equivalence to legacy devices.
UK MHRA Draft guidance on real-world data in external control arms. Open to real-world external controls in high unmet need and rare disease. High-quality RWD audit trails; transparent protocol pre-specification.

For rare disease devices seeking regulatory access through specialized US pathways, external control strategies frequently align with requirements covered in our analysis of humanitarian device exemption for rare-disease devices.


Canonical Device Precedent: Prosthetic Heart Valve OPCs

The history of prosthetic heart valve regulation provides the canonical model for how historical clinical data matures into a formally established Objective Performance Criterion.

Historical Origin and FDA Establishment

Prior to the 1990s, every novel cardiac valve prosthesis required a concurrent randomized controlled trial against an approved valve. However, as valve replacement technology matured and complication rates stabilized, requiring RCTs for minor valve design iterations exposed thousands of control patients to redundant clinical trial protocols without scientific necessity.

In 1994, working in conjunction with ISO standards committees, FDA CDRH established standardized Objective Performance Criteria (OPC) for replacement heart valves, derived from pooled analyses of thousands of patient-years of historical trial data.

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                   CANONICAL HISTORICAL TIMELINE OF HEART VALVE OPCS                     │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                            │
    ┌───────────────────────┬───────────────┴───────────────┬───────────────────────┐
    ▼                       ▼                               ▼                       ▼
┌──────────────────┐    ┌──────────────────┐    ┌──────────────────┐    ┌──────────────────┐
│  PRE-1994: RCT   │    │  1994: ISO/FDA   │    │ 2006: CHEN ET AL.│    │2014: GRUNKEMEIER │
│  REQUIREMENT     │    │   OPC LAUNCH     │    │  FDA EVALUATION  │    │ BAYESIAN FRAMEWORK│
├──────────────────┤    ├──────────────────┤    ├──────────────────┤    ├──────────────────┤
│Every valve model │    │Pooled historical │    │FDA reveals 15    │    │Bayesian stopping │
│required 1:1 RCT  │    │data yields fixed │    │valves approved   │    │rules integrate   │
│against legacy    │    │complication rate │    │via single-arm    │    │continuous OPC    │
│valves            │    │benchmarks        │    │OPC trials        │    │monitoring        │
└──────────────────┘    └──────────────────┘    └──────────────────┘    └──────────────────┘
  • Chen et al. (2006): In an FDA perspective published in Annals of Thoracic Surgery, FDA authors documented that between 1994 and 2006, the agency approved 15 cardiac valve prostheses through single-arm clinical investigations comparing device safety rates directly against established OPC benchmarks.
  • Grunkemeier et al. (2006, 2014): Publications in Annals of Thoracic Surgery and The Journal of Thoracic and Cardiovascular Surgery defined linearized complication rates (% per patient-year) across seven primary valve adverse events (thromboembolism, valve thrombosis, major hemorrhage, paravalvular leak, endocarditis, structural valve deterioration, and non-structural dysfunction). Grunkemeier and colleagues further formalized a Bayesian stopping framework for monitoring single-arm trials against OPC thresholds.

The heart valve experience proves that external control benchmarks succeed when the regulatory authority, clinical community, and industry collaborate to standardize outcome definitions and pool historical evidence.


Frequently Asked Questions

Is an Objective Performance Criterion the same thing as an external control arm?

No. An Objective Performance Criterion (OPC) is a single, fixed numerical target rate or threshold (e.g., "1-year major adverse event rate $< 8.0%$") established by FDA guidance or consensus standards for a well-characterized device class. An external control arm is a constructed cohort of individual historical or real-world patients whose data is statistically matched or weighted against investigational trial subjects.

Does the FDA externally-controlled-trials draft guidance apply to medical devices?

No. The FDA draft guidance issued on February 1, 2023 (Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products, FR Doc 2023-02094) was published by CDER, CBER, and the Oncology Center of Excellence. It covers human drugs and biological products only. Medical device sponsors must refer to CDRH's 2013 pivotal investigation guidance (Section 7.6.1), ICH E10, and device-specific PMA precedents.

How common are external and historical controls in device PMA approvals, and how often are they justified?

According to a 2025 cross-sectional study in JAMA Network Open by Mooghali et al., 59.1% (52 of 88) of original high-risk therapeutic device PMA approvals granted by FDA between 2019 and 2023 relied on nonconcurrent controls (39 performance goals, 6 historical controls, 6 OPCs, and 1 mixed design). However, only 3.8% (3 of 79 analyses) provided formal statistical justification for the selected control, demonstrating a major regulatory vulnerability in historical device submissions.

Can a real-world registry serve as the external control for a single-arm device pivotal trial?

Yes, provided the registry data collection is prospective or rigorously audited, eligibility criteria match the trial protocol, key prognostic covariates are completely captured without systematic missingness, and the statistical analysis plan (incorporating propensity matching or weighting) is pre-specified in the IDE prior to inspecting device trial outcomes.

What is a synthetic control arm, and will FDA accept one for a Class III device?

A synthetic control arm (SCA) is a constructed cohort created by applying causal inference algorithms (such as propensity score matching or target trial emulation) to real-world data from electronic health records, claims databases, or historical trial repositories. The FDA evaluates SCAs case-by-case; acceptance is highest in rare disease indications or high unmet medical needs where a concurrent RCT is unfeasible.


Recommended Reading
FDA Pediatric Medical Device Regulation: Pathways, Section 515A, and Waivers
Regulatory Clinical Evidence2026-07-26 · 18 min read

Summary & Action Plan for Device Clinical Leaders

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                   5-STEP ROADMAP FOR EXTERNALLY CONTROLLED DEVICE TRIALS               │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                            │
    ┌───────────────────────┬───────────────┼───────────────┬───────────────────────┐
    ▼                       ▼               ▼               ▼                       ▼
┌──────────────────┐    ┌──────────────┐┌──────────────┐┌──────────────────┐    ┌──────────────────┐
│ STEP 1: AUDIT    │    │STEP 2: CHECK ││STEP 3: LOCK  ││STEP 4: EXECUTE   │    │STEP 5: SUBMIT    │
│ UNFEASIBILITY    │    │DATASET SCOPE ││SAP & MATCHING││EXCHANGEABILITY   │    │PRE-SUBMISSION    │
├──────────────────┤    ├──────────────┤├──────────────┤├──────────────────┤    ├──────────────────┤
│Document ethical  │    │Determine if  ││Pre-specify   ││Run propensity    ││Present DAG, SAP, │
│or practical      │    │patient-level ││control source││overlap, SMD, &   ││and constancy     │
│barriers to       │    │data or OPC   ││& algorithms  ││E-value           ││argument in formal│
│concurrent RCT    │    │exists        ││before study  ││diagnostics       ││Q-Submission      │
└──────────────────┘    └──────────────┘└──────────────┘└──────────────────┘    └──────────────────┘

When designing a pivotal medical device trial that departs from a concurrent randomized control, clinical and regulatory leaders should follow this systematic workflow:

  1. Establish the Unfeasibility Justification: Formally document ethical barriers (such as surgical sham risk), rare disease prevalence figures, or clinical impracticability in the protocol background.
  2. Determine Control Architecture (OPC vs. Constructed Arm): Evaluate whether a CDRH-established OPC or performance goal exists for the device class. If no fixed benchmark exists, verify whether patient-level historical or registry data can be accessed.
  3. Finalize SAP Prior to Outcome Inspection: Pre-specify the external control data source, inclusion/exclusion rules, matching algorithm (propensity scoring, IPTW, or MAIC), and primary statistical hypothesis in the SAP before unblinding device trial data.
  4. Conduct Rigorous Exchangeability Diagnostics: Execute propensity score overlap checks, verify that standardized mean differences across baseline covariates are $< 0.10$, and perform E-value sensitivity analyses to quantify potential unmeasured confounding.
  5. Engage CDRH via Pre-Submission (Q-Submission): Present the proposed external control design, DAG causal framework, and statistical justification to the relevant CDRH review division prior to IDE filing to lock in regulatory alignment.