MedDeviceGuideMedDeviceGuide
Back

IVD Reference Intervals: Established Values vs Local Verification

When a laboratory may adopt an IVD manufacturer's reference interval, when CLIA requires verification or establishment, and what ISO 15189 adds only for accredited laboratories.

Ran Chen
Ran Chen
Global MedTech Expert | 10× MedTech Global Access
Published 2026-09-22Last reviewed 2026-09-2228 min read

When an IVD manufacturer prints biological reference intervals in its package insert or instructions for use, the receiving laboratory has an immediate operating question: May the receiving clinical laboratory adopt those established values directly into its Laboratory Information System (LIS), or does regulatory compliance mandate an empirical local verification study before reporting patient results?

The same practical question arises under US laboratory rules and under ISO 15189 accreditation, but those are separate duties rather than one international rule: established values are the manufacturer's premarket claim, not the clinical laboratory's operating clearance. Under the United States Clinical Laboratory Improvement Amendments of 1988 (CLIA) regulations, a laboratory that introduces an unmodified, FDA-cleared or approved test system must, before reporting patient test results, verify that the manufacturer's reference intervals (normal values) are appropriate for the laboratory's specific patient population (42 CFR 493.1253(b)(1)(ii)). If the laboratory modifies the assay protocol, applies an alternative specimen matrix, introduces an assay not subject to FDA clearance (such as an in-house laboratory-developed test), or adopts an assay lacking manufacturer performance specifications, CLIA requires the laboratory to establish its own reference intervals (42 CFR 493.1253(b)(2)(vi)) and comprehensively document all activities (42 CFR 493.1253(c)).

On the accreditation side, ISO 15189:2022 clause 7.3.5 is an accreditation requirement, not a CLIA equivalent and not self-executing law. Where it applies, the laboratory must define biological reference intervals and clinical decision limits, record their scientific and demographic basis, and ensure they reflect the patient population served while evaluating potential clinical risk. The note to clause 7.3.5 states that biological reference values provided by the manufacturer can be used by the laboratory if the population base is verified and deemed acceptable. The same note is quoted in the accreditation section below. In parallel, device manufacturers have their own labeling duties: the US FDA mandates that IVD package inserts disclose expected value ranges, establishment methodology, and underlying study populations under 21 CFR 809.10(b)(11), while Regulation (EU) 2017/746 (IVDR) includes expected values in normal and affected populations in clinical performance (Annex I, Section 9.1(b)). The instructions for use must include, where relevant, reference intervals in normal and affected populations (Annex I, Section 20.4.1(aa)).

Navigating this operational junction requires a precise understanding of the distinct legal duties assigned to IVD manufacturers versus testing laboratories, the consensus verification protocols codified in Clinical and Laboratory Standards Institute (CLSI) guidelines, the audit criteria enforced by College of American Pathologists (CAP) inspectors, and the significant statistical limitations inherent in the conventional 20-sample verification shortcut.

What a Reference Interval Claim Is — and Who Owns the Evidence

To maintain audit readiness and clinical accuracy, laboratory professionals must distinguish between descriptive physiological benchmarks and prescriptive medical action thresholds. Under guidelines established by the International Federation of Clinical Chemistry and Laboratory Medicine (IFCC) and codified in CLSI EP28-A3c (Defining, Establishing, and Verifying Reference Intervals in the Clinical Laboratory; Approved Guideline — Third Edition, reaffirmed in April 2020), a reference individual is a person selected on the basis of well-defined biological, clinical, and lifestyle criteria who represents a healthy state or defined physiological baseline. A collective group of these individuals constitutes the reference population, from which a representative reference sample group is recruited.

When biological specimens collected from this sample group are analyzed under standardized preanalytical and analytical conditions, the resulting values are termed reference values. From these values, statistical analysis—most commonly non-parametric ranking—determines the reference limits, which typically bound the central 95% of the healthy distribution between the 2.5th and 97.5th percentiles. The numerical span between these limits is the central 95% reference interval. A fundamental mathematical consequence of this definition is that about 5% of people in the reference population fall outside the limits, roughly 2.5% below the lower limit and 2.5% above the upper limit. That share is part of the definition of a central 95% interval. It is not, by itself, evidence that the interval is wrong.

A critical distinction must be drawn between descriptive reference intervals and prescriptive clinical decision limits (CDLs). While a reference interval reflects the empirical distribution of an analyte in a healthy cohort, a clinical decision limit is an evidence-based threshold derived from clinical outcome studies, randomized trials, or clinical consensus panels that indicates high risk of specific morbidity or triggers therapeutic intervention. Prominent examples include:

  • Lipid Panels: The EFLM 2016 recommendation names total cholesterol as an analyte for which a consensus decision limit replaces a reference interval, so the laboratory does not need to determine or verify a central-95% reference interval for that claim. The numeric cut-offs belong to the lipid guideline the laboratory actually uses. This article does not treat any one cholesterol panel as current law.

  • Glycated Hemoglobin (HbA1c): Glycated hemoglobin is the other EFLM 2016 example. Diagnostic cut-offs come from the clinical guideline in force. ADA and WHO both publish a diabetes threshold, but they do not share one prediabetes range, so a laboratory should not attribute a single prediabetes band to both.

  • High-Sensitivity Cardiac Troponins: A 99th-percentile upper reference limit, including the limit used with high-sensitivity cardiac troponin, is still a reference limit. A consensus document may adopt that limit as a decision threshold. That use does not put troponin in the cholesterol or HbA1c category, and it does not exempt the manufacturer or the laboratory from establishing or verifying the limit.

For the EFLM examples, total cholesterol and glycated hemoglobin, the laboratory need not determine or verify a conventional central-95% reference interval. The useful check is analytical accuracy and precision near the decision limit, including the limit of quantitation where that limit matters, as discussed in our technical guide to IVD analytical performance validation (LoD, LoQ, and precision).

What the IVD Manufacturer Must Establish Before a Lab Ever Sees the Interval

Before an IVD reagent kit or automated analytical system reaches commercial distribution, the manufacturer must generate comprehensive empirical evidence establishing expected values as part of its premarket technical dossier.

In the United States, commercial IVD labeling is governed by 21 CFR 809.10. Specifically, 21 CFR 809.10(b)(11) mandates that package inserts for in vitro diagnostic products incorporate a dedicated section titled "Expected values". This regulation requires the manufacturer to:

  • State the range(s) of expected values obtained with the product from studies of various populations;

  • Indicate explicitly how the range was established (e.g., non-parametric ranking, parametric modeling, or consensus threshold); and

  • Clearly identify the population(s) on which the range was established, including demographic parameters, clinical specimen types, and geographic sourcing.

FDA teaching slides describe, rather than add a second legal duty: the expected-values section may carry different reference ranges for different populations, for example men and women or pediatric individuals, and for a qualitative test system it may describe the cut-off and how people distribute around it. The binding sentence remains 21 CFR 809.10(b)(11), which requires the ranges, how they were established, and the populations studied. Furthermore, 21 CFR 809.10(b)(12) obliges manufacturers to state specific performance characteristics—including accuracy, analytical precision, analytical specificity, and sensitivity relative to generally accepted reference methods across both normal and abnormal specimens.

To fulfill these premarket labeling mandates, manufacturers routinely declare conformity to FDA-recognized consensus standards. Under Recognition Number 7-224 (entered into Recognition List 034 on January 30, 2014, and maintained in current CDRH consensus listings), the FDA formally recognizes CLSI EP28-A3c as an acceptable protocol for establishing and verifying reference intervals in clinical laboratory submissions. A manufacturer may declare conformity to EP28-A3c in a premarket submission. Recognition lets FDA use that declaration for the elements of the submission the standard covers. It does not itself prove that a particular donor study was adequate, and it does not make the guideline a legal requirement.

Under European Union law, the EU In Vitro Diagnostic Medical Devices Regulation (Regulation (EU) 2017/746, IVDR) establishes strict premarket requirements in Annex I (General Safety and Performance Requirements). In Annex I, Section 9.1(b) legally defines device clinical performance to include "expected values in normal and affected populations". Furthermore, Annex I, Section 20.4.1, item (aa) mandates that instructions for use include, where relevant, "reference intervals in normal and affected populations". Those items feed the performance evaluation under Annex XIII, including the performance evaluation report, which notified bodies assess. There is no separate legally named 'Clinical Performance Report' paired with that report. The IVDR performance-evaluation structure is covered in our IVDR performance evaluation guide and foundational overview of IVD devices regulatory compliance.

CLIA's Two Lists: When a US Lab Verifies and When It Must Establish

In the United States, clinical laboratory testing operations are governed by federal regulations enforced by the Centers for Medicare & Medicaid Services (CMS) under CLIA. The central regulatory anchor governing test introduction is codified at 42 CFR 493.1253 (Standard: Establishment and verification of performance specifications). This regulation, not the CLIA statute itself, sets two lists based on whether the test system is unmodified commercial technology or a customized assay.

The Verification Track for Unmodified FDA Cleared/Approved Systems (42 CFR 493.1253(b)(1))

Paragraph (a) exempts any test system the laboratory was already using before 24 April 2003. For a later introduction of an unmodified, FDA-cleared or approved test system, the regulation limits the laboratory's initial validation burden to verification. Under 42 CFR 493.1253(b)(1), the laboratory must, before reporting patient test results, demonstrate that it can achieve performance specifications comparable to those established by the manufacturer for accuracy, precision, and reportable range.

For reference intervals specifically, 42 CFR 493.1253(b)(1)(ii) establishes the binding duty:

  • "Verify that the manufacturer's reference intervals (normal values) are appropriate for the laboratory's patient population."

Under this verification track, the laboratory is not legally obligated to recruit hundreds of healthy subjects to construct a distribution de novo. Instead, it must execute a structured, documented evaluation showing that the manufacturer's reference intervals are appropriate for the laboratory's patient population. A completed 20-sample file can meet that documentation duty without proving the limits are optimal.

The Establishment Track for Modified Assays and LDTs (42 CFR 493.1253(b)(2))

The second statutory list applies whenever a testing facility departs from the manufacturer's validated product claims. Under 42 CFR 493.1253(b)(2), a laboratory must establish its own performance specifications prior to reporting patient test results if the laboratory:

  • Modifies an FDA-cleared or approved test system;

  • Introduces a test system that is not subject to FDA clearance or approval (including laboratory-developed tests, in-house developed assays, and research-use-only reagents adapted for diagnostic use); or

  • Employs a test system in which manufacturer performance specifications are not provided.

Under this establishment track, the laboratory assumes the full evidentiary and scientific burdens of an in vitro diagnostic developer. Specifically, 42 CFR 493.1253(b)(2)(vi) compels the laboratory to establish its own "reference intervals (normal values)" together with the other characteristics in paragraph (b)(2): accuracy, precision, analytical sensitivity, analytical specificity including interfering substances, reportable range, and any other characteristic required for test performance. Calibration and control procedures are a separate duty in paragraph (b)(3), and calibration verification is in 42 CFR 493.1255. For the boundary between a commercial test system and a laboratory modification, see our guide to LDT regulatory compliance and CLIA oversight as well as our analysis of design verification vs validation for medical devices.

The regulation does not publish an exhaustive modification list. CAP's checklist gives a change in specimen type or collection device as a common example that takes a test off the unmodified FDA-cleared or approved track. Other departures from the instructions for use are assessed on the facts. Examples that laboratories commonly have to justify include:

  • Matrix and Specimen Alterations: Testing plasma, urine, or cerebrospinal fluid when the assay is cleared exclusively for serum;

  • Operating Parameter Changes: Altering incubation temperatures, modifying reagent-to-sample dilution ratios, or adjusting photometric read times;

  • Platform Adaptations: Applying an open-channel reagent kit to an automated analyzer platform not validated by the reagent manufacturer; or

  • Specimen Collection Tubes: Utilizing collection tubes with alternate anticoagulants or separator gel barriers not specified in the package insert.

Finally, 42 CFR 493.1253(c) establishes the overarching recordkeeping mandate: "The laboratory must document all activities specified in this section." The 24 April 2003 carve-out is in paragraph (a), and it covers any test system the laboratory was already using before that date, not only unmodified FDA-cleared systems and not as a clause inside paragraph (c). Every verification protocol, donor intake log, raw instrument run, outlier evaluation, and laboratory director approval must be permanently archived and available for inspection by CMS, state survey agencies, or deemed accreditation bodies.

The Verification Study: 20 Local Samples and the Acceptance Rule

While CLIA regulations establish the legal requirement to verify reference intervals, the regulation does not prescribe the mathematical study design. In clinical practice, accredited laboratories implement consensus protocols codified in CLSI EP28-A3c and its accompanying implementation guide CLSI EP28IG (Verification of Reference Intervals in the Medical Laboratory Implementation Guide, First Edition, published April 2022).

The 20-Sample Reference Verification Protocol

CLSI EP28-A3c outlines a standardized protocol for verifying a pre-established or transferred reference interval. Rather than recruiting a large population cohort, the receiving laboratory collects and analyzes biological specimens from a minimum of 20 apparently healthy reference individuals representative of the laboratory's local patient population for each relevant demographic partition (e.g., 20 adult males and 20 adult females).

Donor selection has to follow the exclusion and partition criteria the protocol actually adopts. The items below are examples used in reference-interval studies. They are not a CLIA checklist and they are not quoted here as mandatory CLSI text:

  • Exclusion Criteria: Examples commonly considered include active disease, pregnancy, recent surgery or blood donation, and medication expected to move the analyte. A 14-day illness window is an illustration, not a number fixed by CLIA or by the EFLM recommendation.

  • Preanalytical Controls: The receiving method's preanalytical conditions should be consistent with the study being transferred: fasting when the original interval assumed it, posture, tube type, and separation timing. An 8–12 hour fast is a common metabolic-profile practice, not a universal legal minimum.

  • Outlier Screening: Screen the 20 results for statistical outliers with the Reed/Dixon or Tukey methods, then replace removed results so the set remains 20, as Hoffmann et al. 2026 summarize the CLSI procedure. Reed's published rule deletes an extreme value when D/R is greater than one-third, not at or above one-third. Those outlier checks are separate from the later count of results that fall outside the candidate limits.

The Binomial Acceptance and Escalation Rule

Once 20 valid reference specimen results are compiled for a partition, the laboratory evaluates them against the manufacturer's published candidate interval [L, U] using the established binomial decision rule:

  • Acceptance (2 or fewer results outside): If no more than 2 of the 20 results fall outside the candidate limits, the EFLM 2016 summary and the CAP checklist example treat the interval as verified for the population studied. That is a binomial acceptance rule. It is not statistical confirmation that the interval is correct. The 2026 review below is the reason not to describe a pass as proof.

  • Second set (EFLM: exactly 3 outside): The EFLM 2016 summary says to test a new set of 20 when exactly 3 of the first 20 fall outside, and to accept the interval if 2 or fewer of that second set fall outside. Hoffmann et al. 2026 describe the CLSI procedure differently: a second set of 20 when 3 or 4 results fall outside. This article follows the EFLM counts and attributes them to that summary. It does not merge the two readings into a new 40-result rule.

  • Do not adopt the candidate interval (EFLM: 4 or more outside, or a failed second set): Under the EFLM 2016 summary, 4 or more results outside the first set, or more than 2 outside the second set, means the laboratory should review the analytical procedure, consider biological and demographic differences, and determine limits with the establishment protocol. The laboratory should not report against the failed candidate interval. Hoffmann et al. 2026 instead send the 3-or-4 case to a second set before that conclusion.

When verification fails, the laboratory director must initiate a structured root-cause investigation:

  1. Review Preanalytical Factors: Examine specimen transport temperature, tourniquet duration, tube barrier gel integrity, and hemolysis/icterus/lipemia (HIL) interference indices.

  2. Evaluate Analytical Traceability and Calibration: Inspect quality control drift, confirm instrument calibration status, verify calibrator metrological traceability under ISO 17511 (see our guide to calibrator and control traceability under ISO 17511), and audit equipment maintenance records as detailed in our guide to equipment calibration management under ISO 13485 clause 7.6.

  3. Examine Demographic Discrepancies: Assess whether local population characteristics (such as benign ethnic neutropenia, localized dietary factors, or altitude) diverge systematically from the manufacturer's study cohort.

  4. Transition to De Novo Establishment: If analytical and preanalytical variables are verified as accurate, the laboratory must establish a custom reference interval de novo.

Transference Without New Subjects: When Adoption Is Allowed

Under specific circumstances, can a clinical laboratory adopt an established reference interval from a manufacturer, peer-reviewed literature, or a multicenter study without collecting any new prospective subject specimens? Both the EFLM and CLSI EP28-A3c Chapter 10 recognize two no-recruitment transference pathways:

1. Subjective Transference (Literature and IFU Transference)

Under subjective transference, the laboratory director conducts an exhaustive administrative and scientific review of the originating reference study. EFLM treats subjective transfer as available only when the original study is consistent with the receiving laboratory. That is a professional-body route, not a CLIA exemption. A US laboratory that reports an unmodified FDA-cleared or approved test still needs a recorded basis that the manufacturer's interval is appropriate for its population (42 CFR 493.1253(b)(1)(ii)). CAP will accept a documented evaluation of the manufacturer's interval when a formal study is not practical. The five EFLM elements to document are:

  • Geographic and Demographic Alignment: The demographic profile of the original reference population (age range, sex distribution, ethnicity, lifestyle factors) is consistent with the patient community the laboratory serves.

  • Preanalytical Consistency: Phlebotomy techniques, specimen collection tube types, anticoagulant formulations, handling protocols, centrifugation parameters, and specimen storage temperatures are consistent with the original study. EFLM asks for consistency, not a claim that every step was identical.

  • Analytical Method Comparability: The receiving laboratory uses analytical performance consistent with the original study, including reagent and detection technology close enough that a bias adjustment is unnecessary, and calibrator traceability the laboratory can actually document.

  • Transparent Study Documentation: The manufacturer or published source must fully disclose its inclusion and exclusion criteria, outlier rejection rules, and statistical sample size.

  • Statistical Robustness: The original reference limits must have been computed using sound statistical methods (e.g., non-parametric ranking on >=120 subjects per partition).

The operational vulnerability of subjective transference is significant: medical device manufacturers rarely publish the complete, granular demographic datasets and raw preanalytical data required to satisfy all five criteria. Consequently, relying on subjective transference alone exposes laboratories to audit citations unless supported by empirical verification data.

2. Method-Comparison Transference

When a laboratory transitions an established assay from an older analyzer platform to a new instrument system (or replaces an existing assay reagent lot), collecting 120 new healthy donor specimens is clinically impractical. Under CLSI EP28-A3c and EP09c guidelines, the laboratory can transfer existing, clinically validated reference intervals to the new system using a method-comparison study.

The laboratory tests a panel of patient split specimens (spanning the full clinical measurement range) across both the existing reference method (x) and the candidate new method (y). Using Deming or Passing-Bablok regression, the laboratory models the relationship:

y = a + b * x

EFLM's comparison route is narrower than a blanket regression. If the slope is near 1, the intercept is small relative to the pre-defined allowable limit, and the measuring ranges match, the existing limits can transfer. If there is a proportional difference and the intercept is still small, the new limits can be recalculated from the regression after residual analysis. A large intercept is a reason not to transfer. Ordinary linear regression is not always the right model, which is why Deming or Passing-Bablok is used only when that model fits. The arithmetic is:

L_new = a + b * L_orig and U_new = a + b * U_orig

This method-comparison approach provides mathematically sound transference without requiring new healthy donor recruitment, provided analytical precision on both platforms is well characterized.

3. Multicenter Common Reference Intervals

In recent years, large professional consortia—such as the Canadian Laboratory Initiative on Pediatric Reference Intervals (CALIPER) and the Nordic Reference Interval Project (NORIP)—have established standardized common reference intervals across extensive populations. However, international accreditation standards stipulate that a laboratory cannot blindly adopt a multicenter common interval: each receiving laboratory must verify that its local instrument calibration, reagent lot performance, and preanalytical processes align with the consortium's baseline, typically by running a local 20-sample verification check.

Accreditation Checkpoints: ISO 15189:2022 and CAP Expectations

Accrediting bodies translate statutory requirements and consensus guidelines into enforceable audit checklists. In laboratory medicine, the two most influential auditing frameworks are ISO 15189:2022 and the College of American Pathologists (CAP) Laboratory Accreditation Program.

ISO 15189:2022 Clause 7.3.5 Audit Requirements

In the fourth edition of ISO 15189 (published in 2022), reference intervals are governed under clause 7.3.5 (Biological reference intervals and clinical decision limits). Lead technical assessors inspect for conformity against six discrete obligations:

  • Clear Definition & Documentation: Biological reference intervals and clinical decision limits, when needed to interpret results, are defined and communicated to users, and their basis is recorded so the intervals reflect the patient population served while considering risk to patients. Clause 7.3.5 does not require that basis to sit in a document titled a quality manual.

  • Patient Population Alignment: Intervals must reflect the specific patient population served by the laboratory, taking into account demographic composition and associated patient clinical risk.

  • Verification of Manufacturer Values: The explicit Note in clause 7.3.5 affirms: "Biological reference values provided by the manufacturer can be used by the laboratory if the population base of these values is verified and deemed acceptable by the laboratory."

  • Periodic Systematic Review: The laboratory must establish and maintain a scheduled periodic review procedure to ensure reference intervals remain valid over time.

  • Change Notification: When intervals or decision limits change, the laboratory communicates the changes to users. Clause 7.3.5 does not prescribe a particular letter format.

  • Impact Assessment for Method Changes: Under clause 7.3.5 and clause 7.3.7.4 (comparability of examination results), whenever an examination procedure, reagent formulation, or analyzer platform is altered, the laboratory must systematically evaluate and document the impact on biological reference intervals.

CAP Laboratory Accreditation Program (LAP) Checklist Requirements

For CAP-accredited laboratories, inspectors audit compliance against the All Common Checklist (specifically checklist requirements within the COM.50000 / COM.50100 family). Key inspection checkpoints include:

  • Verification / Establishment per Analyte and Matrix: The laboratory must verify or establish reference intervals for each analyte and specimen matrix (e.g., serum, lithium-heparin plasma, urine) prior to clinical reporting.

  • Acceptable Evidence of Compliance: The checklist note gives one verification example: samples from 20 healthy representative individuals, considered verified if no more than two results fall outside the proposed interval (CLSI EP28-A3c). Evidence of compliance is a record of the reference-interval study, or records of verification of the manufacturer's stated interval when a formal study is not practical, or another method approved by the laboratory or section director. Director approval is the third of those evidence options, not a separate certificate layered onto every 20-sample file. The training excerpt uses the COM.50000 and COM.50100 numbers; confirm them against the checklist edition in force.

  • Mandatory Re-Evaluation Triggers: CAP inspectors verify that reference intervals are systematically re-evaluated whenever: (1) a new analyte is introduced; (2) analytical methodology, instrument platform, or reagent chemistry changes; or (3) a meaningful demographic shift occurs in the served patient population.

The Configuration-to-Claim-to-Evidence Map

To provide an auditable operational blueprint for regulatory affairs, quality assurance, and laboratory management, the following decision tree and configuration matrix map every common laboratory testing scenario to its governing legal authority, permissible reporting claim, required verification route, and minimum sample size.

flowchart TD
    A["Test system the laboratory will report"] --> B{"Unmodified FDA-cleared or approved system, and not exempt under 493.1253(a)?"}
    B -- "No: modified, non-FDA, or no manufacturer specifications" --> C["Establish reference intervals under 42 CFR 493.1253(b)(2)"]
    B -- "Yes" --> E{"EFLM decision-limit analyte, such as cholesterol or HbA1c?"}
    E -- "Yes" --> F["Check accuracy and precision at the decision limit. Do not substitute a central 95 percent interval"]
    E -- "No" --> G{"All five EFLM subjective-transfer conditions documented as consistent?"}
    G -- "Yes. This is not a CLIA exemption" --> H["Keep the population-base record. US reporting still needs a 493.1253(b)(1)(ii) basis"]
    G -- "No or incomplete" --> I["Binomial verification summarized from CLSI EP28-A3c"]
    I --> J["20 apparently healthy local reference individuals per partition, after outlier replacement"]
    J --> K{"Results outside the candidate limits, EFLM 2016 counts"}
    K -- "2 or fewer" --> L["EFLM and CAP example: interval accepted for that population"]
    K -- "Exactly 3" --> M["Second set of 20"]
    M --> N{"Second set: 2 or fewer outside?"}
    N -- "Yes" --> L
    N -- "No" --> C
    K -- "4 or more" --> C
    C --> O["CLSI floor: at least 120 reference individuals per partition"]
    O --> P["Nonparametric central 95 percent interval"]
    P --> Q["Document activities under 42 CFR 493.1253(c)"]
    L --> Q
    F --> Q
    H --> Q
EFLM 2016 decision counts: a second set of 20 only when exactly 3 results fall outside. Hoffmann et al. 2026 describe a second set when 3 or 4 fall outside. Subjective transfer is not a CLIA exemption.
Operational ConfigurationRegulatory & Standard BasisPermissible Reporting ClaimRequired Evidence RouteGuideline sample-size conventionPrimary Binding Citation
Unmodified FDA-cleared or approved commercial IVD assay42 CFR 493.1253(b)(1) for a nonwaived unmodified FDA-cleared or approved system; ISO 15189:2022 only where accreditedManufacturer-established reference interval adopted locallyEmpirical verification of manufacturer's stated intervalCLSI/EFLM convention: 20 reference individuals per partition. Not a CLIA minimum42 CFR 493.1253(b)(1)(ii); ISO 15189:2022 7.3.5
Commercial IVD assay with modified specimen matrix or protocol42 CFR 493.1253(b)(2) establishment duty. Modified FDA-cleared tests are commonly treated as high complexity, which is a categorization point separate from the reference-interval paragraphLaboratory-established custom reference intervalFull de novo establishment study (non-parametric percentiles)CLSI EP28-A3c floor: at least 120 reference individuals per partition. Not a number in CLIA42 CFR 493.1253(b)(2)(vi); 42 CFR 493.1253(c)
Laboratory-Developed Test (in-house developed assay)42 CFR 493.1253(b)(2) for a US non-FDA or in-house method. IVDR Article 5(5) is a separate EU in-house regime and does not set the 120-subject floorLaboratory-established proprietary reference intervalFull de novo establishment with documented inclusion/exclusionCLSI EP28-A3c floor: at least 120 reference individuals per partition. Not a number in CLIA42 CFR 493.1253(b)(2)(vi); CLSI EP28-A3c
Analyzer platform replacement (same analyte and reagent chemistry)Method-comparison transference under CLSI EP28 Chapter 10; ISO 15189:2022 clause 7.3.7.4 where comparability differences must be checked against intervalsTransferred existing reference interval (adjusted if biased)Method comparison regression (Deming / Passing-Bablok) on split patient samplesSplit patient samples sized to the laboratory's method-comparison protocol. CLIA does not set that countCLSI EP28-A3c Chapter 10; ISO 15189:2022 7.3.7.4
Demographic or clinical shift in served patient populationCAP COM.50100 change-in-population trigger. If the laboratory still reports a manufacturer interval, 42 CFR 493.1253(b)(1)(ii) still asks whether that interval fits the populationRe-verified or re-established population intervalTargeted demographic verification or indirect LIS data audit20 reference individuals per affected partition if the binomial check is repeated. The 400-subject figure in Hoffmann et al. 2026 is an establishment cohort size, not a routine-data minimum42 CFR 493.1253(b)(1)(ii); CAP COM.50100
Multicenter harmonized reference interval adoption (e.g. CALIPER)EFLM multicenter prerequisites, plus local verification. ISO 15189:2022 clause 7.3.5 where accreditedConsortium harmonized reference intervalLocal verification check ensuring analytical alignment with consortiumEFLM: each laboratory validates the common interval in its own environment, commonly with the 20-person checkCLSI EP28-A3c; ISO 15189:2022 7.3.5
Analyte governed by clinical consensus decision limits (e.g., HbA1c)Clinical Practice Guidelines (ADA, NCEP, ACC/AHA)Consensus decision limit. Not a locally verified central-95% reference intervalAnalytical accuracy and precision verification near the decision cut-offAccuracy and precision near the decision limit, using the laboratory's ordinary analytical-validation protocolEFLM 2016 for the laboratory. 21 CFR 809.10(b)(11) still requires the manufacturer to state expected values, which may be decision limits rather than a central-95% interval

When Verification Fails — and the Statistical Case Against the 20-Sample Shortcut

While the CLSI EP28-A3c 20-sample protocol represents the accepted baseline for accreditation inspections, peer-reviewed literature in laboratory medicine has demonstrated that this conventional shortcut suffers from severe statistical flaws.

The Statistical Vulnerabilities of the 20-Sample Rule

A comprehensive critical review published in Critical Reviews in Clinical Laboratory Sciences (Hoffmann et al., 2026) evaluates the empirical and mathematical reliability of the 20-sample verification test:

  • High False-Rejection Rate (Type I / Alpha Error): Under a central 95% reference interval, the probability that a randomly drawn healthy individual exceeds the limits is p = 0.05. According to the binomial cumulative distribution, the probability of 3 or more of 20 results falling outside the limits, when the interval truly fits, is about 7.5% (the binomial probability is 0.075). Hoffmann et al. 2026 use that 7.5% false-rejection figure. Consequently, approximately 1 out of every 13 valid, appropriate reference intervals will fail verification purely due to random sampling variation.

  • Virtually Zero Power Against Overly Wide Intervals: The critical safety hazard in clinical diagnostics is adopting an overly wide reference interval, which can mask subtle pathological elevations. Statistical simulation studies demonstrate that at n = 20, the test has essentially zero statistical power to detect too-wide intervals. In the Beck simulation summarized by Hoffmann et al. 2026, 100% of the too-wide intervals were accepted. In that simulation the too-wide case was a central 99.8% interval, and 29% of the too-narrow intervals were also accepted by the 2-of-20 rule.

  • The Double-Sampling Fallacy: Collecting a second set of 20 samples when 3 outliers appear reduces the false-rejection rate below 1%, but mathematical analyses (such as Goossens et al., summarized in the 2026 review) show that double-sampling drastically inflates false-acceptance rates (Type II error), increasing the likelihood of approving inaccurate intervals.

  • True Sample Size Requirements: Hoffmann et al. 2026 separate the errors: beta is the false-acceptance rate, and power is 1 minus beta. They report that n = 73 still reaches only about 80% power when an interval is roughly 50% too wide, and that realistic detection calls for sample sizes approaching n = 100 reference individuals. That is a statistical argument in the review, not a CLIA minimum.

Documented Real-World Clinical Failures

These statistical limitations have concrete real-world consequences documented across multicenter laboratory investigations:

  • Serum Free Light Chain Misclassification: Hoffmann et al. 2026 cite Cotten et al. for a narrower point: manufacturer reference intervals for serum free immunoglobulin light chains failed transference on three of four instrument platforms and misclassified patient results. This article does not add a disease-specific rate the review does not state.

  • Widespread Local Discrepancies: Hoffmann et al. 2026 cite Plagov et al. for the finding that 88% of intervals differed from manufacturer values when tested with local samples, and that the 20-sample approach accepted inaccurate intervals in most cases.

  • Extensive Replacement in Routine Practice: Hoffmann et al. 2026 also cite a large Brazilian laboratory in which manufacturer or literature-derived intervals could not be verified, and had to be replaced, for 8 of 11 analytes studied.

The Emerging Indirect Verification Alternative

Because recruiting 120 reference individuals per partition is often impractical, and because the 20-sample check is statistically weak, laboratory medicine has developed indirect verification methods powered by routine clinical laboratory data. Rather than recruiting healthy volunteers, indirect methods leverage thousands of anonymized test results stored in Laboratory Information Systems (LIS) and electronic health records (EHR).

Open-source algorithms such as reflimR, refineR, and the VeRIf workflow apply advanced statistical modeling—including Box-Cox power transformations, truncated maximum likelihood estimation, and non-parametric Gaussian mixture modeling—to extract the latent non-pathological distribution from mixed outpatient datasets. Hoffmann et al. 2026 describe two quantitative acceptance approaches. Neither is a CLSI or IFCC rule:

  • Equivalence Limits: reflimR uses an equivalence limit based on permissible uncertainty. The 2026 review says a scaling factor of about 1.3 was chosen because it looked plausible. That is not a total-allowable-error limit set by CLIA, ISO, or CLSI.

  • Uncertainty Margins (VeRUS): refineR's VeRUS margins are described as theoretical 90% confidence intervals of the 2.5th and 97.5th percentiles for an n = 120 nonparametric sample. The review warns that overlapping confidence intervals are not as strong as a formal equivalence test, and that the two criteria diverge for broad, right-skewed analytes such as bilirubin.

Conclusion and Actionable Audit Checklist

In in vitro diagnostics, biological reference intervals bridge analytical measurements and clinical decision-making. Device manufacturers bear the legal obligation to rigorously establish expected values and transparently document study demographics in product labeling under 21 CFR 809.10(b)(11) and EU IVDR Annex I Section 20.4.1(aa). Clinical laboratories bear the reciprocal obligation under CLIA 42 CFR 493.1253 and ISO 15189:2022 to independently verify that these established claims hold true for their specific patient population.

For an inspection, the file should show the basis on which the laboratory adopted, verified, or established the interval. A practical file includes:

  • Manufacturer Reference Documentation: The complete manufacturer package insert or IFU carrying the stated expected values and study population description.

  • Verification Protocol: A protocol that states the inclusion and exclusion criteria actually used, the preanalytical conditions, and the acceptance rule being applied, including whether the laboratory is following the EFLM count (2 or fewer outside) or the Hoffmann 2026 reading of a second set at 3 or 4 outside.

  • Raw Testing Records: Instrument printouts or LIS audit trails recording raw values, instrument serial numbers, calibrator lot numbers, and quality control status for all 20 reference specimens per partition.

  • Outlier & Statistical Calculations: The Reed/Dixon or Tukey outlier screen, kept separate from the count of results outside the candidate limits.

  • Escalation Records: If initial verification was inconclusive, complete documentation of the second 20-specimen run or the root-cause investigation leading to de novo establishment.

  • Laboratory Director Formal Approval: Review and approval recorded in the laboratory's validation system before patient results are reported. CAP describes a signed approval statement for the verification data, or director approval when the laboratory uses a method other than a study or a manufacturer-interval verification. CLIA does not require a document titled a certificate.

  • Scheduled Periodic Review: A periodic review, which ISO 15189:2022 clause 7.3.5 requires without setting an annual clock, plus the CAP COM.50100 triggers: a new analyte, a change of analytic methodology, and a change in patient population.

By maintaining clear boundaries between manufacturer establishment duties and local laboratory verification mandates, diagnostic developers and clinical laboratories protect both regulatory compliance and patient safety.