XGene CMC IntelligenceXGene Intelligence

Biosimilar CMC Strategy — 351(k) Pathway and Analytical Similarity Tiers

SpecificationsAnalytical MethodsBiologics

"The analytical similarity assessment does not include a finger-printing approach to structural characterization — the submitted data compare only the attributes in the proposed biosimilar specification against the reference product…

By Khaled Aamer, PhD · Founder, XGene LLC Aug 22, 2026 10 min read
On this pageArticle overview

    “The analytical similarity assessment does not include a finger-printing approach to structural characterization — the submitted data compare only the attributes in the proposed biosimilar specification against the reference product specification limits, rather than demonstrating structural similarity attribute by attribute relative to the reference product batch distribution.” This deficiency captures the fundamental misunderstanding in many first-generation biosimilar submissions: analytical similarity is not about passing a specification — it is about being structurally similar.

    That sentence, drawn from the pattern of FDA Complete Response Letters and deficiency communications over nearly a decade of 351(k) experience, represents a conceptual failure that no amount of additional analytical data can retroactively fix once the study design is already locked. The study design itself embeds the misunderstanding. Comparing a proposed biosimilar’s characterization results to the reference product’s release specification limits is the wrong comparator. Specification limits are set to ensure minimum acceptable quality — they are not a description of what the reference product actually looks like lot to lot, batch to batch, across the years of its commercial history. The natural variability of the reference product is the comparator. If you do not characterize that variability first, in sufficient depth, before designing your similarity study, you are building an evidentiary package on a foundation that FDA will reject at the first substantive review.

    Section 351(k) of the Public Health Service Act (PHS Act) — added by the Biologics Price Competition and Innovation Act of 2009 (BPCIA) and codified at 42 U.S.C. § 262(k) — establishes the statutory pathway for biosimilar licensure. It is worth being precise about that citation, because the pathway is routinely and incorrectly attributed to the Federal Food, Drug, and Cosmetic Act. Biosimilars are biologics, licensed under the PHS Act’s Section 351 framework, not drugs approved under the FD&C Act’s NDA or ANDA provisions, and a CMC team that blurs that distinction tends to import small-molecule bioequivalence logic into a program that runs on a fundamentally different statutory and scientific standard. Section 351(k) does not simply ask whether the proposed product is “similar enough” to the reference product in a generic, commonsense way. It asks whether the proposed product is highly similar to the reference product notwithstanding minor differences in clinically inactive components, and whether there are no clinically meaningful differences between the two in terms of safety, purity, and potency. Those are precise, demanding, operationally specific standards. The phrase “highly similar” is not aspirational language. It is a regulatory conclusion that must be supported by a totality-of-evidence package in which analytical similarity data occupy the central role.

    The guidance architecture underneath that evidence package has shifted substantially since the program’s early years, and a CMC team building a 351(k) submission today needs to be working from the current stack, not the documents that trained most of the field. FDA’s original 2015 guidance on Quality Considerations in Demonstrating Biosimilarity of a Therapeutic Protein Product to a Reference Product (April 30, 2015) established the three-tier analytical similarity framework, and the companion 2015 guidance on Scientific Considerations in Demonstrating Biosimilarity to a Reference Product (April 2015) supplied the totality-of-evidence standard that still governs 351(k) review. Neither document is current guidance any longer. FDA finalized the long-pending 2019 draft guidance, Development of Therapeutic Protein Biosimilars: Comparative Analytical Assessment and Other Quality-Related Considerations, in September 2025; the final version formally replaces the 2015 Quality Considerations guidance, carries the Tier 1/2/3 architecture forward in substance, and refines the expectations around state-of-science analytical methods for higher-order structure characterization, the design of clinical pharmacology studies, and the role a robust analytical similarity package can play in reducing or eliminating the need for clinical efficacy studies. FDA went further in March 2026, withdrawing the 2015 Scientific Considerations guidance outright because, in the agency’s own language, it “no longer represents the FDA’s current thinking” after a decade of biosimilar review experience spanning more than eighty approved products. In its place, FDA has been rebuilding the totality-of-evidence framework through an October 2025 draft guidance narrowing the circumstances in which a comparative efficacy study is needed, and a March 2026 revision to its biosimilar development Q&A guidance that relaxes certain comparative pharmacokinetic study requirements. These guidances — current and superseded alike — are not aspirational documents; they describe what reviewers in CDER’s Office of Pharmaceutical Quality and the relevant therapeutic review division will look for when they open your Module 3.

    The three-tier framework — introduced in the original 2015 Quality Considerations guidance and carried forward, with refinements, into FDA’s September 2025 final guidance on the development of therapeutic protein biosimilars — has become the organizing principle of every analytically credible biosimilar submission. Tier 1 encompasses those quality attributes most sensitive to clinical outcome — functional attributes, binding activity, cell-based potency, Fc effector function where relevant — and requires pre-specified equivalence testing using a 90% confidence interval approach with an equivalence margin of 0.80 to 1.25. Tier 2 covers physicochemical and purity attributes that are important but less directly linked to clinical outcome, and applies a quality range approach in which the boundaries of the quality range are defined by the reference product lot distribution, not by specification limits. Tier 3 encompasses general properties less likely to affect clinical outcome and uses descriptive comparison. The hierarchy is not arbitrary. It is a risk-based architecture in which the stringency of the statistical acceptance criteria scales with the clinical relevance of the attribute.

    What makes this framework technically demanding is the word “pre-specified.” The equivalence margins for Tier 1, the quality range boundaries for Tier 2, and the attribute assignment to each tier must all be established before the first biosimilar lot is tested for similarity purposes. They must be derived from the reference product characterization dataset. If you reverse that order — if you characterize your biosimilar first and then characterize the reference product and design your acceptance criteria — you have introduced the possibility of retrospective rationalization into your study design, and FDA reviewers will recognize the sequence of events. Pre-specification is not a procedural formality. It is what makes the analytical similarity exercise scientifically credible.

    The state-of-science characterization methods expected in a modern biosimilar submission go well beyond what was standard in the first wave of 351(k) applications filed between 2015 and 2018. Intact mass spectrometry and subunit mass analysis establish the molecular weight and gross proteoform profile. Peptide mapping with UV and MS detection provides sequence confirmation and identifies post-translational modification sites. Glycan profiling — released glycan analysis by fluorescence and MS — characterizes the N-linked and, where relevant, O-linked glycan distribution. Hydrogen-deuterium exchange mass spectrometry provides higher-order structural information by interrogating backbone amide solvent accessibility. Circular dichroism and Fourier-transform infrared spectroscopy characterize secondary structure content and tertiary folding. Disulfide bond mapping confirms the connectivity pattern. Immunochemical methods — ELISA-based binding, surface plasmon resonance, or biolayer interferometry — characterize epitope binding and receptor engagement. Each of these methods has its own validation requirements under FDA’s 2015 guidance on Analytical Procedures and Methods Validation for Drugs and Biologics — a guidance that remains in effect and now operates alongside ICH Q2(R2), adopted by ICH in November 2023 and issued as final FDA guidance in March 2024, which supplies the harmonized international framework for analytical procedure validation lifecycle management — and the validation data for similarity-specific methods must be included in the submission.

    ICH Q5E, the international guideline on comparability of biotechnological and biological products subject to changes in their manufacturing process, provides an important conceptual parallel. The comparability exercise described in Q5E — demonstrating that a post-change product retains the quality, safety, and efficacy profile of the pre-change product — is structurally analogous to the biosimilar analytical similarity exercise. The key difference is that in the Q5E context the manufacturer knows both the pre- and post-change processes and can design the comparability study with full knowledge of the product’s manufacturing history. In the biosimilar context, the applicant does not have access to the reference product’s manufacturing process, its in-process controls, its raw material specifications, or its historical characterization database. The applicant must reconstruct the reference product’s variability portrait entirely from commercial lots obtained through procurement.

    This is why the number of reference product lots matters so much, and why lots from mixed expiry date cohorts are required. A lot close to its expiry date represents a product that has aged at the far end of its approved shelf life. It may show degradation profiles — deamidation, oxidation, glycation, aggregation — that are meaningfully different from a freshly released lot. If your reference product characterization dataset includes only recently released lots, you have not characterized the full span of variability that the reference product exhibits across its commercial lifetime. FDA’s current final guidance is explicit on this point: it recommends a minimum of ten reference product lots — generally matched by a comparable number of biosimilar lots in the similarity study itself — and it expects those lots to span a range of expiry dates so that the quality range you derive actually reflects the product’s real-world variability envelope. A quality range derived from five recently released lots with similar expiry dates is not a quality range — it is a characterization of a single point in the product’s commercial history, and the acceptance criteria derived from it will be challenged.

    The immunogenicity dimension of the analytical similarity package is often underweighted in early program planning. Clinical pharmacology studies — single-dose PK studies comparing the proposed biosimilar to the reference product, typically using a sensitive patient population or healthy volunteers — are almost invariably required. The design of those studies, including the selection of endpoints, the sample size justification, and the immunogenicity monitoring plan, must be coherent with the analytical similarity data. A biosimilar whose glycan profile shows a shifted fucosylation distribution relative to the reference product lot distribution should anticipate questions about the impact on Fc receptor engagement and potentially on immunogenicity. Those questions should be anticipated in the analytical similarity package, not encountered for the first time during the clinical review.

    The totality of the evidence — analytical similarity data from a well-designed, pre-specified three-tier program; state-of-science structural characterization; reference product variability established from an adequate lot set; validated methods; coherent clinical pharmacology and immunogenicity design — is what 351(k) requires. A submission that presents biosimilar specification results against reference product specification limits is not a totality-of-evidence package. It is a specification comparison, and it does not answer the question Congress embedded in the statute. Structural similarity, demonstrated attribute by attribute relative to the reference product batch distribution — that is what analytical similarity means.

    ──────────────────────────────────────────────────────────────────────────────── THE XGENE BIOSIMILAR ANALYTICAL SIMILARITY ARCHITECTURE

    XGene structures every 351(k) analytical similarity program around three integrated components, each designed to be completed in sequence before the first biosimilar lot is characterized for similarity purposes.

    COMPONENT 1 — REFERENCE PRODUCT PROCUREMENT AND CHARACTERIZATION PLAN Procurement of ≥10 reference product lots with mixed expiry dates (minimum two expiry date cohorts spanning at least 50% of approved shelf life). Comprehensive state-of-science characterization of all lots: intact mass, subunit analysis, peptide mapping (UV + MS), N-linked glycan profiling, HDX-MS for higher-order structure, CD and FTIR for secondary/tertiary structure, disulfide mapping, immunochemical binding panel, and cell-based potency assay(s). All characterization is completed before any biosimilar lot enters the similarity testing phase.

    COMPONENT 2 — PRE-SPECIFIED ANALYTICAL SIMILARITY STUDY DESIGN Attribute assignment to Tier 1, Tier 2, or Tier 3 is finalized based on reference product characterization data and published scientific literature on mechanism of action and structure-activity relationships. All Tier 1 equivalence margins (90% CI, 0.80–1.25) and Tier 2 quality range boundaries (derived from reference product lot distribution, typically mean ± k×SD with k selected based on lot number and desired coverage probability) are locked in the pre-specified Data Analysis Plan before biosimilar lot characterization begins. The DAP is version-controlled and date-stamped to establish pre-specification.

    COMPONENT 3 — TOTALITY-OF-EVIDENCE INTEGRATION AND BLA MODULE 3 AUTHORING Statistical similarity testing output (Tier 1 equivalence test results with 90% CI; Tier 2 quality range assessment; Tier 3 descriptive tables) is integrated with the full structural characterization dataset and method validation data into a coherent Module 3 narrative that explains not only what the data show but why the program was designed as it was, what the reference product characterization revealed about natural variability, and how the biosimilar’s analytical profile is positioned within that variability. Clinical pharmacology study design is cross-referenced to the analytical similarity findings to demonstrate consistency between the analytical and clinical evidence streams.

    Deliverable: A Module 3 analytical similarity section that reads as a pre-specified scientific study — because it was one — rather than as a post-hoc data assembly. ────────────────────────────────────────────────────────────────────────────────