XGene CMC IntelligenceXGene Intelligence

ISO IDMP 11238 — Substance Identification and UNII Registration

StabilitySolid StateIDMP / SPORBiologics

A pharmaceutical company that does not have UNII registrations for all of its active substances has a regulatory data foundation problem that will surface at every submission — FDA, EMA,…

By Khaled Aamer, PhD · Founder, XGene LLC Aug 22, 2026 10 min read
On this pageArticle overview

    ISO IDMP 11238 — Substance Identification and UNII Registration

    A pharmaceutical company that does not have UNII registrations for all of its active substances has a regulatory data foundation problem that will surface at every submission — FDA, EMA, and the global IDMP infrastructure all require unique substance identifiers, and the companies that have not built substance identification as a systematic regulatory competency are managing this risk reactively.

    I have watched this play out in a predictable pattern. A regulatory team is preparing a major submission — an NDA, an MAA, a variation — and the data review reveals that the active substance does not have a validated UNII code, or that the UNII in the FDA submission does not match the substance record in the EMA’s SPOR Substance Management Service, or that a salt form has been used interchangeably with the free base without recognizing that they are distinct UNII registrations. The submission is imminent. The fix is not imminent. The gap between those two facts defines the crisis.

    This is a solvable problem. But it is only solved upstream — in the product development workflow, at the moment of first synthesis or first formulation decision — not at the filing deadline. The thesis I want to make explicit today is this: UNII registration is not a one-time task assigned to a regulatory coordinator. It is a systematic competency that must be embedded in the product development workflow from the first synthesis, because the downstream consequence of an unregistered substance is a submission validation failure at the most consequential moment.

    WHY UNII IS THE FOUNDATION OF GLOBAL PHARMACEUTICAL REGULATORY DATA

    ISO 11238:2018 — “Health informatics — Identification of medicinal products — Data elements and structures for the unique identification and exchange of regulated information on substances,” which superseded the original 2012 edition — published under the broader ISO IDMP family of standards, defines the framework for identifying and describing medicinal product substances at the level of specificity required for global regulatory exchange. The standard does not simply ask “what is this substance called?” — it asks for a complete, unambiguous scientific characterization that allows two independent parties to confirm they are discussing the same molecular or biological entity.

    The UNII — Unique Ingredient Identifier — is the FDA’s implementation of this principle in an operational regulatory data system. Assigned through the FDA’s Substance Registration System, known as FDASIS, and maintained in collaboration with the United States Pharmacopeia, UNII codes are ten-character alphanumeric strings that serve as globally recognized identifiers for FDA-regulated substances. They are not proprietary codes for internal FDA use. They are internationally recognized as the ISO 11238-compliant identifier for substances in FDA-regulated products, cross-referenced by PubChem at NCBI, integrated into the WHO INN database, and accepted as external identifier cross-references within the EMA’s SPOR SMS system.

    Understanding why UNII functions as the foundation of regulatory data requires understanding what ISO 11238 demands of substance identification. The standard establishes distinct identification requirements across its recognized substance classes — chemical, protein/biological, nucleic acid, polymer, and structurally diverse entities (the class under which botanical and herbal materials are registered) — and each class demands a different type of scientific data. The four practitioner-facing categories most relevant to a typical CMC program are worked through below.

    For chemical substances — the category that encompasses most small molecule drugs — ISO 11238 requires a molecular formula, an InChI string (the IUPAC International Chemical Identifier that encodes molecular structure unambiguously), a stereochemistry descriptor, and typically a CAS registry number. This is not merely naming a compound. It is defining the precise molecular entity, including stereochemical configuration, salt form, and hydration state, such that a regulatory authority anywhere in the world can confirm with certainty that the substance in a submission is the entity that has been registered.

    The stereochemistry dimension is operationally important and routinely underestimated. Different stereoisomers of the same base molecule have distinct UNII codes. Racemic mixtures have distinct UNII codes from their individual enantiomers. Salt forms — the sodium salt, the hydrochloride salt, the free acid — each have distinct UNII codes. This is not regulatory formalism for its own sake. Different stereoisomers have different pharmacological profiles. Different salt forms have different physicochemical properties that affect bioavailability, stability, and manufacturability. Regulatory science requires that these distinctions be made explicit in submission data, and UNII is the mechanism through which they are made explicit in a machine-readable, globally exchangeable format.

    For biological substances — monoclonal antibodies, therapeutic proteins, recombinant enzymes — ISO 11238 recognizes that a molecular formula is insufficient for identification. A monoclonal antibody cannot be characterized by a chemical formula alone. Its identity under ISO 11238 requires sequence data, glycosylation pattern characterization, species of origin, and a manufacturing method summary sufficient to establish that the biological substance is the specific molecular entity described. This requirement reflects a fundamental truth about biologics: two antibodies with identical amino acid sequences produced by different cell lines under different manufacturing conditions may not be the same substance from a regulatory standpoint. ISO 11238 encodes this scientific reality into the identification framework.

    For herbal substances, the standard requires identification of the part of the plant used, the species designation, and the preparation type — recognizing that an extract of Echinacea root is not the same as an extract of Echinacea aerial parts, and that a dry extract is not the same as a tincture, even when derived from the same plant species.

    For nucleotide sequences — increasingly relevant as oligonucleotide therapeutics and mRNA-based medicines move through development pipelines — the standard requires sequence data and modification type characterization, recognizing that modified nucleotides alter pharmacological and stability properties in ways that demand distinct identification.

    The UNII system operationalizes all of these requirements through the FDASIS registration process. When a sponsor submits a substance registration through FDASIS, they are providing the scientific data package that allows FDA and USP to assign a UNII that carries the ISO 11238-compliant characterization of that specific entity. The UNII then propagates into the FDA’s internal data systems, into public-facing tools including FDASIS and PubChem, into the WHO INN database for internationally nonproprietary name cross-referencing, and serves as the external identifier bridge to EMA’s SPOR SMS.

    The dual-registration reality — that submissions to FDA require UNII cross-referencing while submissions to EMA require SPOR SMS registration — creates a data management obligation that many companies do not address systematically until they are simultaneously managing FDA and EMA filings. By that point, discovering that a substance is inconsistently characterized across the two systems, or that a UNII registered in FDASIS has not been cross-referenced in the SPOR SMS record, introduces a data reconciliation task under submission pressure that should have been resolved at IND/CTA filing or earlier.

    The companies that manage substance identification well treat UNII registration as a precondition for any regulatory filing — not a prerequisite that can be satisfied in parallel with submission preparation. A UNII registration for a straightforward small molecule through FDASIS typically takes four to eight weeks. Complex biologics requiring sequence, glycosylation, and manufacturing method data can require eight to sixteen weeks. These timelines are not compatible with retroactive resolution at a filing deadline.

    THE FDASIS REGISTRATION PROCESS: DATA REQUIREMENTS AND TIMELINE

    The FDASIS submission process for UNII registration is governed by technical requirements that reflect ISO 11238’s substance category distinctions. For a chemical substance, the submission package must include the molecular formula, the InChI or SMILES structure string, the stereochemistry specification, the CAS registry number where one exists, and the substance names — systematic, common, INN where assigned, and trade names. The level of structural specificity required means that ambiguous or incomplete characterization of the substance — a common situation in early development when final salt form decisions have not been made — will delay registration.

    This is a practical argument for establishing the intended regulatory substance identity as early as possible in development. The salt form selected for clinical development is a regulatory substance that needs a UNII. If that decision changes between Phase 1 and Phase 3 — as it sometimes does when a different salt form demonstrates superior stability or bioavailability — the new salt form requires a new UNII registration, and the regulatory history must clearly document the relationship between the two substance identities.

    The FDASIS public database and PubChem serve as the verification tools. Before submitting a registration request, a sponsor should confirm that a UNII does not already exist for their substance — FDASIS and PubChem both provide searchable UNII lookup by name, CAS number, InChI, and structure. Duplicating a registration for an already-registered substance creates data inconsistency problems in downstream regulatory systems. Cross-referencing the UNII identified through FDASIS and PubChem against the WHO INN database should also be routine practice for any substance that has been assigned or is seeking an INN.

    BIOLOGICAL SUBSTANCE IDENTIFICATION UNDER ISO 11238: WHY IT CANNOT BE TREATED LIKE A SMALL MOLECULE

    The operational mistake I see most often in biological substance registration is attempting to register a biologic with the data package appropriate for a small molecule. This approach reflects a misunderstanding of what ISO 11238 requires for biological substances and why those requirements exist.

    A therapeutic protein is not characterized by a formula. It is characterized by its sequence — amino acid sequence, disulfide bond pattern, post-translational modifications — and by the manufacturing context in which that sequence is expressed and processed. Glycosylation pattern is not incidental information for a biologic. The glycan structures on a therapeutic antibody affect its pharmacokinetics, its immunogenicity risk, and its effector function. Two biologics with identical amino acid sequences but different glycosylation profiles produced by different cell lines are, under ISO 11238 and under regulatory science, different substances.

    The UNII registration for a biologic therefore requires sequence data in a format compatible with the FDASIS system, characterization of relevant post-translational modifications, and manufacturing method information sufficient to establish the substance identity in context. This data package is substantially more complex than a chemical substance registration, which is why the registration timeline for complex biologics can extend to sixteen weeks, and why sponsors of biologic programs who leave UNII registration to the submission preparation phase face genuine operational risk.

    The SPOR SMS cross-reference obligation compounds this. For a biologic intended for both FDA and EMA submissions, the biological substance identification data must satisfy both FDASIS registration requirements and SPOR SMS technical specifications for biological substance characterization. Verifying consistency between the two registrations — same sequence data, same glycosylation characterization, same manufacturing method representation — is a data governance task that requires deliberate process ownership, not ad hoc verification at filing time.

    THE XGENE PHARMACEUTICAL SUBSTANCE REGISTRY AND IDMP COMPLIANCE PROGRAM

    The XGene Pharmaceutical Substance Registry and IDMP Compliance Program establishes substance identification as a systematic regulatory competency embedded in the product development workflow — not a filing-period corrective action.

    The program delivers:

    Complete Portfolio UNII Registration Audit — A structured review of every active substance, salt form, excipient, and biological entity across the product portfolio. Each substance is verified in FDASIS, cross-referenced in PubChem, and assessed against SPOR SMS registration status. Gaps, inconsistencies, and missing registrations are documented with remediation priority assignments.

    New Substance UNII Registration Protocol — A defined process for initiating UNII registration at first synthesis or formulation decision, including data package preparation requirements by substance category (chemical, biological, herbal, nucleotide), FDASIS submission procedures, and timeline management aligned to the development milestone schedule.

    SPOR SMS Cross-Reference Maintenance — A process for establishing and maintaining UNII as an external identifier cross-reference within the SPOR SMS substance record, ensuring consistency between FDA and EMA substance identification data for every product in the portfolio.

    Excipient UNII Registry for PQ-CMC Compliance — A maintained internal registry of UNII codes for every excipient in the formulation portfolio, structured to meet PQ-CMC structured data requirements for inactive ingredients in FDA submission content.

    Biological Substance ISO 11238 Identification Data Package Template — A standardized data package template for biological substance UNII registration, incorporating sequence data structure, glycosylation characterization format, manufacturing method summary format, and species of origin documentation aligned to ISO 11238 biological substance requirements.

    Internal Substance Registry — A maintained internal data asset linking UNII, SPOR SMS ID, CAS registry number, INN, and proprietary substance identifiers for every substance in the portfolio, serving as the single source of truth for substance identification data across CMC, regulatory, and pharmacovigilance functions.