PQ-CMC / KASA — Introduction to FDA’s Structured CMC Data Initiative
FDA is building a structured CMC data submission system that will fundamentally change how Module 3 packages are created, reviewed, and processed — and most pharmaceutical companies are not yet…
On this pageArticle overview
FDA is building a structured CMC data submission system that will fundamentally change how Module 3 packages are created, reviewed, and processed — and most pharmaceutical companies are not yet aware that the first structured data requirements are already in effect for certain submission types.
That statement is not a prediction. It is a description of where the regulatory landscape already stands.
The FDA Pharmaceutical Quality — Chemistry, Manufacturing, and Controls program, known as PQ-CMC, is the agency’s initiative to convert Module 3 submission content from unstructured narrative documents into machine-readable, structured data formatted according to HL7 FHIR profiles developed specifically for pharmaceutical quality. The companion infrastructure, KASA — Knowledge-Aided Assessment and Structured Application — is the FDA reviewer-side system that will consume, parse, and cross-reference that structured data automatically, reducing the manual extraction burden that has historically added weeks and uncertainty to CMC review cycles. Together, PQ-CMC and KASA represent the most significant transformation in CMC regulatory submission architecture since the introduction of the eCTD in the early 2000s.
The companies that understand this now and begin aligning their data systems today will have a structural competitive advantage in review speed and deficiency reduction. The companies that treat this as a future-state concern will face costly, time-consuming retroactive data restructuring when submission deadlines arrive.
What PQ-CMC Is and Why FDA Built It: The Knowledge-Aided Assessment Architecture
To understand why FDA built PQ-CMC, you have to start with what reviewers have been doing manually for decades. A CMC reviewer receives a Module 3 submission — potentially thousands of pages across 3.2.S and 3.2.P sections — and must extract, compare, and cross-reference specifications, batch analyses, stability data, and product composition information from narrative documents, tables embedded in Word files converted to PDF, and supporting data appendices. There is no consistent field structure, no machine-readable format, and no automated mechanism for the reviewer to check whether the acceptance criterion stated in the specification section matches the values reported in the batch analysis section or aligns with the stability trending data provided in the long-term stability summary. Every cross-reference is manual. Every inconsistency is discovered only when a human reviewer happens to look in the right place.
FDA’s PQ-CMC initiative was designed to solve exactly that problem. By requiring sponsors to submit Module 3 data in structured HL7 FHIR format — using defined resource types including MedicinalProductDefinition, Ingredient, ManufacturedItemDefinition, RegulatedAuthorization, and SubstanceDefinition — the agency creates a submission layer where individual data elements are machine-readable, consistently structured, and directly processable by KASA without manual extraction.
The data model is precise by design. A specification test in a PQ-CMC-compliant submission is not a sentence that reads “Assay: 98.0% to 102.0% by HPLC.” It is a structured set of fields: test name, method reference, acceptance criterion operator, acceptance criterion numeric value, acceptance criterion units, and validation status. Each field is a discrete, addressable data element. KASA can pull all acceptance criteria across all specifications in a submission in seconds, compare them against the batch analysis results reported in corresponding structured data elements, and flag inconsistencies automatically before a reviewer ever opens the file.
This is what “knowledge-aided assessment” means in practice. It is not artificial intelligence in the promotional sense. It is systematic, structured data comparison enabled by a consistent data architecture that sponsors and FDA agreed to build together.
The scope of PQ-CMC covers the core content domains of Module 3: drug substance specifications, drug product specifications, stability data, batch analyses, and product composition. The program has been implemented in phases, with pilot programs beginning in 2020 and formal requirements expanding over subsequent years as FDA published updated Implementation Guides through HL7 and issued guidance through the Federal Register. The phased expansion is designed to ultimately cover all structured content across Module 3, not just specifications and batch analyses. FDA published structured data submission expectations in the Federal Register in 2020, signaling the formal transition from pilot to program.
The data standard itself is grounded in ISO/IEC 21090, the international standard for healthcare data types, which provides the underlying data type definitions that HL7 FHIR profiles for pharmaceutical quality reference. PQ-CMC is therefore not a standalone FDA invention — it is an FDA application of internationally recognized data standards, developed in coordination with HL7 and aligned to the ICH M2 eCTD framework.
The eCTD Sections Where Structured CMC Data Is Now Expected
A common point of confusion in CMC teams first encountering PQ-CMC is the assumption that structured CMC data replaces the eCTD. It does not. The eCTD remains the required submission format. PQ-CMC structured data files are submitted within the eCTD structure, specifically as additional data files accompanying the narrative document content that sponsors have always provided. Both are mandatory. They are separate and complementary, not interchangeable.
Within the eCTD, the sections where structured PQ-CMC data is expected or required — depending on submission type and current phase of FDA’s rollout — map directly to the content domains FDA identified as the highest-value targets for automation. The drug substance specification content in Module 3.2.S.4.1 is a primary target because specification data has a highly consistent logical structure — test, method, criterion — that translates directly into structured FHIR resources. Drug product specification content in 3.2.P.5.1 carries the same structure and the same requirement logic. Batch analysis data in 3.2.S.4.4 and 3.2.P.5.4 are targeted because they generate the comparison data that KASA needs to cross-reference against specifications. Stability data supporting summaries, covered in 3.2.S.7 and 3.2.P.8, are included in the phased scope because stability trending is one of the most manually intensive reviewer tasks and one of the highest-value automation targets.
The submission types where structured data requirements are already in effect or formally expected include certain NDA filings, BLA filings, and ANDA submissions, with scope determined by the specific section content and the IG version in effect at time of submission. The HL7 Pharmaceutical Quality — CMC Implementation Guides, published through HL7.org and updated through the PQ-CMC pilot feedback process, are the definitive reference for which FHIR resource types apply to which eCTD sections and what data elements are required versus optional within each resource.
This is the operational detail that most CMC teams miss. The question is not simply “does PQ-CMC apply to my submission?” The question is “which sections of my submission trigger structured data requirements under the current IG version, and do my source data systems — my LIMS, my ELN, my document management system — have the capability to generate FHIR-formatted output for those sections?” For the majority of pharmaceutical companies operating today, the honest answer to the second question is no, because those systems were built to generate narrative reports and PDF exports, not machine-readable FHIR bundles.
Retrofitting narrative CMC content to produce FHIR output is not a formatting exercise. It requires a data model redesign. The source data must be structured at the point of generation — at the instrument, in the LIMS, in the ELN — not retroactively converted from a finished Word table. That distinction defines the difference between a company that is genuinely PQ-CMC-ready and a company that has a compliance gap it has not yet recognized.
What Happens to Submissions Without PQ-CMC Data in the Post-Pilot Era
The transition from pilot program to formal expectation changes the consequences of non-compliance materially. During the pilot period, FDA collected structured data voluntarily from participating sponsors, used those submissions to refine the KASA system, and accepted non-structured submissions from non-participants without formal review consequence. That era is closing.
As FDA’s phased implementation timeline advances and structured data requirements become formally applicable to specific submission types, submissions that arrive without the required PQ-CMC data components will generate deficiencies. The deficiency is not that the narrative content is wrong — it may be perfectly accurate. The deficiency is that the required structured data layer is absent, meaning KASA cannot process the submission as designed, and the reviewer must either revert to fully manual extraction or request the structured data as a response to deficiency.
In practical terms, a missing PQ-CMC structured data component in an otherwise complete NDA or BLA submission creates a deficiency that extends review timelines, consumes FDA reviewer time, and forces the sponsor to generate structured data outputs under deadline pressure from systems that may not have been built to produce them. That is the worst possible scenario — not because the science is wrong, but because the data architecture was not designed in advance.
The deeper risk is organizational. CMC teams that treat PQ-CMC as a regulatory affairs documentation project rather than a cross-functional data architecture initiative will discover, when the deficiency arrives, that the gap cannot be closed quickly. Generating compliant FHIR output requires LIMS configuration, ELN data capture redesign, or the implementation of middleware capable of transforming source system data into valid FHIR bundles against the HL7 PQ-CMC IG profiles. None of those changes happen in the weeks between a deficiency letter and the response deadline.
The companies building PQ-CMC-compatible data infrastructure now — assessing their LIMS, ELN, and DMS systems for FHIR output capability, mapping their data models against the current HL7 IG, and designing their data capture workflows to generate structured outputs at the source — will submit complete packages the first time. They will not receive structured-data deficiencies. Their submissions will move through KASA-assisted review faster because the reviewer receives processable data rather than documents requiring manual re-extraction. That is the competitive advantage this initiative creates for organizations that act before the deadline, not after it.
The HL7 PQ-CMC FHIR Implementation Guide: Current Standard Status and What CMC Teams Need to Know Today
The technical foundation of PQ-CMC is the HL7 FHIR Implementation Guide for Pharmaceutical Quality — Chemistry, Manufacturing, and Controls, officially titled “HL7 FHIR Implementation Guide: Pharmaceutical Quality — CMC.” As of 2026, the IG is published as a Standard for Trial Use (STU) , which is a critical status distinction that shapes how sponsors should approach implementation. A STU designation means the standard is stable enough for testing and pilot implementation but is not yet final. Feedback from real-world implementation is actively being used to revise the standard before final release. CMC teams waiting for a “final” version before acting will miss the implementation window entirely — because the feedback period is itself the opportunity to shape the standard to your systems.
The IG’s most recent published release (Release 1.0, updated in 2024) defines the FHIR R5 resource structure for five core CMC data domains, each mapping directly to eCTD sections identified earlier in this article. The resource types are precise and worth memorizing for anyone building the technical bridge between your LIMS/ELN and FDA’s KASA. MedicinalProductDefinition captures the identity and description of the drug product or substance. Ingredient breaks down each component — including excipients and active ingredient — as a structured, addressable entity. ManufacturedItemDefinition describes the manufactured item (the drug substance or product). RegulatedAuthorization attaches regulatory identifiers (NDA, BLA, ANDA numbers) to the product record. And SubstanceDefinition profiles the drug substance itself, including its physicochemical properties, impurity limits, and reference standard information. Together, these five resource types create a machine-readable representation of what today requires hundreds of pages of narrative text and PDF tables.
The data elements within each resource are where the IG’s scope becomes operationally visible. For a drug substance specification in 3.2.S.4.1, the IG requires that each acceptance criterion be represented as a discrete FHIR element with attributes for test name ( testCode ), method reference ( method ), comparison operator ( comparator ), numeric value ( value ), and unit of measure ( unit ). Validation status ( isValidated ) and stability‐indicating designation ( isStabilityIndicating ) are separate data elements. This is not an encoding of the specification document — it is a complete re-architecture of the specification as data . A LIMS that outputs a PDF report of a specification table cannot meet this requirement. A LIMS that can generate a FHIR bundle conforming to the IG’s profiles can.
The scope of the IG’s current STU release is intentionally bounded. The data domains covered are drug substance and drug product specifications, batch analysis results for release and stability, stability study design parameters and timepoint data, and product composition (ingredient lists with strength and reference information). These domains correspond to the eCTD sections FDA identified as the highest‐value automation targets: 3.2.S.4.1, 3.2.P.5.1, 3.2.S.4.4, 3.2.P.5.4, 3.2.S.7, 3.2.P.8, and the composition sections 3.2.S.1 and 3.2.P.1. The IG does not yet cover process validation data in 3.2.S.2.5 or 3.2.P.3.5, nor does it cover stability protocols or analytical method validation reports. Those content domains are expected to be added in future release cycles, but no public timeline for their inclusion has been announced. Sponsors should plan for phased expansion and design their source data systems to accommodate new resource types as they are published.
The IG’s STU status has direct consequences for implementation strategy. Because the standard is subject to revision based on implementation feedback, sponsors who lock in a rigid data pipeline to the current IG version without building modular, IG‐version‐aware transformation layers risk expensive rework when the next release adds new required fields or changes resource structures. The most defensible technical approach — and the one that aligns with FDA’s stated expectation — is to implement a middleware layer that ingests source system data in its native format, maps it to the current IG profiles, and can be reconfigured as the IG evolves. Hard‐coding FHIR output generation directly into your LIMS or ELN is not recommended unless those systems are designed to be updated rapidly in response to standard changes.
For CMC professionals, the IG should not be read as a technical specification alone. It should be read as a regulatory expectation document. FDA has explicitly stated that submissions using the IG’s profiles for the covered data domains will be accepted for review, and that KASA is being calibrated to process FHIR bundles conforming to the IG. A LIMS that cannot generate those FHIR bundles is not merely a technology gap — it is a submission readiness gap with regulatory consequences. The question for your organization today is not whether you will implement FHIR output. The question is whether you will implement it on your timeline — with adequate testing and integration — or under deadline pressure from a deficiency letter.
Open the current IG — available at the HL7 FHIR Implementation Guide for Pharmaceutical Quality — CMC — and ask two questions about each of the five core resource types: does your LIMS or ELN currently capture all the required data elements for that resource, and does your data architecture have a mechanism to transform that data into valid FHIR R5 format? The gap between yes and no defines your organization’s PQ-CMC readiness baseline.
XGene PQ-CMC Readiness Assessment and Implementation Framework

COMPONENT 1 — Current Submission Type Assessment Determine which of your active or planned INDs, NDAs, BLAs, and ANDAs trigger PQ-CMC structured data requirements under the current HL7 Implementation Guide version and FDA phased rollout schedule. Not every submission type is in scope today — but knowing precisely which ones are is the starting point for all downstream decisions.
COMPONENT 2 — Data System Audit Conduct a systematic audit of your LIMS, ELN, and document management systems to determine current FHIR output capability. The question is not whether these systems hold the right data — they almost certainly do. The question is whether they can export it in valid HL7 FHIR bundle format conformant to the PQ-CMC Implementation Guide profiles. For most systems currently in deployment, the answer requires middleware or configuration work to become yes.
COMPONENT 3 — Gap Analysis Against HL7 PQ-CMC IG Map your current data models — how specification tests, acceptance criteria, batch analysis results, and stability data are structured in your source systems — against the required FHIR resource structures in the current HL7 PQ-CMC IG. This analysis identifies exactly where data fields are missing, where nomenclature is inconsistent with required terminology, and where validation status fields are not captured at the source.
COMPONENT 4 — Data Model Redesign Roadmap Based on the gap analysis, develop a sequenced data model redesign roadmap that addresses source system changes in priority order, aligned to which submission types have the earliest structured data deadlines. This roadmap specifies LIMS configuration changes, ELN capture workflow modifications, middleware requirements, and validation protocols for FHIR output testing.
COMPONENT 5 — Implementation Plan Aligned to FDA Phased Rollout Build an implementation plan that phases your internal data system changes to match the FDA PQ-CMC rollout schedule, ensuring that structured data capability is in place — and tested against the current IG — before the first in-scope submission is filed. Implementation includes staff training on structured data concepts, integration testing of FHIR output against FDA validator tools, and a submission readiness review before first use.
PQ-CMC readiness is not a last-step regulatory check. It is a data architecture initiative that must be initiated well upstream of submission planning.
