XGene CMC IntelligenceXGene Intelligence

ALCOA+ Data Completeness — The Deletion Problem FDA Finds First

SpecificationsStabilityData Integrity / ALCOA+FDA Warning LettersFDA 483

FDA investigators are trained to look for what is missing from your data records before they look at what is present — and the deletion of failed runs, rejected results,…

By Khaled Aamer, PhD · Founder, XGene LLC Aug 22, 2026 17 min read
On this pageArticle overview

    FDA investigators are trained to look for what is missing from your data records before they look at what is present — and the deletion of failed runs, rejected results, or inconvenient data points is the data integrity violation that most reliably transforms a 483 observation into a Warning Letter.

    That sentence reflects one of the most consistent and consequential patterns in FDA data integrity enforcement over the past decade, and it points directly to the structural vulnerability that most pharmaceutical laboratories carry into every inspection cycle without recognizing it: GMP data systems that are technically compliant in their record-keeping architecture but operationally incomplete in their data retention practice. The audit trail captures what was entered. It does not always capture what was deleted, why it was deleted, or whether the deletion was scientifically justified and management-authorized. The injection sequence file shows runs 001 through 047. No documentation explains why runs 023, 031, and 038 do not appear in the corresponding results report. The stability data package transmitted to the regulatory authority contains time points at 0, 3, 6, 9, and 12 months. It does not contain the 6-month data for the 40°C/75% RH condition that was flagged as out-of-trend and then excluded without a formal investigation. These are not hypothetical scenarios. They are representative of the patterns FDA investigators extract from instrument data files and electronic systems in routine inspections, and they are the patterns that, once documented, very rarely resolve without significant remediation effort and regulatory consequence.

    What Completeness Requires: The Full Data Set Principle in FDA Enforcement Positions

    The ALCOA+ principle of completeness is deceptively simple in its statement and deeply demanding in its operational application. Completeness requires that all data generated in the conduct of a GMP activity is retained and available for review — not only the data that supports the desired conclusion, not only the data from runs that meet predefined acceptance criteria, not only the time points that fall within specification, and not only the analytical results that are consistent with the product’s approved profile. All data means the failed injection sequences, the aborted system suitability runs, the retested samples where the first result was anomalous, the stability time points where the reading was out-of-trend, and the process batch records for batches that were rejected rather than released. The regulatory basis for this requirement sits across multiple provisions. Under 21 CFR 211.192(a), all laboratory records — including those for tests that fail specifications — must be retained, reviewed, and, where failures are noted, investigated for cause. Under 21 CFR 211.180, records must be retained for the period specified in the regulation — and destruction of records before the applicable retention period constitutes a separate and independent violation. Under 21 CFR 211.68(b), automatic, mechanical, and electronic equipment used in the generation and management of GMP records must produce records that are accurate and complete, and the completeness of those records is subject to verification through audit trail review during inspection.

    The FDA Guidance for Industry: Data Integrity and Compliance with Drug CGMP (2018) operationalizes the completeness principle directly. The guidance is explicit that all raw data — the original instrument output, the original electronic files, the complete audit trail — must be retained in a form that is attributable, legible, contemporaneous, original, and accurate. It further specifies that original data includes all laboratory data whether or not the data is used to support disposition decisions, and that the absence of data from a record when that data should exist is itself a data integrity finding. The MHRA GMP Data Integrity Definitions (2018) aligns with this position and adds the practical test: if data exists in a system and is not presented in the corresponding report, the gap between the system record and the reported record is the starting point for an investigator’s inquiry. WHO TRS 996 Annex 5 (2016) frames the same principle in terms of the data governance requirement: organizations must establish controls that prevent the selective retention of data supporting a favorable outcome while eliminating or concealing data that does not.

    The enforcement consequence of completeness failures is more severe than that of most other ALCOA+ violations because completeness failures directly implicate product quality decisions. When data is missing from a record, there is no way to independently verify that every result generated for a batch was reviewed before disposition was decided. An analyst who deletes a failing HPLC result and retests until a passing result is obtained does not produce a batch record showing a failing result followed by investigation followed by authorized retesting. They produce a batch record showing a single passing result that appears, on its face, to satisfy every specification requirement. The product may have been released on the basis of that record. Multiple batches may have been released on the basis of the same practice. FDA investigators understand this, and they approach completeness assessments not as administrative audits of filing completeness but as investigations into whether the data record actually supports the disposition decisions made from it.

    The HPLC injection sequence is the most consistent and technically specific tool FDA investigators use to identify completeness gaps in the pharmaceutical laboratory. The injection sequence file is a system-generated record that assigns a sequential number to every injection performed on the instrument from the time the sequence was initiated. Unlike the analyst-generated results report, the injection sequence file cannot be edited without leaving a visible audit trail signature, and the sequential numbering cannot be re-ordered without creating a visible gap. When an FDA investigator sits down with a laboratory chromatography data system — whether Empower, Chromeleon, OpenLAB, or any other platform — one of the first queries they run is a sequence file completeness review: they look at the injection numbers assigned to every run executed on the system over the inspection period and compare them against the results reported in the batch records and analytical worksheets. Missing injection numbers — numbers that exist in the sequence file but are absent from the corresponding results report, with no documented explanation in the audit trail — are flagged immediately as a completeness gap. The gap may reflect an aborted run, a system suitability failure that was appropriately investigated, or an instrument malfunction that was documented and assessed. But if the documentation explaining the missing injection is not present and contemporaneous, the gap appears to the investigator as selective deletion — and the investigation that follows proceeds from that assumption.

    Aborted runs carry a specific documentation requirement that many laboratories manage inadequately. An aborted run is one in which the analytical sequence was started and then terminated before completion, for any reason — instrument malfunction, reagent failure, incorrect sample preparation, system suitability failure, or analyst decision. Every aborted run must be documented in the laboratory record with the reason for the abort, the identity of the analyst who initiated and terminated the run, and the disposition of any partial data generated before termination. If the partial data met a specification limit and was retained rather than aborted, that data must be reviewed as part of the complete data set, not selectively excluded because it is inconvenient to the result. The FDA 2018 guidance is specific: the existence of an aborted run is not itself a violation, but the absence of documentation explaining an aborted run is a completeness violation, and the pattern of aborted runs that consistently precede passing results is a data integrity finding warranting investigation.

    Retest data completeness follows the same principle with higher regulatory visibility. Under 21 CFR 211.192(a), retesting of samples that yielded initial out-of-specification results must be conducted under a written procedure, the initial failing result must be retained and included in the batch record, and all retest results must be reported regardless of whether they support release. The practice of retesting until a passing result is obtained and then reporting only the passing result — without documenting the original failure, the investigation of the failure cause, and the scientific justification for accepting the retest result in lieu of the original — is one of the most documented data integrity patterns in FDA Warning Letters. It is not a gray area. It is not an area where investigator discretion determines the outcome. It is a defined completeness violation that generates 483 observations in nearly every inspection where it is identified, and it generates Warning Letters when the pattern is systematic.

    How Deletion, Overwrite, and Test-Until-You-Pass Become Data Integrity Findings

    The technical mechanism by which deletion, overwrite, and selective retention become FDA findings is the audit trail. Every GMP-compliant electronic system is required, under 21 CFR 211.68(b) and 21 CFR Part 11, to maintain a secure, computer-generated, time-stamped audit trail that records the date and time of operator entries and actions that create, modify, or delete electronic records. The audit trail requirement means that deletion is not invisible — but the completeness of the audit trail, and the ability of investigators to extract and interpret it, determines whether deletion activity is identified and assessed. A laboratory in which the audit trail is technically enabled but operationally unreviewed is a laboratory in which deletion can accumulate without detection until an FDA investigator conducts the review that internal quality systems failed to perform.

    The specific data integrity enforcement pattern that FDA has documented across multiple Warning Letters involves one of three mechanisms: direct deletion of electronic raw data files from the system, overwrite of original data with revised values without retention of the original, or generation of data outside the GMP electronic system — on personal computers, in manual worksheets, or in parallel unofficial records — followed by selective entry of only the favorable results into the official system. Each mechanism leaves a different audit trail signature. Direct deletion of a raw data file should leave a deletion event in the system audit trail with a timestamp, user ID, and original file identifier — if the audit trail is configured to capture deletion events, and if the deleted records are recoverable. If the system is configured to purge deleted records rather than mark them as deleted and retain them, the file is gone and the audit trail entry is the only evidence of its prior existence. Overwrite of original data generates an audit trail event showing the original value, the revised value, the timestamp of the revision, and the identity of the user who made the change — the absence of a justification entry explaining the revision is the basis for a completeness finding. Generation of data outside the official system is the hardest to detect without physical inspection of laboratory equipment and personal computing devices, but it is also the finding that most consistently appears in Warning Letters because investigators who identify unofficial data generation typically find a pattern that spans multiple analysts, multiple products, and multiple years.

    Stability data completeness is the dimension of this problem with the most direct regulatory submission consequence. A stability program supporting a drug application must include all time-point data generated at all conditions and intervals specified in the approved protocol. When stability data is submitted to FDA in a New Drug Application or supplement, the agency reviewer expects that the submitted data set is complete — that no time points have been excluded, that no out-of-trend results have been suppressed, and that the statistical trend analysis performed on the data represents the actual performance of the product under the storage conditions tested. When FDA inspectors compare the data in an approved application against the data in the sponsor’s stability data system and find time points that are in the system but not in the submission, or time points that were retested after an initial out-of-trend result without documentation of the OOT investigation, the finding is not limited to a 21 CFR 211.192 observation. It becomes a question of whether the approved shelf life is supported by the complete data set — and if the complete data set, including the excluded time points, does not support the approved shelf life, the product quality implication extends to every batch released under that shelf life determination.

    The Metadata Layer: What FDA Investigators Extract from Instrument Data Files

    The metadata embedded in electronic instrument data files is the technical foundation of FDA’s data completeness investigation capability in the modern inspection environment. Metadata — the data about data — includes the file creation timestamp, the last modification timestamp, all audit trail entries, the user login identity associated with each action, the instrument ID and method parameters in use at the time of each run, and, in systems that capture it, the sequence of commands issued by the operator to the instrument control software. When an FDA investigator accesses a chromatography data system during an inspection, they are not reviewing the printed results report. They are interrogating the underlying electronic records — the .raw files, the .lcd files, the .dat files, or whatever native format the specific CDS uses to store the instrument output — for the metadata layer that the printed report does not contain.

    The specific metadata elements that are most frequently referenced in FDA 483 observations and Warning Letters arising from data completeness reviews are: the file creation time of raw data files relative to the batch manufacturing timeline (files created outside the period when the batch was being manufactured are a red flag); the user identity associated with data generation relative to the shift record and access control log (data files generated under user credentials inconsistent with the shift record require explanation); the injection sequence numbers embedded in the raw data relative to the numbers appearing in the results report (gaps in the sequence are the primary injection sequence completeness check); and the processing parameters applied to the raw data file, including integration parameters, at the time the result was calculated (retrospective changes to integration parameters that consistently move a borderline result across a specification limit are a data manipulation finding independent of the completeness analysis).

    The completeness check for injection sequences has become one of the most standardized and systematic components of FDA’s laboratory data integrity inspection protocol. Investigators request the complete sequence file for every analytical method relevant to the batches under review, extract the injection number range, and compare the extraction to the results reported. The check takes minutes in a well-organized CDS and yields immediate visibility into any run that was performed but not reported. The check also extends to the system audit trail: for every injection number that is missing from the results report, the investigator looks for an audit trail entry explaining the disposition of that run. Documented aborted runs with contemporaneous explanations are reviewed and may or may not generate observations depending on the pattern and frequency. Undocumented missing injections — injections that are in the sequence file, absent from the results report, and unaccounted for in the audit trail — are completeness violations with no available defense.

    The data migration completeness dimension is a relatively newer enforcement focus that has grown in prominence as laboratories transition between chromatography data systems. When a laboratory migrates data from one CDS platform to another — from a legacy system to a current GxP-compliant platform — the migration must be validated to confirm that all data transferred completely and accurately. Every raw data file that existed in the source system must be verifiably present in the destination system, with metadata integrity confirmed. Every audit trail entry from the source system must be preserved and accessible in the destination system or in an accessible archive. Migration validation gaps — cases where the data migration was performed without a validated comparison of source and destination data sets — create the same completeness vulnerability as internal deletion: FDA investigators who review migrated data systems and find that the metadata does not match the source system documentation, or that raw data files are present in printed reports but absent from the electronic archive, treat the gap as a completeness violation requiring investigation.

    The XGene Data Completeness Assurance and Audit Program

    XGene Framework for ALCOA+ Data Completeness — The Deletion Problem FDA Finds First
    XGene Framework

    Eliminating the Deletion Problem Before FDA Arrives

    The XGene Data Completeness Assurance and Audit Program is a structured completeness verification program built against the specific completeness gap patterns that FDA investigators document under 21 CFR 211.192(a), 21 CFR 211.68(b), and 21 CFR 211.180. The program operates across five integrated workstreams designed to identify and close completeness gaps before they become inspection observations.

    Workstream 1 — Audit Trail Completeness Verification for All GMP Electronic Systems: Every GMP electronic system that generates or stores analytical data — chromatography data systems, LIMS, process control systems, environmental monitoring systems — is subject to a periodic audit trail completeness review conducted by a QA function independent of the generating laboratory. The review confirms that the audit trail is active, that it captures all creation, modification, and deletion events with timestamp and user attribution, that deleted records are recoverable with attribution intact, and that no data exists in printed reports that is absent from the corresponding electronic archive. The review is conducted on a frequency proportional to the risk tier of the system and the inspection history of the facility, with critical GMP laboratory systems reviewed no less than quarterly. Findings are tracked in a CAPA system with resolution timelines proportional to regulatory risk.

    Workstream 2 — Injection Sequence Completeness Review: For every HPLC and other chromatographic system in the GMP laboratory, the XGene program conducts a systematic injection sequence review against the corresponding batch records and analytical results reports on a defined periodic basis — not deferred until inspection. The review extracts the complete injection sequence file for the review period, identifies every injection number in the sequence, and confirms that each injection number either appears in a documented results report or is explained by a contemporaneous audit trail entry documenting an aborted run, a failed system suitability, or an instrument event. Missing injection numbers without contemporaneous documentation are identified as completeness gaps, investigated for cause, and remediated before the next inspection window. The review output is a documented completeness certification for each system and each review period — a document that can be produced to an FDA investigator as evidence of proactive completeness management.

    Workstream 3 — Aborted Run Documentation and Stability Data Completeness Audit: The aborted run documentation workstream establishes a defined procedure for the documentation of every started-and-stopped analytical run, with mandatory fields for reason, analyst identity, time of abort, and disposition of any partial data. Laboratory SOPs are reviewed and updated to confirm that the aborted run documentation requirement is unambiguous, that training records confirm analyst understanding, and that the audit trail review workstream captures undocumented aborted runs as completeness findings. The stability data completeness audit is conducted across all active stability programs on an annual basis and confirms that every time point specified in the approved stability protocol has been tested, that every time point result — including out-of-trend results — has been documented in the stability data system and the corresponding regulatory submission, and that all OOT investigations have been completed and documented before data is reported.

    Workstream 4 — Development Data Completeness Review for CMC Submissions: Data submitted in CMC modules of regulatory applications is verified against the generating system for completeness before submission. Every analytical data table in a Module 3 submission is traced to its source records in the GMP or development electronic system, and the completeness of the source record set — including all runs, all time points, and all results regardless of whether they support the application narrative — is confirmed before the submission package is finalized. Data migration events associated with CMC submission preparation — laboratory moves, system transitions, data consolidation from multiple sites — are subject to the data migration completeness verification protocol, confirming source-to-destination completeness by file count, file integrity check, and audit trail continuity.

    Workstream 5 — QA Release Process with Systematic Completeness Check Before Batch Disposition: The batch release process is structured so that QA review of the batch record includes a defined completeness verification step before disposition is authorized. The completeness check confirms that the analytical record for the batch contains results for all tests specified in the approved specification, that all results — including any anomalous results, retested samples, and investigation outcomes — are present in the record, and that no injection sequences associated with the batch contain undocumented gaps. The check is documented in the batch record as a distinct QA action, creating a contemporaneous record of completeness review that predates the disposition decision. The output of this workstream is a release record in which completeness of the analytical data set is affirmatively documented by QA — not assumed from the absence of a flag, but confirmed by a defined check performed before every disposition decision.

    The output of the XGene Data Completeness Assurance and Audit Program is a GMP data environment in which every analytical result in every batch record is traceable to a complete source record, every injection sequence is accounted for in the corresponding results documentation, every aborted run is documented contemporaneously, every stability time point is present in both the data system and the submission package, and QA has affirmatively verified data completeness before every batch disposition decision. That is the data completeness standard FDA investigators expect to find when they access the chromatography data system during an inspection — and it is the standard that determines whether the inspection closes as a voluntary action indicated or advances to the Warning Letter process.

    The thesis that runs through every data completeness Warning Letter is the one stated at the outset: selective retention of passing results and deletion of failing results produces a data record that appears compliant while the product it describes may not be. Data completeness is not a documentation requirement that exists to satisfy an administrative preference for complete files. It is the foundational assurance mechanism that makes every other element of the GMP quality system meaningful — because a release decision made from an incomplete data set is not a quality decision. It is a guess, made from selectively favorable evidence, about a product that may or may not be what the record represents. FDA investigators have seen that pattern enough times, across enough facilities and enough product categories, that they have built their inspection methodology around finding it. The facilities that survive that methodology are the ones that have built their internal data management programs around the same systematic completeness discipline before the investigator arrives.