Skip to content
Katalyze announces $10.5M seed round. Read the press release

Guide

The complete guide to pharmaceutical document digitization

Published September 15, 20267 min read

Tubes in a rack, close

In brief

Pharmaceutical document digitization is the process of reading GxP records such as executed batch records and certificates of analysis into a defined schema, with every value linked back to the page it came from. A scan by itself is a picture, not data. Done well, digitization produces structured data that reviewers can check against the source, that meets the regulatory expectation for accurate and complete records, and that other work can build on. This guide covers how digitization works, what OCR is and how to evaluate its accuracy, and how to build trust with digitization technology.

Key takeaways

  • Digitization is not scanning: a scanned page is an image, a digitized record is data with a schema and a source reference for every value.
  • FDA's 2018 data integrity guidance accepts electronic copies of paper records as true copies only when they preserve the content and meaning of the original, including the metadata needed to reconstruct the activity.
  • Extraction quality depends on the document type and image quality, so a defensible process pairs automated extraction with defined checks and expert review of anything uncertain.
  • Every extracted value keeps a link to the page and the position it was read from, which is what makes the data usable in a GxP environment.
  • Digitized records are most valuable when they connect to other data, so a batch's records, its materials' certificates and its deviations can be read together.

What is pharmaceutical document digitization?

Pharmaceutical document digitization turns the information inside operational records into structured data. The records include executed batch records, certificates of analysis, lab reports, recipes and process records, deviation records, SOPs and tech transfer documentation. The output is not a PDF. It is a set of fields and tables in a schema the operation defines, each value tied to the place on the page where it was read.

Three things separate digitization from scanning:

  • Structure. A scan keeps the page. Digitization keeps the values: the lot number, the hold time, the yield, the test result and its specification.
  • Lineage. Every extracted value keeps a link to its source, so a reviewer can see exactly where it came from.
  • Review. Values that fail a rule, or could not be read with confidence, are flagged for a person to resolve before the data is used.

A useful test: if a reviewer cannot click from a value to the line it came from, the record has been stored, not digitized.

Why paper and PDF records are still the norm

Much of the evidence behind a batch still lives on paper or in files no system reads. Executed batch records are often printed forms completed by hand on the floor, with corrections struck once and initialed, margin notes and signatures at every review. Certificates of analysis arrive as PDFs from suppliers, each with its own layout. Contract partners send batch records and reports as scans.

These records are complete, controlled and legally required. They are also hard to use. Answering a simple question, such as which batches used a given raw material lot and what the supplier reported for it, can mean pulling paper from an archive and retyping values into a spreadsheet. The data exists; it is not available in a form anyone can query.

Common GMP records and why their contents are hard to query.
RecordWhat it holdsWhy it's hard to use as data
Executed batch recordHeader fields, process tables, hold times, yields, corrections, margin notes, signaturesHandwritten values on printed forms, dozens of pages per batch
Certificate of analysisMaterial, lot, tests, specifications, results, methods, release signatureLayout changes with every supplier
Lab reportSample, instrument, results table, traces, analyst and reviewerResults and traces split across tables and images
Deviation recordDescription, impact assessment, actions, approvals, referenced batchFree text that points at other records
SOP and transfer documentationProcedure steps, revision history, parameters site to siteLong documents with nested tables and versions

What regulators expect of a digitized record

No regulation requires a manufacturer to digitize paper records. Several set the bar that a digitized copy has to meet if it will be relied on.

Predicate rules come first. 21 CFR 211.188 requires a batch production and control record for each batch with complete information about its production and control, including the identity of the people who performed and checked each significant step. 21 CFR 211.180 allows records to be kept "as original records or as true copies," provided they are accurate reproductions.

True copies must preserve meaning. FDA's guidance Data Integrity and Compliance With Drug CGMP: Questions and Answers (December 2018) says electronic copies can serve as true copies of paper or electronic records if they preserve the content and meaning of the original, including the metadata needed to reconstruct the activity. MHRA's 2018 GxP data integrity guidance defines a true copy as one verified, by a dated signature or a validated process, to hold the same information as the original, including its context and structure.

Data must be ALCOA+. FDA's guidance expects data to be attributable, legible, contemporaneously recorded, original or a true copy, and accurate. PIC/S PI 041-1 (2021) and WHO TRS 1033 Annex 4 (2021) add complete, consistent, enduring and available. For extracted data, that means knowing who or what captured each value, keeping the link to the original, and being able to show the value is accurate.

Electronic records bring Part 11. Where digitized data becomes the record relied on for a GMP decision, 21 CFR Part 11 applies: validated systems, secure time-stamped audit trails, limited access and authority checks. FDA's 2003 Part 11 scope guidance interprets the rule narrowly and points firms back to the predicate rules.

Regulatory expectations that apply to digitized GMP records.
ExpectationSourceWhat it means for a digitized record
Complete batch information21 CFR 211.188Every required field is captured, including who performed and checked each step
True copies preserve content and meaningFDA data integrity Q&A (2018), Q9The copy keeps the context needed to reconstruct the activity
Verified true copyMHRA GxP data integrity guidance (2018)A dated signature or a validated process confirms the copy matches
ALCOA+FDA (2018), PIC/S PI 041-1 (2021), WHO TRS 1033 Annex 4 (2021)Each value is attributable, accurate, linked to its original and available for the retention period
Audit trails and access control21 CFR 11.10Changes to extracted data are recorded and do not obscure earlier entries

How document digitization works, step by step

A governed digitization process moves each record through the same five steps. The order matters, because each step depends on the one before it.

  1. Extract. Read the text, tables, symbols, handwritten fields and margin notes on every page, from paper, scans and PDFs.
  2. Assemble. Organize what was read into the operation's schema, using its own document structures, terminology and field definitions, with every value linked to the page it came from.
  3. Check. Run defined rules against the assembled data: required fields present, values within expected ranges, units consistent, totals that add up, signatures where signatures belong.
  4. Review. Show each flagged item beside the page it came from so an expert can resolve it, then approve the record before anything moves downstream.
  5. Deliver. Send the approved, structured output to a configured destination, or connect it with other data for analysis.

This is the sequence Document Intelligence follows. The point of the check and review steps is that nothing uncertain passes silently. A value the system could not read with confidence is surfaced for review rather than guessed.

Batch records: the hardest and most valuable place to start

Executed batch records are the densest records in a plant. A single batch can run to dozens of pages of printed forms completed by hand, with values, times, initials, corrections and comments. They hold the process history that yield analysis, deviation investigation and batch release all depend on.

They are also where extraction is most likely to go wrong. Handwriting varies. A correction must be read as a correction: the struck value, the new value, the initials and the date are all part of the record. A margin note may be the one piece of context that explains a deviation. A digitization process for batch records needs to capture all of it and to show a reviewer the original for anything it was unsure of.

Certificates of analysis: supplier data you are required to verify

A certificate of analysis is a supplier's or a lab's statement of what a lot is: the tests run against a specification, each result, the method and a release signature. Every supplier formats it differently.

21 CFR 211.84(d)(2) allows a manufacturer to accept a supplier's report of analysis in place of full testing only if it performs at least one specific identity test on each component and establishes the reliability of the supplier's results through appropriate validation at appropriate intervals. That validation is much easier when certificate data is structured. With certificates in a schema, a team can trend a supplier's reported results across lots, compare them with in-house testing and spot a lot that is in specification but out of trend.

That is also why certificates are the usual starting point for Raw Material Characterization: the analysis starts with the certificate data being readable.

Accuracy, confidence and human review

No extraction method reads every page perfectly. Performance depends on the document type, the layout and the image quality. A defensible process is designed around that fact rather than hiding it.

  • Confidence. Each value carries an indication of how confidently it was read, so low-confidence values can be routed to review.
  • Deterministic checks. Rules written for the operation catch errors that confidence alone would miss, such as a yield greater than the theoretical maximum or a date before the batch started.
  • Expert review. Reviewers see the flagged value next to the source image and decide. Their decision is recorded.
  • Measured performance. Accuracy is measured on the operation's own documents during implementation, by document type, rather than assumed from a general benchmark.

When evaluating any approach, ask to see what happens to a value it cannot read. The answer should be that it is flagged with its source, never silently filled in.

OCR, intelligent document processing and schema-based extraction

The terms are often used interchangeably. They describe different depths of work.

How the three approaches differ on the same batch record.
ApproachWhat it producesWhere it falls short for GxP records
OCRMachine-readable text from an imageNo structure: a lot number is just characters on a line, with no field or unit
General intelligent document processingFields from common business documents such as invoices and formsTrained on generic layouts; weak on handwritten process tables, corrections and scientific notation
Schema-based extraction for pharmaceutical recordsValues in the operation's own schema, linked to their source, checked and reviewedRequires the schema and rules to be defined for each document type up front

From digitized records to connected context

A digitized batch record is useful on its own. It becomes far more useful when it can be read with everything related to it: the certificates for the materials charged, the lab results for the samples taken, the deviation raised on the third day and the recipe it was meant to follow.

That is the step from digitization to context. In Katalyze, document-derived data joins system data brought in through Connectors in the Context Layer, which organizes it with an ontology specific to the operation and keeps every value linked to its original source. Customer systems remain the source of truth; nothing has to be migrated. Agentic Solutions such as Deviation Investigation and Yield Optimization then work on that context, and the operation's experts review and approve what they produce.

How to evaluate a document digitization approach

Use a short, practical checklist during evaluation. Every item should be demonstrated on the operation's own records.

  1. Run it on a real executed batch record with handwriting, corrections and margin notes, and on certificates from at least three suppliers.
  2. Check that every extracted value links to the exact place it was read.
  3. Ask how uncertain values are identified and what the reviewer sees.
  4. Review the rules: who writes them, how they are versioned and how failures are shown.
  5. Confirm how corrections, strike-throughs and signatures are represented in the output.
  6. Ask how the process supports the operation's validation approach and where its records and audit trail live.
  7. Confirm where the output goes and whether source systems stay the source of truth.
  8. Ask what else the data can be used for once it is structured.

Questions

Is a scanned PDF of a batch record a digitized record?
No. A scan preserves an image of the page, which may serve as a true copy if it is verified, but the values inside it are not data. Digitization extracts those values into a schema and links each one back to the page it came from.
Can digitized copies replace paper originals?
Regulators allow true copies in place of originals when the copy is verified to preserve the content and meaning of the original, including its metadata. FDA's 2018 data integrity guidance and MHRA's 2018 guidance both describe this. Whether to destroy originals is a decision for the operation's own quality system and risk assessment.
Which documents should a team digitize first?
Most teams start with executed batch records or certificates of analysis, because both are high in volume, hard to query and needed for recurring work such as batch review, supplier qualification and investigations. The right first set is the one behind a specific operational question the team needs to answer.
Does digitization require replacing an MES or eBR system?
No. Digitization works on the records that already exist, including records that come from suppliers and contract partners. It complements electronic systems rather than replacing them, and those systems remain the source of truth.
How accurate is automated extraction of handwritten records?
It depends on the document type, the handwriting and the image quality, which is why accuracy should be measured on the operation's own records. A governed process flags any value read with low confidence, or that fails a rule, for expert review instead of passing it on.

Sources

  1. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry, U.S. FDA, December 18, 2018
  2. 21 CFR Part 211, Current Good Manufacturing Practice for Finished Pharmaceuticals (211.84, 211.180, 211.188)., U.S. Code of Federal Regulations. , September 25, 2026
  3. 21 CFR Part 11, Electronic Records; Electronic Signatures, U.S. Code of Federal Regulations., September 25, 2026
  4. Part 11, Electronic Records; Electronic Signatures: Scope and Application. Guidance for Industry., U.S. FDA, August 1, 2003
  5. 'GXP' Data Integrity Guidance and Definitions, Revision 1, MHRA, March 1, 2018
  6. PI 041-1, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments., PIC/S, July 1, 2021
  7. TRS 1033, Annex 4: Guideline on data integrity, WHO, October 10, 2021

In this guide

  1. True copies, scanned batch records, and data integrityA scanned batch record can stand in for the paper original only if it is a true copy: a copy verified, by a dated signature or a validated process, to preserve the original's full content and meaning, including the metadata needed to reconstruct the activity. FDA's 2018 data integrity guidance and MHRA's 2018 guidance both set this expectation, and 21 CFR 211.180 allows GMP records to be kept as true copies. A true copy is still an image, though; digitizing its contents into data is a separate step with its own integrity requirements.
  2. How to digitize certificates of analysis from suppliersDigitizing a supplier certificate of analysis (CoA) means extracting the material, lot, tests, specifications, results, and methods into one schema, with each value linked to the certificate it came from, whatever layout the supplier used. The steps are to define one schema for every supplier, map each supplier's terms to it, extract, check the results against the specification, and review anything flagged. Structured certificate data makes it practical to trend a supplier's results across lots and to support the verification of supplier data that 21 CFR 211.84 requires.
  3. How to digitize paper batch records without losing the sourceDigitizing a paper batch record means reading every value on the executed record, including handwritten entries, corrections and signatures, and loading into a defined schema, with each value linked to the location it came from. The steps are to define the schema, capture a good image, extract, check against rules, have an expert review flags, and deliver the approved data. Keeping the link to the source at every step is what makes the result usable in regulated work.
  • Tube caps under Ink
    September 23, 2026What FDA's draft AI guidance asks for, and what it leaves out

    FDA's draft guidance Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, issued in January 2025 and still a draft as of September 2026, sets out a seven-step, risk-based framework for establishing that an AI model is credible for a defined context of use. It applies to AI that produces information supporting regulatory decisions about a drug's safety, effectiveness, or quality. It excludes AI used in drug discovery and AI used for operational efficiencies that do not affect patient safety.

  • A floral macro in blue and orange
    September 21, 2026Why unstructured records slow regulated operations

    In pharmaceutical operations, much of the evidence behind a batch sits in unstructured records: paper batch records, supplier certificates of analysis, lab reports and deviation records that no system can query. The records are complete and controlled, but answering a question across them means finding, reading, and retyping them by hand. That slows investigations, supplier qualification, and batch review, and it keeps experts on assembly work. Turning those records into structured, source-linked data removes that step without replacing the systems that hold them.

Alyse Gonthier, PhDHead of Content

PhD in biomaterials; Science communication enthusiast

LinkedIn
Where this goes next

Document Intelligence

Turns records from suppliers, CROs, CDMOs and your labs into contextualized, source-linked data.

See How It Works
Highly Regulated

A weekly letter from Katalyze.

Book a demo

See Katalyze on your operation.

Bring the job you want to move forward and the records you'd like connected. Thirty minutes on how Katalyze would do the work, and where your experts come in.

Or, just a quick chat to run through the product, your call.

Pharmaceutical document digitization guide | Katalyze