# Why unstructured records slow regulated operations

By Alyse Gonthier, PhD · Published September 21, 2026

Canonical: https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations

## In brief

In pharmaceutical operations, much of the evidence behind a batch sits in unstructured records: paper batch records, supplier certificates of analysis, lab reports and deviation records that no system can query. The records are complete and controlled, but answering a question across them means finding, reading, and retyping them by hand. That slows investigations, supplier qualification, and batch review, and it keeps experts on assembly work. Turning those records into structured, source-linked data removes that step without replacing the systems that hold them.

## Key takeaways

- An unstructured record holds its information in pages or free text rather than in fields a system can query, which is the norm for paper batch records, supplier certificates, and many lab reports.
- The cost of unstructured records shows up as time: every cross-record question, such as which batches used a given material lot, becomes a manual search.
- Scanning a record does not make it structured; its values have to be extracted into a schema and linked to their source.
- FDA's 2018 data integrity guidance expects data to be attributable, legible, contemporaneous, original or a true copy, and accurate (ALCOA), so extracted data has to keep its link to the original.
- Structured, source-linked records can be read together with data from existing systems without migrating anything, and those systems remain the source of truth.

## Contents

- [What counts as an unstructured record](https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations#what-counts-as-an-unstructured-record)
- [Where the time goes](https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations#where-the-time-goes)
- [Why scanning and OCR are not enough](https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations#why-scanning-and-ocr-are-not-enough)
- [What regulators expect of the result](https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations#what-regulators-expect-of-the-result)
- [What changes when records become data](https://www.katalyzeai.com/blog/why-unstructured-records-slow-regulated-operations#what-changes-when-records-become-data)

## **What counts as an unstructured record**

A structured record keeps its information in defined fields: a LIMS result, an ERP inventory entry, a row in a database. An unstructured record keeps it on a page. In pharmaceutical operations that includes:

- Executed batch records completed by hand on printed forms
- Certificates of analysis from suppliers, in a different layout from each one
- Lab reports that combine results tables, traces and signatures
- Deviation records written mostly as free text
- SOPs, recipes and tech transfer documents with nested tables
- Records sent by contract manufacturers and labs as PDFs or scans

These records are not a failure of digitization. Many are controlled and signed, and some come from outside the organization. They’re simply not data yet.

## **Where the time goes**

The cost of unstructured records is rarely a line item. It shows up as the hours between a question and its answer.

| Question | What Answering It Takes Today |
| --- | --- |
| Which batches used raw material lot 2231-B? | Search batch records for the lot number, page by page |
| Has this supplier's reported assay drifted over the past two years? | Retype results from each certificate into a spreadsheet |
| What happened on batch BR-0412 before the deviation? | Pull the executed record, read it and cross-check the lab results |
| Did the same issue occur at the other site? | Request records from another team and repeat the search |

*Common questions and the manual work unstructured records turn them into. (Lot and batch numbers are illustrative only).*

Each question is reasonable and recurring. Each one turns an expert into a data scavenger for hours or days.

## **Why scanning and OCR are not enough**

Scanning makes a record easier to store and find. It does not make its contents usable. Optical character recognition turns the image into text, but a string of characters is not a lot number in a field with a unit and a link to where it was read. OCR also struggles with what matters most in GxP records: handwriting, corrections, tables that span pages, scientific characters, and margin notes.

What turns a record into data is extraction into a schema, checks against rules and review of anything uncertain, with every value keeping its link to the page. 

## **What regulators expect of the result**

Structured data drawn from GxP records has to meet the same integrity expectations as the records themselves. FDA's 2018 data integrity guidance expects data to be attributable, legible, contemporaneously recorded, original or a true copy, and accurate (ALCOA). It accepts electronic copies of paper records as true copies only when they preserve the content and meaning of the original, including the metadata needed to reconstruct the activity.

In practice that means three things for any extraction process:

- Every value keeps its link to the page, table or field it came from.
- Uncertain values are flagged for review rather than filled in.
- The review and approval are recorded.

## **What changes when records become data**

Document Intelligence reads the text, tables, handwritten fields, and margin notes in these records, organizes them into the operation's schema with every value linked to its page, runs defined checks, and routes anything flagged to the operation's experts for review. Approved data can go to a configured destination or join the Context Layer, where it is read alongside data brought in through Connectors from systems such as SharePoint, Veeva Vault, and SAP. Nothing is migrated, and those systems remain the source of truth.

The questions in the table above become queries. The answer to "which batches used lot 2231-B" links to every page where the lot appears. A supplier's certificate results can be trended across lots. The evidence for an investigation can be assembled from records that used to sit in an archive.

The larger change is where expert time goes. When finding and reconciling records takes less of the day, scientists and quality teams can spend more of it deciding what the evidence means.

## Questions

### What is an unstructured record in pharma?

It is a record whose information sits on a page or in free text rather than in fields a system can query. Paper batch records, supplier certificates of analysis, lab reports and deviation narratives are common examples.

### Is a scanned PDF structured data?

No. A scan is an image of the record. Its values become structured data only when they are extracted into a schema, checked and linked back to where they were read.

### Do we have to replace existing systems to use unstructured records?

No. Extracted data can be connected with data from existing systems without migrating it, and those systems remain the source of truth.

## Sources

- [Data Integrity and Compliance With Drug CGMP: Questions and Answers.](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/data-integrity-and-compliance-drug-cgmp-questions-and-answers), U.S. FDA, December 1, 2018
- [21 CFR Part 211 (211.180, 211.188)](https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211), U.S. Code of Federal Regulations, September 25, 2026

## Read next

- [How to digitize paper batch records without losing the source](https://www.katalyzeai.com/blog/how-to-digitize-paper-batch-records-without-losing-the-source): Digitizing a paper batch record means reading every value on the executed record, including handwritten entries, corrections and signatures, and loading into a defined schema, with each value linked to the location it came from. The steps are to define the schema, capture a good image, extract, check against rules, have an expert review flags, and deliver the approved data. Keeping the link to the source at every step is what makes the result usable in regulated work.
- [The complete guide to pharmaceutical document digitization](https://www.katalyzeai.com/guides/the-complete-guide-to-pharmaceutical-document-digitization): Pharmaceutical document digitization is the process of reading GxP records such as executed batch records and certificates of analysis into a defined schema, with every value linked back to the page it came from. A scan by itself is a picture, not data. Done well, digitization produces structured data that reviewers can check against the source, that meets the regulatory expectation for accurate and complete records, and that other work can build on. This guide covers how digitization works, what OCR is and how to evaluate its accuracy, and how to build trust with digitization technology. 

[See How It Works](https://www.katalyzeai.com/product/document-intelligence)
