Why unstructured records slow regulated operations

In brief
In pharmaceutical operations, much of the evidence behind a batch sits in unstructured records: paper batch records, supplier certificates of analysis, lab reports and deviation records that no system can query. The records are complete and controlled, but answering a question across them means finding, reading, and retyping them by hand. That slows investigations, supplier qualification, and batch review, and it keeps experts on assembly work. Turning those records into structured, source-linked data removes that step without replacing the systems that hold them.
Key takeaways
- An unstructured record holds its information in pages or free text rather than in fields a system can query, which is the norm for paper batch records, supplier certificates, and many lab reports.
- The cost of unstructured records shows up as time: every cross-record question, such as which batches used a given material lot, becomes a manual search.
- Scanning a record does not make it structured; its values have to be extracted into a schema and linked to their source.
- FDA's 2018 data integrity guidance expects data to be attributable, legible, contemporaneous, original or a true copy, and accurate (ALCOA), so extracted data has to keep its link to the original.
- Structured, source-linked records can be read together with data from existing systems without migrating anything, and those systems remain the source of truth.
What counts as an unstructured record
A structured record keeps its information in defined fields: a LIMS result, an ERP inventory entry, a row in a database. An unstructured record keeps it on a page. In pharmaceutical operations that includes:
- Executed batch records completed by hand on printed forms
- Certificates of analysis from suppliers, in a different layout from each one
- Lab reports that combine results tables, traces and signatures
- Deviation records written mostly as free text
- SOPs, recipes and tech transfer documents with nested tables
- Records sent by contract manufacturers and labs as PDFs or scans
These records are not a failure of digitization. Many are controlled and signed, and some come from outside the organization. They’re simply not data yet.
Where the time goes
The cost of unstructured records is rarely a line item. It shows up as the hours between a question and its answer.
| Question | What Answering It Takes Today |
|---|---|
| Which batches used raw material lot 2231-B? | Search batch records for the lot number, page by page |
| Has this supplier's reported assay drifted over the past two years? | Retype results from each certificate into a spreadsheet |
| What happened on batch BR-0412 before the deviation? | Pull the executed record, read it and cross-check the lab results |
| Did the same issue occur at the other site? | Request records from another team and repeat the search |
Each question is reasonable and recurring. Each one turns an expert into a data scavenger for hours or days.
Why scanning and OCR are not enough
Scanning makes a record easier to store and find. It does not make its contents usable. Optical character recognition turns the image into text, but a string of characters is not a lot number in a field with a unit and a link to where it was read. OCR also struggles with what matters most in GxP records: handwriting, corrections, tables that span pages, scientific characters, and margin notes.
What turns a record into data is extraction into a schema, checks against rules and review of anything uncertain, with every value keeping its link to the page.
What regulators expect of the result
Structured data drawn from GxP records has to meet the same integrity expectations as the records themselves. FDA's 2018 data integrity guidance expects data to be attributable, legible, contemporaneously recorded, original or a true copy, and accurate (ALCOA). It accepts electronic copies of paper records as true copies only when they preserve the content and meaning of the original, including the metadata needed to reconstruct the activity.
In practice that means three things for any extraction process:
- Every value keeps its link to the page, table or field it came from.
- Uncertain values are flagged for review rather than filled in.
- The review and approval are recorded.
What changes when records become data
Document Intelligence reads the text, tables, handwritten fields, and margin notes in these records, organizes them into the operation's schema with every value linked to its page, runs defined checks, and routes anything flagged to the operation's experts for review. Approved data can go to a configured destination or join the Context Layer, where it is read alongside data brought in through Connectors from systems such as SharePoint, Veeva Vault, and SAP. Nothing is migrated, and those systems remain the source of truth.
The questions in the table above become queries. The answer to "which batches used lot 2231-B" links to every page where the lot appears. A supplier's certificate results can be trended across lots. The evidence for an investigation can be assembled from records that used to sit in an archive.
The larger change is where expert time goes. When finding and reconciling records takes less of the day, scientists and quality teams can spend more of it deciding what the evidence means.
Questions
- What is an unstructured record in pharma?
- It is a record whose information sits on a page or in free text rather than in fields a system can query. Paper batch records, supplier certificates of analysis, lab reports and deviation narratives are common examples.
- Is a scanned PDF structured data?
- No. A scan is an image of the record. Its values become structured data only when they are extracted into a schema, checked and linked back to where they were read.
- Do we have to replace existing systems to use unstructured records?
- No. Extracted data can be connected with data from existing systems without migrating it, and those systems remain the source of truth.
Sources
- Data Integrity and Compliance With Drug CGMP: Questions and Answers., U.S. FDA, December 1, 2018
- 21 CFR Part 211 (211.180, 211.188), U.S. Code of Federal Regulations, September 25, 2026
Read next
Document Intelligence
Turns records from suppliers, CROs, CDMOs and your labs into contextualized, source-linked data.
See How It Works



