Guide
Agentic AI in GxP environments: A practical guide

In brief
An AI agent in a GxP environment that carries out a defined operational task, such as assembling the evidence for a deviation investigation, within a governed workflow that limits its sources and tools, records every run and leaves the decision with a qualified person. Regulators are setting expectations for this type of work. The EU GMP Annex 22 draft (July 2025) keeps generative AI out of critical GMP applications but allows it elsewhere with a qualified person checking the output, so the practical question is how the work is controlled, not whether AI is present.
Key takeaways
- An AI agent differs from a chat assistant because it performs a defined task end to end, such as gathering records, applying rules and drafting a report, rather than answering one question at a time.
- FDA's January 2025 draft guidance on AI in regulatory decision-making excludes AI used for internal operational efficiencies that do not affect patient safety, so it is not the governing framework for most operational agents.
- The EU GMP Annex 22 draft, published for consultation on July 7, 2025 and still a draft as of September 2026, says generative AI and large language models should not be used in critical GMP applications, and requires a qualified person to check outputs in non-critical ones.
- Software cannot be "GxP compliant" on its own; compliance belongs to the regulated process, and what a vendor can offer is software built for GxP environments that supports validated processes.
- The controls a quality team should expect are approved sources and tools, source lineage, a recorded execution history, surfaced uncertainty, access control, and expert approval as a gated step.
What an AI agent is, in operational terms
An AI agent is a system that works toward a defined goal by taking a series of steps: finding information, calling tools, checking results, and producing an output. In pharmaceutical operations, the useful version of that idea is narrow. The agent does a specific job the operation has defined, against sources the operation has approved, and hands its work to a person who decides.
Consider a deviation raised on a batch. Before anyone can investigate the cause, someone has to assemble the batch record, the in-process results, the raw material lots and their certificates, the equipment logs, and any related deviations. An agent can do that assembly, compare the batch with similar batches, check the evidence against the relevant procedures, and draft a summary of what it found and what it could not find. The investigator reviews the evidence, tests the explanations, and determines the cause.
That division of labor is the whole idea: agents take on the assembly work, and experts keep the judgment.
Agents, copilots, and chat assistants
| Chat assistant | Copilot | Governed agent | |
|---|---|---|---|
| What it does | Answers a question | Suggests the next action inside one application | Performs a defined task across several sources and steps |
| Sources | Whatever it was trained on, plus what the user pastes in | The data in the host application | Only the sources approved for the task |
| Output | Text in a conversation | A suggestion the user accepts or rejects | A draft record, analysis, or plan with evidence linked to its sources |
| Record of the work | A chat history, if kept | Usually the host application's logs | An execution record of every run |
| Who decides | The user, informally | The user, one suggestion at a time | A qualified reviewer, as a defined approval step |
A chat assistant is useful for drafting and exploring. It is not a controlled process. The governed agent is the version that belongs near GMP work, because every element of its work can be inspected.
What regulators have said so far
No regulator has published a final rule specific to AI agents in GMP manufacturing. Several documents set the direction, and a quality organization should know what each one covers.
| Document | Status (Sept. 2026) | What it says the matters here |
|---|---|---|
| EU GMP Annex 22, Artificial Intelligence | Draft; consultation July 7 to October 7, 2025 | Covers static models in critical GMP applications. Generative AI and LLMs should not be used in critical applications. In non-critical uses, a qualified person must check outputs. |
| Revised EU GMP Annex 11 and Chapter 4 | Drafts, same consultation | Strengthen computerised system lifecycle management, quality risk management, and data integrity controls for paper, digital, and hybrid records |
| FDA draft guidance on AI to support regulatory decision-making | Draft, January 2025 | A seven-step credibility framework built on context of use and model risk. Excludes AI used for internal operational efficiencies that do not affect patient safety. |
| EMA reflection paper on AI in the medicinal product lifecycle | Final, September 2024 | A risk-based approach across the lifecycle, from discovery to post-authorization. |
| EMA and FDA Guiding Principles of Good AI Practice in Drug Development | Published January 14, 2026 | Ten shared principles, explicitly including manufacturing. |
| ISPE GAMP Guide: Artificial Intelligence | Published July 2025 | An industry framework for developing and using AI-enabled computerized systems in GxP environments. |
Two points are easy to miss. First, the Annex 22 draft does not forbid generative AI in GMP; it keeps it out of critical applications, the ones with a direct impact on patient safety, product quality, or data integrity, and puts a qualified person in charge of checking it elsewhere. Second, FDA's draft guidance is about AI that produces information supporting a regulatory decision, such as a submission. Its scope excludes AI used for operational efficiencies like internal workflows. That does not mean operational AI is unregulated; it means the existing GMP framework, including the predicate rules and Part 11, is where the expectations come from.
The post in this guide, What FDA's draft AI guidance asks for, and what it leaves out, covers the FDA draft in detail.
Why "GxP compliant AI" is the wrong question
Vendors often describe software as "GxP compliant." The claim does not hold up, and quality teams know it. Compliance is a property of a regulated process and the organization running it. A tool can make that process easier to control and to demonstrate, but it cannot make the process compliant by itself.
An analogy helps. A restaurant earns its health grade as a whole. A supplier can sell the restaurant excellent knives that are easy to clean, and that helps, but the knives are not graded; the kitchen is. Software in a GxP environment is the same. The accurate claims are that it is built for GxP environments, that it supports validated processes, and that it gives the operation the records it needs to show how work was done.
Where an individual workflow is validated for a specific intended use, that validation belongs to the operation and that use. It does not transfer to a platform as a whole.
The controls a quality team should expect
Whatever the regulatory status of a given document, the controls that make agentic work inspectable are consistent. A quality organization evaluating any agentic system should expect each of these.
- Approved sources and tools. The agent is constrained to the data and tools approved for the task and nothing else.
- Source lineage. Every result keeps its path back to the original record, table, field or note.
- Execution records. Each run is recorded with an owner: the sources used, the evidence assembled, the uncertainty surfaced, the draft prepared, the review, and the approval.
- Surfaced uncertainty. Missing or conflicting information is shown rather than papered over, so an unsupported conclusion does not pass silently.
- Access control. Who can run, change, and approve each workflow is defined and enforced.
- Expert approval. A qualified person reviews and approves the result as a defined step, and that approval is recorded.
- Controlled change. Changes to models, prompts, tools, and workflow steps are versioned and assessed before they are used.
These are the controls Agentic Solutions are built around. The customer's systems remain the source of truth, and the execution record is available for the quality team to inspect.
Audit trails and execution records
21 CFR 11.10(e) requires secure, computer-generated, time-stamped audit trails of operator entries and actions that create, modify, or delete electronic records. An audit trail answers "who changed this record, and when."
Agentic work needs a second kind of record. An execution record answers "how was this result produced": which sources the agent was allowed to use, what it retrieved, which rules it applied, what it could not find, what it drafted, and who approved it. The two records work together. The audit trail protects the integrity of the record; the execution record makes the reasoning behind a result inspectable.
| Audit trail | Execution record | |
|---|---|---|
| Question it answers | Who created, changed, or deleted a record, and when | How a result was produced, from which evidence, and who approved it |
| Regulatory anchor | 21 CFR 11.10(e); EU GMP Annex 11 | No single rule; supports ALCOA+ expectations for AI-assisted outputs |
| What it contains | User, timestamp, action, old, and new values | Approved sources, evidence retrieved, rules applied, uncertainty surfaced, draft, review, and approval |
Human in the loop versus expert approval
"Human in the loop" is often used to mean that a person is somewhere nearby. That is not enough in GMP work. A signature on a result the reviewer could not realistically check is not a review.
Expert approval is more specific. The workflow defines the point at which a qualified person must review, what they are shown, what they are approving, and what happens if they reject it. The reviewer sees the draft beside its evidence and the uncertainty the agent surfaced. Their decision is recorded as part of the run. The Annex 22 draft points the same way: it asks that the operator's role and responsibilities be described as part of the intended use.
The value does not stop at approval. When agents take on the assembly, the reviewer's time goes into the questions that need expertise: whether an explanation holds, what the finding means for the process, and what to do next.
Validating and changing an agentic workflow
Validation principles do not change because a workflow uses AI. GAMP 5 Second Edition (2022) emphasizes critical thinking and a risk-based approach scaled to the intended use, and ISPE's AI guide extends that to AI-enabled systems. For an agentic workflow, a risk-based approach usually means:
- Define the intended use and whether the application is critical in the Annex 22 sense.
- Identify the risks: wrong source, missed evidence, incorrect rule, unsupported conclusion.
- Test the workflow on representative cases, including cases where evidence is missing, and confirm the uncertainty is surfaced.
- Confirm the approval step, the access controls and the execution record work as intended.
- Put changes to models, prompts, tools, and steps under change control, with an impact assessment before release.
The Annex 22 draft expects that a retrained model is introduced only as a controlled change. The same discipline applies to a new model version or a revised prompt in an agentic workflow.
Where to start
Most operations start with a job where the assembly work is heavy and the decision is clearly human: digitizing and reviewing records, assembling evidence for investigations, or characterizing raw material variability. These are useful, bounded, and easy to inspect. In an ideal design, each job builds context that the next job can reuse, because every solution runs on the same platform and Context Layer.
Questions
- Can generative AI be used in GMP operations?
- Under the EU GMP Annex 22 draft, generative AI and large language models should not be used in critical GMP applications, meaning those with a direct impact on patient safety, product quality, or data integrity. In non-critical applications they can be used with a qualified person responsible for checking the output. Annex 22 was still a draft as of September 2026.
- Does FDA's January 2025 draft AI guidance apply to AI used inside manufacturing operations?
- Its scope covers AI that produces information to support regulatory decisions about a drug's safety, effectiveness, or quality. It excludes AI used for operational efficiencies, such as internal workflows, that do not affect patient safety. Operational uses remain subject to the existing CGMP framework, including the predicate rules and 21 CFR Part 11.
- Can an AI platform be GxP compliant?
- No software is GxP compliant on its own, because compliance belongs to the regulated process and the organization running it. A platform can be built for GxP environments and support validated processes, and a specific workflow can be validated for a specific intended use by the operation.
- What should a quality team ask an AI vendor first?
- Ask which sources the agent can use and who approves them, how each result links to its evidence, what happens when evidence is missing, what the execution record contains and who can see it, how changes to models and prompts are controlled, and where the approval step sits.
- Does an AI agent replace the investigator or the reviewer?
- No. In a governed workflow the agent assembles evidence, applies defined rules and drafts findings. A qualified person reviews the evidence, decides, and approves, and that approval is recorded.
Sources
- Stakeholders' consultation on EudraLex Volume 4, Chapter 4, Annex 11 and new Annex 22, European Commission, October 7, 2025
- Draft Annex 22: Artificial Intelligence, European Commission, July 1, 2025
- Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products. Draft guidance., U.S. FDA, January 1, 2025
- Reflection paper on the use of Artificial Intelligence in the medicinal product lifecycle, EMA, September 30, 2024
- EMA and FDA set common principles for AI in medicine development, EMA, January 14, 2026
- GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems, Second Edition, ISPE, July 1, 2022
- GAMP Guide: Artificial Intelligence, ISPE, July 1, 2025
- 21 CFR Part 11, Electronic Records; Electronic Signatures (11.10), U.S. Code of Federal Regulations, October 6, 2026
- Artificial Intelligence in Drug Manufacturing. Discussion paper, U.S. FDA, March 1, 2023
In this guide
- What FDA's draft AI guidance asks for, and what it leaves outFDA's draft guidance Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, issued in January 2025 and still a draft as of September 2026, sets out a seven-step, risk-based framework for establishing that an AI model is credible for a defined context of use. It applies to AI that produces information supporting regulatory decisions about a drug's safety, effectiveness, or quality. It excludes AI used in drug discovery and AI used for operational efficiencies that do not affect patient safety.
- EU GMP Annex 22 and the revised Annex 11: what the drafts ask of AI systemsEU GMP Annex 22, published as a draft for consultation on July 7, 2025 alongside revised drafts of Annex 11 and Chapter 4, is the first GMP text written specifically for artificial intelligence. It applies to static AI models used in critical GMP applications, says generative AI and large language models should not be used in those applications, and requires a qualified person to check outputs where AI is used in non-critical ones. The consultation closed on October 7, 2025, and as of September 2026 none of the three documents had been adopted.
Read next
Agentic Solutions
Katalyze performing a defined operational job, within governed workflows.
See the Solutions




