AI in GMP: What EU Annex 22 and the Revised Annex 11 Mean for Your Models
For the first time, a GMP text is being written specifically for artificial intelligence. It is still a draft, but it already tells you which kinds of models regulators are comfortable with and which they are not.
01Why this is the trend to watch
Visual inspection, predictive maintenance, deviation triage and batch review are all being handed to machine learning models. A model is a computerised system whose behaviour came from data rather than code, and that breaks a lot of assumptions behind the validation approaches covered earlier in this series.
On 7 July 2025 the European Commission and PIC/S published three linked drafts: a revised Annex 11 on computerised systems, a revised Chapter 4 on documentation, and a brand new Annex 22 on artificial intelligence, the first GMP text dedicated to the topic.3 The consultation closed on 7 October 2025 after drawing roughly 1,300 comments, and EMA held a stakeholder workshop on 30 June and 1 July 2026.1,6 A final version is widely expected around the end of 2026, but no adoption date has been published.1,4,5
That timing is the reason to read the drafts now. The data integrity, audit trail and supplier oversight expectations in Annex 11 and Chapter 4 are already shaping inspections, and the Annex 22 text shows how models will be judged once it lands.
ISPE GAMP Guide: Artificial Intelligence
Published in July 2025 as a roughly 290-page guide for developing and using AI-enabled computerised systems in GxP settings, covering data governance, model risk management and change control. ISPE sells it directly, so check its store if the Amazon search does not show a copy.7,8
Find it on Amazon →02The regulatory landscape
| Document | Body | What it adds |
|---|---|---|
| Draft Annex 22 — Artificial Intelligence | European Commission / EMA / PIC/S | First GMP annex on AI, aimed at models used in critical GMP applications1,5 |
| Draft revised Annex 11 — Computerised Systems | European Commission / EMA / PIC/S | Grows from about 5 pages to 19, with detail on security, access management and audit trails3 |
| Draft revised Chapter 4 — Documentation | European Commission / EMA / PIC/S | Codifies ALCOA++ data integrity principles3 |
| ISPE GAMP Guide: Artificial Intelligence (July 2025) | ISPE | Industry framework for the full AI lifecycle, building on GAMP 5 Second Edition7,8 |
| Draft guidance on AI for regulatory decision-making (Jan 2025) | U.S. FDA | Separate U.S. track covering AI used to support regulatory decisions3 |
Annex 22 is meant to sit on top of Annex 11, not replace it: Annex 11 covers computerised systems in general, and Annex 22 adds the model-specific layer.5 The same 2026 revision wave also touches Chapter 1, Annex 15 and other texts, so AI is one part of a broader reset of the EU quality system rules.4
03The AI model lifecycle in GMP
The draft treats a model like any other validated system, with a lifecycle that runs from definition through retirement. Click each stage to expand it.
State exactly what the model decides, on what inputs, and how good it has to be before it is trusted. This is the user requirements step from the computer system validation post, applied to a model.
- Performance targets are set before testing, not after seeing results
- The targets are usually benchmarked against the human or process the model replaces
The model is trained on one data set and tested on a separate, representative one it has never seen. As the draft is generally read, test data must be independent of training data, and results are judged against the pre-set criteria.
- Test data should reflect the real range of inputs, including awkward cases
- Metrics such as sensitivity and specificity are reported, not just overall accuracy
The draft expects monitoring for performance drift, revalidation when a model is retrained or its input data changes materially, and a defined process for retiring a model while keeping its records.1
- Retraining is a change, so it goes through change control
- Model records stay retrievable after retirement, consistent with ALCOA++
GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition)
The 2022 second edition added Appendix D11 on AI and machine learning and is the base the new GAMP AI Guide builds on, so it is the practical starting point for validating any model as a computerised system.7
Find it on Amazon →04What is in scope and what is not
The most discussed feature of the draft is which model types it will accept in critical applications. Switch tabs to compare them.
Static, deterministic machine learning. This is what the draft covers: models whose functionality came from training data, whose parameters are frozen during use, and which give identical outputs for identical inputs.2 A trained image classifier that checks vials for particles, locked after validation, is the textbook example.
Dynamic or continuously learning models. Models that keep adapting in use fall outside what the draft would accept for critical GMP applications, because the validated state could change without anyone deciding it should.2
Generative AI and large language models. The draft excludes these from scope with a plain instruction that they should not be used in critical GMP applications. Industry pushback has led EMA to reconsider the boundary, and the outcome is one of the main things to watch in the final text.2,5
Non-critical uses. The restriction targets critical applications. Using an LLM to draft a training summary or search SOPs, with a qualified person reviewing the output, is a different risk conversation, handled under Annex 11, supplier oversight and your own risk assessment rather than Annex 22 alone.
05Model performance calculator
Acceptance criteria for a classification model usually come from a confusion matrix. Enter your test results to see sensitivity, specificity and accuracy, plus a 95% lower confidence bound, because a small test set can flatter a model.
Confusion matrix evaluator interactive
Positive means "defect present" for an inspection model. Sensitivity = TP ÷ (TP + FN). Specificity = TN ÷ (TN + FP). The lower bound is a Wilson score interval at 95% confidence.
This is a simplified illustration of how acceptance metrics and sample size interact. Real acceptance criteria must be set in advance, justified against the process being replaced, and tested on independent, representative data. Missed defects usually matter more than false rejects, so criteria are rarely symmetric.
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow — Aurélien Géron
A widely used ML primer, and the right preparation for validating models you did not build. It explains train and test splits, confusion matrices, precision and recall, and overfitting, which are the ideas behind the calculator above.
Find it on Amazon →06AI readiness self-check
Readiness checklist
07Where programs will fail inspection
- Treating a vendor's AI feature as someone else's problem. If a model is embedded in software you use for GMP decisions, you still own the validation, and supplier oversight is part of Annex 11.
- Judging a model on overall accuracy. With rare defects, a model can be highly accurate while missing most of them. Sensitivity and specificity, with sample size, tell the real story.
- Letting models change silently. A vendor update or a retraining run is a change, and the validated state has to be reassessed before the new version is used.
- Moving the acceptance criteria after seeing results. This is the same failure described in the method validation and change control posts, and it is just as damaging for models.
08Specimen quality forms
An AI model intended-use and risk classification record, and a model validation and drift monitoring log. They give a starting structure to adapt to your own computer system validation procedure.
Form AI-01 — Model Intended Use & Risk Classification Record
Specimen only — not a controlled document.
| Question | Answer | Comment |
|---|---|---|
| Model type (static / dynamic / generative) | ||
| Used in a critical GMP application? | ||
| Human review of every output? | ||
| Acceptance criteria (metric and target) |
Form AI-02 — Model Validation & Drift Monitoring Log
Specimen only — for periodic performance checks of a live model.
| Review date | Sensitivity | Specificity | Input data change? | Action |
|---|---|---|---|---|
These specimen forms illustrate typical content only. Your quality system's document control procedure takes precedence over this format.
Data Integrity and Data Governance: Practical Implementation in Regulated Environments — R.D. McDowall
The draft Annex 11 and Chapter 4 raise the bar on audit trails, access control and ALCOA++, and those requirements apply to a model's records and training data too. This is the data integrity reference recommended in the computer system validation post.
Find it on Amazon →09References
- Eupry. "EU GMP Annex 22: What does it mean for AI in pharma?" eupry.com
- Pharmaceutical Technology. "Europe Tried to Ban Generative AI From Critical GMP. The Ban May Not Survive. It Does Not Matter." pharmtech.com
- MFLRC. "EU GMP Annex 11 Revision: Computerised Systems, Data Integrity and AI Rules Arriving in 2026." mflrc.com
- MFLRC. "The EU GMP Guide Is Being Rewritten: Every Annex and Chapter Changing Through 2028." mflrc.com
- QMSdesk. "EU GMP Annex 22 artificial intelligence: what the draft asks" (status as of 24 September 2026). qmsdesk.com
- Herrmann, D. "Annex 22 and AI Validation: What the Draft GMP Annex Means for Your Systems." daniel-herrmann.io
- ISPE. "ISPE Announces the Availability of ISPE GAMP® Guide: Artificial Intelligence." July 2025. ispe.org
- ISPE Pharmaceutical Engineering. "New GAMP® Guide Addresses Challenges Posed by AI-Enabled Computerized Systems." September–October 2025. ispe.org
No comments:
Post a Comment