Sommaire
- Digitalisation: Considerations for Data Integrity & Data Review
- Intelligence artificielle en environnement GMP : où en est-on ?
- Biopharma Industry Readiness for AI Implementation. Key Questions & Answers
- Écarts majeurs en inspection visuelle : enseignements des inspections ANSM 2023-2025
- Use of a Justified Sample Size and Criteria for the Validation of Visual Inspection Processes
- Factoring Sustainability Procurement decisions: Environmental Aspects
- Hidden costs in Environmental Monitoring: Why Quality is more expensive than you think?
- Pistes d’optimisation de l’inspection visuelle manuelle, au delà de la base réglementaire
- Dépasser la barrière du financement de la transition environnementale de la production pharma.
Biopharma Industry Readiness for AI Implementation. Key Questions & Answers.
Artificial intelligence is working its way into every conversation and every conference. Pilot projects multiply, vendors reframe their offerings, and industry guides pile up from every direction. At the same time, a new regulatory landscape is taking shape from FDA to EMA, through the AI Act.
The noise is real. But behind the hype, how ready is our industry, really? How do we cut through the documentation flood and tell principle from practice? And most importantly, how will AI genuinely fit into the validation, quality, and manufacturing practices we already know?

This Q&A offers a few points to reflect on, based on a discussion with two Data Scientists, working full time in the biopharmaceutical industry to better understand what AI really brings, and where its limits still are in real-world contexts.
Q1. What does AI adoption look like across the biopharmaceutical manufacturing sector ?
The biopharmaceutical is really eager to adopt AI across its entire value-creation chain. But when you look at where investment actually goes, most investments go into administrative, supply chain, IT, R&D and sales functions. Manufacturing and GxP activities, while growing and accelerating, still face significant challenges and remain largely in an exploratory phase.
The OECD data illustrates this well. According to the latest OECD report 2026(1), 26% of EU pharmaceutical companies declared using at least one AI technology in manufacturing in 2024. That is more than double the manufacturing average of 11%, up from 7% in 2021. And 7.6% of pharma companies are using AI to optimise production processes, making it the second highest rate across all sectors, just behind oil and gas at 10.7%.
Within that GxP perimeter, adoption is most mature where AI’s pattern-recognition strengths meet well-defined, high-volume operational problems: predictive maintenance, PAT-driven in-process control, and computer-vision quality inspection. An active frontier is bioprocess development, where AI accelerates experimental navigation, and where generative tools are starting to support knowledge management tasks (quality knowledge bases, deviation investigations, …etc). Core GxP activities such as batch release, by contrast, remain human-gated, and rightly so. Applications touching critical quality attributes demand a level of explainability and validation rigour for which the industry is still building the frameworks
The next phase will not be decided by technology itself, but by three more practical factors: data infrastructure, because many sites still operate on fragmented, siloed environments; talent, because deploying AI in a GMP context requires a rare mix of data science, process knowledge and regulatory literacy; and culture, because the shift toward risk-proportionate, outcome-focused oversight asks Quality teams to rethink their own role.
Q2. Now that the landscape is mapped: what does CSV look like in the context of AI?
Obviously, AI is implemented through pieces of software, and is thus subject to CSV. An important starting point is that AI-CSV builds on the existing CSV framework rather than replacing it. The arrival of AI does not make traditional validation approaches obsolete; it makes them more important than before.
AI systems are almost always embedded in conventional software, and biopharma is no exception. They usually operate as one component embedded within broader IT ecosystems (MES, LIMS, QMS, HMI, …etc), each of which keeps its own validation requirements, unchanged. So the task is not to reinvent CSV. It is to extend it to account for what AI brings with it!
These adaptations play out at three levels, summarized in Figure 1. They will also structure the rest of our discussion: technical first, then organisational, then human.
The first level is technical, and it revolves around data and models. A model’s behaviour is directly shaped by the quality, representativeness and completeness of the data it is trained or run on. Data is not a secondary product; it becomes a first-class compliance artefact, and Data Governance shifts from best practice to a direct component of computerised system validation (see A3P GST n°22(2))
The second level is organisational. AI introduces new categories of records that need to live inside the quality system: versioned models, datasets and prompts, and in higher-risk cases, explainability reports, …etc. This cascades into SOPs, supplier management and change control.
The third level is human, and it comes down to critical thinking. AI-compliance can no longer be reduced to producing documentation and executing test scripts: it demands rigorous risk assessment, a clear evaluation of fitness for intended use, and human oversight mechanisms that are explicitly designed in terms of robustness. Practitioners and users must apply real judgment rather than blindly rely on model outputs. This is exactly what FDA’s Computer Software Assurance (CSA) framework(3) has put at the heart of validation — and rather than pulling regulated industries away from CSA, AI makes its principles indispensable.
Q3. Before we discuss the pitfalls: could it be that AI and pharma validation are actually more aligned than we assume?
For Machine Learning (ML) experts, it may seem surprising that ML and GxP would be difficult to reconcile — and that ML validation in particular is seen as difficult.
Why is validation perceived as difficult? After all, ML fundamentally focuses on performance specification, bias avoidance, and objectively measured generalisation capabilities. An ML expert does not want a model that performs on the data at hand, but one that can be proven to generalise to future, unseen data: isn’t it exactly the purpose of system validation in GxP environments? What ML expert calls validation of generalization capabilities is very close to what quality managers would call fitness for intended use.
And thus, in that sense, AI/ML and GxP are close in spirit, as reflected in several guidance documents. Building an ML model requires to specfying business-relevant metrics, not in abstract statistical terms, and associated performance targets. It is a validation procedure close to, for example qualifying a piece of equipment or assessing operator competencies. The way they are implemented may differ, but they rely on the same fundamental principles. Think for example of the building of a reference bank of vials with defects, that operators must be able to identify with a certain level of sensitivity and specificity, without having had access to this specific bank during training.
Planning, executing and documenting a validation strategy for models is the main purpose of the EMA Annex 22 draft(4), but are also present in slightly older documents such as PAT guidance, and corresponding ICH Q2 (R2)(5) and Q14(6). PAT has already been applied to validate ML models, just as one would validate a more traditional analytical method.
The question becomes however more difficult when it comes to other parts of AI such as generative AI.
Q4. So where does generative AI break this alignment, and what does draft Annex 22 actually say about it?
Gen AI produces rich, free-form artefacts (documents, images, analyses). For that reason, the validation of Gen AI systems is more challenging. Take the vial defect detection example again. A decision-support system outputs a prediction (binary or categorical), which can be assessed against a curated ‘ground truth’ using metrics such as accuracy, recall or precision. Target performances can be defined a priori, quantified objectively. Even when a statistical performance is obtained (e.g. “95+/-2%” accuracy), specifications are relatively straightforward to define, measure, document. The output of a decision-support system is rarely more than a single number, or at most a handful of values (for example metabolite modelling in cell culture).
With Gen AI outputting far more complex artefacts, specifying precisely the quality metrics and the associated target levels becomes a task nearly as difficult as building the generative model itself. How can one specify all the metrics (and target levels) related, for example, to an AI-generated contamination risk analysis?
The draft Annex 22 simply prohibits generative AI, on the grounds that its outputs are most often non-deterministic: given the same input, the system may return a different output each time. This is not systematic, but it remains difficult to avoid with Gen AI, so the concern is legitimate.
Even if Gen AI in GxP environments were limited to deterministic approaches, it would still be very difficult to clearly define the expected output specifications.
EMA has received many comments pushing back against an outright exclusion of Gen AI on principle. The advocates of this may very well succeed. Yet, it remains to be discussed how the industry will validate Gen AI systems and define appropriate usage guardrails. This does not mean, of course, that these approaches have no role to play in manufacturing at all, but the way one can validate these methods in critical GxP contexts is not solved today. Validating more traditional ML system in the spirit of the draft Annex 22 should, on the contrary, pose no fundamental issue.
Q5. These limitations point to a root challenge: what makes data truly “GxP-ready” for AI systems?
A recurring element in all recent pieces of regulation and guidance is the attention that must be paid to data. Data obviously play a key role in AI systems, throughout their complete lifecycle. The A3P Guide n°22(2), the draft GMP Annex 22(4), and even in a broader context the Good Machine Learning Practices (GMLP)(13) position paper of the FDA stresses the need to specify, contextualize, curate, and version the data that will be used to build, then to test and validate, AI systems.
Fortunately, Biopharma companies have access to ever-growing sources of data (sensors, lab tests, operator records,…etc). A modern biotech process can generate tens of thousands of data points per produced lot. And biopharma has a long tradition of producing records, many of which are audit-ready, and in the best cases even ALCOA+ compliant.
However, a record is not the same as data, and this distinction is too often overlooked in the industry. As a result, data from different silos cannot easily be reconciled, making it hard to build datasets, and more broadly FAIR(12) (or VERIDICUS(2)) principles are not implemented.
Consequently, datasets themselves are not properly documented, auditable, or versioned, yet they are the basis for the validation of AI systems.
Q6. From a traceability perspective: what new deliverables must a VSI project involving AI include?
In fact, the familiar validation documents remain the backbone: URS, Functional Specifications, risk analysis, IQ/OQ/PQ, traceability matrix, validation plan and report. What changes is how they are filled out, and what new categories of deliverables emerge alongside them. Several existing documents need to be enriched to deal with AI-specific concerns. The URS, for instance, has to cover new requirements, such as: model performance, acceptance thresholds, explainability expectations, fallback rules.
The risk analysis also needs to incorporate new failure modes that traditional software did not present: data drift, algorithmic bias, hallucinations in the case of generative AI, and supplier dependency when the model comes from a third party.
On top of these evolving documents, three new families of deliverables emerge — covering what traditional CSV was not designed for: data, the model itself, and performance over time.
- Datasets become qualified deliverables: provenance, quality, representativeness and locked versioning must be documented. This can take the form of data sheets, a data governance plan, a data preparation report, and versioning records for training and test sets.
- Models themselves become qualified deliverables. The associated deliverables typically include a model card (a standardised description of the model’s purpose, architecture, development choices and version history), a clear rationale for hyperparameter selection, explainability documentation aligned with the level of GxP criticality, and reports from challenge or shadow-run exercises.
- Monitoring-related documents complete this set. A monitoring plan designed to detect drift, explicit retraining criteria and SOPs, and periodic performance reviews all belong in the validated documentation set.
- Figure 2 show how these deliverables map across the system lifecycle (Concept → Operational), and where AI-specific documents sit alongside traditional CSV ones.
Nota Bene: Not every deliverable listed above is required on every project. In keeping with a risk-based approach, they can be activated case by case, proportionate to the criticality and intended use of the AI system.
Q7. … And from a Quality System perspective: what does the integration of AI fundamentally change?
Once an organisation runs several AI initiatives in parallel, managing them as independent projects is no longer effective. A enterprise-wide governance is required.
Strategically, this begins with an AI Policy signed by senior management. The policy sets the company’s position on AI risks, defines authorised use cases, draws clear boundaries, and establishes how AI governance aligns with the existing Pharmaceutical Quality System (PQS).
Operationally, most organisations end up building a cross-functional AI Governance Committee. Its role is to authorise new use cases, review risk classifications at predefined intervals, track and ensure the execution of action plans, and arbitrate on the future of deployed systems.
The real difficulty is that AI initiatives are not limited to GxP — they cut across the whole company. The challenge is then to find the overlaps between AI governance and the existing PQS, in order to avoid a double governance structure, contradictory procedures and blind spots.
The good news is that ISO/IEC 42001:2023[8], the AI Management System standard, shares a common backbone with ISO 9001. Integration is then a matter of mapping each ISO 42001 requirement onto the corresponding PQS building block, and enriching rather than duplicating.
For example:
- The AI Impact Assessment from ISO 42001, for instance, can be integrated into the existing change control process.
- The AI system inventory can extend the computerised system inventory already in place.
lays out the main QMS building blocks that need to be revisited in this effort — change control, supplier management, risk management, training, and the internal audit and management review cycle, among others.
For example:
- Change control incorporates new triggers, such as model retraining, silent supplier updates, …
- Risk management takes on new impact categories, including data drift, human bias, …
- Supplier management introduces an AI-specific questionnaire covering data governance, data sovereignty, explainability, update frequency, …
Q8. Now that we have the governance structure: how do we preserve effective human accountability when AI is part of the process?
One of the biggest AI challenges is not technological but human: how we relate to it. This is a cognitive bias known as automation bias, i.e. the tendency to accept AI-generated recommendations too quickly and without critical evaluation.
That is why the AI Act’s(7) four modes of human control deserve attention. It forces us to describe precisely what automation we are actually allowing.
– Human-In-Command
The most common and robust mode, typical of a user working with a general-purpose LLM: the human keeps full supervisory authority (scope, prompting, arbitration of outputs), and no individual decision is delegated to the model.This does not mean there is no risk. A recent FDA warning letter (April 2026)(9) highlights issues related to overreliance on AI and inappropriate use. However, in such cases, human responsibility remains full and absolute.
– Human-In-The-Loop (HITL)
HITL is the most often cited mode in generative AI, and sometimes one of the most misunderstood. In this operating model, the AI generates a recommendation, and a human reviews it before any action is taken. It sounds good in theory, but does not work on its own: HITL must not be a cosmetic safeguard. It only works if the reviewer has enough time, the right skills, and the context to challenge the model.
– Human-On-The-Loop (HOTL)
HOTL shifts to oversight through indicators, alerts and sampling rather than case-by-case validation. It suits lower-criticality or high-volume scenarios where reviewing every output is neither practical nor proportionate.
– Full Automation
It is the most sensitive mode and should be reserved for well-bounded, non-critical tasks. It is clearly unsuitable for activities with direct or indirect GxP impact, since it fully delegates responsibility to the AI system.
In conclusion, the choice of mode is not an architectural detail. It is the primary driver of an AI system’s criticality. This is why FDA guidance(10) and the draft Annex 22(4) insist so heavily on the Context of Use: the intended type of decision, the user population, and the conditions under which human override is activated, and it must be revisited whenever any of those parameters change.
In the spirit of ICH Q9(11), the erosion of human accountability must itself be recognised as a new risk to manage. Organisations should be able to produce documented evidence that meaningful review (and not just a signature at the end of a workflow) actually took place.
If we fail to frame this human-machine cohabitation properly, we expose ourselves to a wave of errors committed under the cover of automation. Excessive trust in AI systems may well become the new “pandora box” of the QMS system, the new “human error”.
Q9. Finally, if we expect Quality teams to govern AI systems responsibly, what skills and culture must we actually build?
The 2026 OECD report(1) latest figures show that AI talent concentration in manufacturing remains below 2% in most EU member states in 2025. It is therefore essential to address both upskilling and reskilling of teams. Training is essential, but it would be a mistake to train everyone to the same depth.We identify three layers to consider.
– The first is the quality foundation, and it does not change: risk-based thinking, GAMP 5 lifecycle, ALCOA+ mindset, …etc. AI does not remove these PQS fundamentals, it raises the bar!
– The second layer is about AI literacy(7). This is what most Quality and Validation professionals need: a working vocabulary and a grasp of the key concepts.
What is a training set versus a test set? What does a performance metric actually tell us? What do we mean by data drift, or by explainability? … etc. The goal is not to turn people into data scientists, but to give them enough fluency to perceive the risks and the limits. This layer is common across the organisation, but could be slightly calibrated by role.
– The third layer is most advanced, and it applies to a small core whose day-to-day responsibilities now includes AI: validation engineers, AI Stewards, CSV experts. They require a more advanced mastery of methodologies, including model risk evaluation, statistical validation techniques, and MLOps practices (= good model deployment practices).
In the future, these points will no longer be a NICE-TO-HAVES, but a MUST-DO. Training should not be limited to quality professionals; it must work both ways. Digital and data teams also need to understand GxP operations to stay relevant. This is essentially a process of cross-acculturation.
We would like to conclude with two points. First, technologies and use cases will keep evolving, so training cannot be a one-off. It must be a continuous process, spanning onboarding, day-to-day operations, and major retraining phases triggered by change control.
And we should think about AI the way we think about Quality: not only as a training curriculum, but as a culture built on shared behaviours and values. Top-down training alone will not be enough. Biopharmaceutical companies need to build a genuine AI culture, just as they have built a Quality culture.
References
- 1. OECD (2026), Progress in Implementing the European Union Coordinated Plan on Artificial Intelligence (Volume 2): Uptake in High-Impact Sectors, OECD Publishing, Paris, https://doi.org/10.1787/3ac96d41-en
- 2. A3P Guide, Mastery of computerized systems integrating Artificial Intelligence in a GxP environment, July 2025
- 3. U.S. Food and Drug Administration. (2026). Computer Software Assurance for Production and Quality Management System Software: Guidance for Industry and FDA Staff. https://www.fda.gov/
- 4. European Commission. (2025). EudraLex Volume 4 – Good Manufacturing Practice: Chapter 4, Annex 11 and Annex 22 (Artificial Intelligence) – Draft guideline for public consultation. https://health.ec.europa.eu/
- 5. European Medicines Agency. (2024). ICH Q2(R2) Validation of Analytical Procedures – Scientific Guideline. https://www.ema.europa.eu
- 6. European Medicines Agency. (2024). ICH Q14 Analytical Procedure Development – Scientific Guideline. https://www.ema.europa.eu
- 7. European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). https://eur-lex.europa.eu/
- 8. International Organization for Standardization (ISO). (2023). ISO/IEC 42001:2023 – Information technology — Artificial intelligence — Management system
- 9. WARNING LETTER FDA, MARCS-CMS 722591 — April 02, 2026, https://www.fda.gov/
- 10. U.S. Food and Drug Administration. (2025). Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products – Draft guidance for industry and FDA staff. https://www.fda.gov/
- 11. International Council of Harmonisation. (2023). ICH guideline Q9 (R1) on quality risk management. https://www.ich.org/
- 12. FAIR data, wikipedia. https://fr.wikipedia.org/wiki/Fair_data
- 13 U.S. Food and Drug Administration. (2025). Good Machine Learning Practice for Medical Device Development: Guiding Principles. https://www.fda.gov/
Partager l’article
Arnaud DUIGOU
Thibault HELLEPUTTE







