Medical device manufacturers working with machine learning finally have an ISO-level companion to ISO 14971
- Date
- July 8, 2026
- Category
Other Regulations
- Description
ISO/TS 24971-2:2026, published in June, is the first Technical Specification to translate ML-specific risks into the vocabulary of medical device risk management. It builds on AAMI TIR 34971:2023, which for three years was the closest thing the field had to consensus guidance, and lifts that content to ISO reach with meaningful additions along the way.
It's a document worth reading carefully.
What the TS does well
The integration with ISO 14971 is clean. The TS does not introduce a new process, nor does it expand ISO 14971's requirements. It layers MLMD-specific considerations onto each clause in order. This is the right architectural choice: it makes the guidance immediately usable inside existing risk management systems rather than forcing a parallel structure. Manufacturers can adopt it without restructuring their QMS.
Annex B is genuinely useful. The expanded table of hazards, foreseeable sequences of events, hazardous situations, harms, and potential risk controls (15 rows covering data quality, overfitting, bias, data storage, overtrust, algorithm model issues, functionality, and diagnostic information) is the most practically valuable content in the document. It's the kind of reference material that can serve as a benchmark for hazard analyses, and it's substantially expanded from the equivalent table in TIR 34971.
Autonomy and post-production get real treatment. Annex D introduces a six-level autonomy scale (from Level 0 Manual Support to Level 5 Full Autonomy with continuous learning) alongside the IEC/TR 60601-4-1 functional decomposition. Clause 10 draws explicit distinctions between concept improvement, retraining, and continuous learning as separate post-production options — each with different risk management implications. These are areas where MLMD risk management actually gets hard, and the TS engages with them rather than gesturing past them.
Where it stops short
The LLM and generative AI exclusion is conceptually leaky. Clause 1 states that the document does not apply to MLMDs employing large language models or generative AI. On strict reading, this is self-contradictory: all LLMs and generative AI models are ML models trained by ML algorithms on training data, which is exactly the definition the TS itself uses. The boundary blurs immediately for hybrid systems (foundation-model classifiers, GAN-based image denoisers, LLMs) used with constrained outputs as classifiers. Manufacturers with such devices will need to draw the line themselves and justify it in their risk files.
Guidance is at the hazard-identification level, not the quantitative parameterization level. The TS tells you what to worry about but rarely how much or how to measure. Sample size determination for target patient populations, annotation reliability metrics, feature-selection risk, drift detection thresholds are all named, but none are developed. This is arguably appropriate for a Technical Specification, but it means the risk file will need to import external rigor that the TS does not demand itself.
The document assumes a certain kind of ML. Discriminative models with bounded output spaces, definable ground truth, testable acceptance criteria, and limited input variability. Devices that fit this archetype (the majority of currently marketed MLMDs) will find the guidance directly applicable. Devices that don't will find it useful but incomplete.
What it means for manufacturers
Don't treat it as a checklist. The TS is a serious document that expects work behind each clause. Reading it as a compliance exercise will produce a risk file that reads like one and won't stand up to reviewer scrutiny. The Annex C question catalog in particular is explicitly not a checklist, and the authors say so.
Update your risk management plan around post-production. Clause 4.4 and Clause 10 are where the TS is most operationally demanding. Retraining triggers, rollback criteria, drift monitoring cadence, versioning of MLMD, ML model, training data and test data are all now expected to be defined in the RM plan for devices where post-production performance monitoring is relevant to safety. If your current RM plan is silent on these, that is the first gap to close.
Read the TS as part of an ecosystem. The document explicitly builds on IMDRF N67, N81, and N88, the FDA / Health Canada / MHRA guiding principles on GMLP and transparency, and Health Canada's 2025 pre-market guidance for ML-enabled medical devices. Reading the TS in isolation will miss half the picture.
A balanced close
Read as a reference document, the TS delivers. Read as an instruction manual, it will frustrate. The gap between the two is where the real risk management work lives, and it is larger than the document lets on. That is not a failure of the authors, but it does mean that a well-formatted risk file citing the TS at every clause is not the same as a risk file that would survive serious scrutiny.
We will be publishing a more detailed analysis in the coming weeks, focusing on the areas where the TS points in the right direction but leaves the operational rigor to the manufacturer. If you are currently building or remediating an MLMD risk file, those are likely the areas that will consume most of your time.