+ Guide · Device software and AI

The next patient, the next update.

Why the software in a medical device can never be tested on every path and what its makers do instead, how an AI model is trained, scored and fitted into a hospital that was not built around it, how a device is defended against attack for years after it ships, and what the FDA, the EU and India ask to see before and after an update.
+ The central question

Software can never be tested on every path, and an AI model learns only from the data its makers chose. What evidence shows either will give the right answer for the next patient, and what keeps that true after an update?

Starts at:everyday familiarity with apps, updates and passwords, and arithmetic with percentages and ratios. No programming, statistics, machine-learning or regulatory background is assumed.Ends at:asking of any software or AI device which patients its evidence covers, which parts of it someone else wrote, what a hospital must check before the first patient, which updates it may receive without a new review, and how it is defended once it ships. Built around:IDx-DR, the first device the FDA authorized to give a screening decision with no clinician reading the image, from a 2017 trial in US primary-care offices to its later clearances; and a second case, insulin pumps recalled in 2019 because a flaw in their wireless link could not be patched.
+ Opening

Four hours of training.

The people holding the camera had signed a statement that they had never taken a picture of the inside of an eye. They worked in family-medicine and primary-care offices, ten of them spread across the United States, and each had been given one standardized training session of four hours, with no refresher afterward. Their task was to photograph the retinas of adults with diabetes who had come in for ordinary care: two pictures of each eye on a Topcon NW400, one centered on the optic disc and one on the fovea, the small pit at the center of sharp vision. About three participants in four needed no drops to widen the pupil. The other quarter were dilated and photographed again.

The pictures went to a program that had been locked before the first participant arrived. Within seconds it returned one of three answers: more than mild diabetic retinopathy detected, not detected, or image quality insufficient. Nobody at the clinic interpreted the photographs. Between January and July 2017, 900 people enrolled.

Every participant who completed the study was photographed a second time, by photographers certified by the Wisconsin Fundus Photograph Reading Center, with a far more demanding protocol: four overlapping stereo pairs of each retina and a scan of the macula by optical coherence tomography. Three experienced readers graded the photographs without seeing the program's answer, and their majority decided which participants truly had the disease; the macular scans were graded separately. Before the trial began, its designers, with input from the FDA, had set the targets it was measured against: sensitivity of 85 percent and specificity of 82.5 percent.

Of the 198 analyzable participants whose reference images showed more than mild disease, the program flagged 173, or 87.4 percent. Of the 621 without it, it cleared 556, or 89.5 percent. Among participants whose reference images could be graded, it gave a usable answer for 96.1 percent. The trial was funded by the developer, IDx, and its first author, the University of Iowa ophthalmologist Michael Abràmoff, had founded the company; the paper discloses the funding and his financial ties to IDx. On 11 April 2018 the FDA granted the system a De Novo authorization and described it as the first device authorized to provide a screening decision "without the need for a clinician to also interpret the image or results."

The trial measured one locked program, run by those operators on that camera, for those patients. It said nothing about the same program in a clinic whose computers send images differently, after the camera is replaced or the model retrained, or when someone sets out to make it fail. A device that will be used for years on people who were never in its trial needs evidence of a different kind, and each later version needs more of it.

+ Before we start

Before we start.

Software in a medical device does what it was written, or trained, to do, on whatever input arrives. A trial like the one above shows how it behaved for 819 analyzable people. A maker, a regulator and a hospital need something more: a reason to believe the next person will get the right answer too, including people the trial never saw, and that the reason will survive the next release. That is the question this guide answers: software can never be tested on every path, and an AI model learns only from the data its makers chose; what evidence shows either will give the right answer for the next patient, and what keeps that true after an update?

Four terms carry the argument. A device software function is software that meets the legal definition of a medical device, whether it runs inside a pump or on a server by itself. An algorithm is the procedure the software follows. An AI model, in this guide, is an algorithm whose decision rules were fitted to example data rather than written line by line. And evidence means records that someone outside the team could check: requirements, test results, study data, monitoring reports.

The answer comes in four parts. Testing samples what software does; it cannot exhaust it, because the paths through even a modest program and the inputs it may meet outnumber any test campaign. Confidence therefore comes from how the software was built: requirements derived from the harm a fault could do and traced to tests, an architecture that keeps the dangerous part small enough to test thoroughly, a full account of the code someone else wrote, and proof that each change broke nothing already verified. Agile teams and automated pipelines can produce that evidence, provided each increment leaves its records. An AI model adds a second limit. Its accuracy is a measurement on a test set, and it holds only for patients like those in that set, scored against a reference someone chose, for the exact model that was locked and fed the inputs it was specified for. The evidence that it will be right for the next patient is a test kept apart from training and drawn from the setting of use, results reported for each group of patients with their uncertainty, a check at each hospital that the specified conditions hold, and monitoring once it is in use. A device can also be made to fail by someone trying, and that threat changes while the code stands still, so a device must be built to be patched and watched for years, starting from a list of what is inside it. What keeps all of this true after an update is change control: every change judged against the claims it could move, and, for AI, a plan agreed with the regulator in advance that says what may change and how the new model must be tested.

One fact shapes every chapter. Software does not wear out. A pump's motor and a camera's sensor age and fail at rates that can be measured; software never does. Every way it can fail was built in on the day it shipped or arrived with an update, waiting for the input that triggers it, and for an AI model the training data is part of what was built.

How to read this guide

The guide runs through 29 chapters in eight parts. Parts I and II explain why software fails differently from hardware and what a software lifecycle does about it, including when much of the code comes from elsewhere and when the team ships every two weeks. Parts III and IV turn to models learned from data: how they are built and scored, and why they do worse on patients unlike those they were tested on. Part V follows a model into a hospital and through its later versions; by its end the central question has its answer. Part VI covers attack and defense, Part VII surveys the products and makers of October 2026, and Part VIII sets out what the FDA, the EU and India require, as of that month. IDx-DR returns throughout as the thread's device, and the recalled insulin pumps return in the security chapters. Every chapter works one example with real numbers, and plates carry their own numbers so that each can be cited alone. These boxes recur:

Insight or solved example

Blue edge. Carries the structural point of a section, or a calculation worked once with real values and stated assumptions.

Caution

Orange edge. Names a common misreading, a trap, or the limit of a claim.

In practice

Green edge. Maps the idea onto the reader's own work: specifying, building or testing device software, evaluating an AI product, deploying it in a hospital, or planning the evidence for a submission.

+ What this chapter established
  • Each chapter closes with what it established, in four lines.

No formula needs more than arithmetic. Proportions such as sensitivity are percentages; uncertainty is given as a range, and the one rule used to size a test, that a run of failure-free trials bounds a failure rate, is worked once with numbers. Counting paths uses powers, written as 10⁹ for a billion. Where a figure is a company's own claim about its product, the text says so, and where a paper's authors work for the company whose product they test, the chapter notes it. The guide uses US spelling.

Colors in the plates follow a fixed key. Requirements, claims and acceptance criteria are gold; software the maker wrote, and a model running as software, is teal, with third-party code hatched in the same teal; data of every kind, from training sets to the images a device receives, is violet; an attacker, an attack path or a vulnerability is red; and a defect, a wrong output or a drift in performance is magenta. The output reported to a clinician or patient is always blue, and one orange mark in each plate points at the detail that matters most. A key strip under every plate lists only the colors that plate uses.

The color key
THE THINGS THE PLATES TELL APART READOUT AND CONSTRUCTION SPECIFICATION requirements, intended use, claims and acceptance criteria CODE software the maker wrote; hatched: third-party code DATA training, tuning, test and field data; inputs THREAT an attacker, an attack path, a vulnerability FAULT defects, wrong outputs, drift, misclassification refer rescreen READOUT the output reported to a clinician or patient, and measured results FOCAL DETAIL the one detail that matters: a single orange mark per plate GRAY hardware, networks, hospital systems; axes and construction DASHED GRAY removed, retired, not tested, or not a device SPECIFICATION CODE READOUT FOCAL DETAIL each plate's key strip lists only the colors it uses
Key — The same color means the same thing in every plate: gold for requirements and claims, teal for software, hatched where a third party wrote it, violet for data, red for an attacker or vulnerability, magenta for faults and wrong outputs, and blue for the output reported to a clinician or patient.
The limits of this guide

This is a reference for understanding, not a development procedure, a regulatory opinion or clinical advice. Products, authorizations, guidances, standards and laws are stated as of October 2026 and will date; several of those named here were drafts or under revision in that month, and the text says which. Where a figure comes from a company about its own product, or from a paper whose authors work for that company, the text says so. Named companies and products are examples of a principle, not recommendations.

+ Part I · Built in on the day it shipped

Why software fails differently.

Between 1992 and 1998 the FDA counted 3,140 recalls of medical devices, and 242 of them, 7.7 percent, were caused by software. Of those 242, 192 were caused by defects introduced when the software was changed after it had first been distributed. Both numbers point to the same peculiarity: software is not a part that ages, so its failures are mistakes made by people, in the first version or in a later one. The three chapters of this part set out what software in a device is trusted to do, why it fails when it does, and why no amount of testing can prove that it will not.

01 — Software in a device

What the code is trusted to decide.

+ The questionWhat does the software in a medical device actually do, and when is the software itself the device?

A camera, a client and a server

The system in the opening scene has three pieces, and only its software is the device the FDA authorized. A commercial retinal camera, the Topcon NW400, takes the photographs. A computer attached to the camera runs a program the maker calls the client, which lets the operator pick two images of each eye and sends them "over a secure internet connection" to a server in a data center. There a second program, the analysis, examines the images and returns one of three answers to the client screen: more than mild diabetic retinopathy detected, not detected, or image quality insufficient. The FDA's decision summary describes this architecture and names the software version reviewed in 2018, IDx-DR 2.0.0.

The camera is a separate product, outside the authorization. What the FDA classified in April 2018 was software, in a new category it called a retinal diagnostic software device, class II, under a regulation written for this authorization. Nothing in that device touches the patient, emits energy or moves. Its effect on health passes entirely through a decision: whether a person with diabetes is referred to an eye specialist now or told to come back in a year.

01 — One camera, software in two places, three answers
THE AUTHORIZED DEVICE: SOFTWARE ONLY (class II, 21 CFR 886.1100) RETINAL CAMERA Topcon NW400 a separate cleared device 2 images per eye: disc- and fovea-centered CLIENT on the camera's computer secure internet connection DATA CENTER SERVICE ANALYSIS images one of three answers operator picks images and sends them shown on the client screen more than mild DR detected: refer not detected: rescreen in 12 months insufficient image quality no clinician reads the image architecture as described in FDA decision summary DEN180001 (2018); software version 2.0.0
SpecificationCodeDataReadoutFocal detail
Plate 01 — The device the FDA authorized in 2018 was software: a client on the camera's computer and an analysis on a server, returning one of three answers. The camera is a separate product, and no clinician interprets the images before the answer is given.

When the software is the device

Most medical devices with software carry it inside: the firmware of an infusion pump, the program that runs a ventilator's valves, the code in an imaging scanner. That software is part of the device and is judged with it. A growing share of devices are software alone, running on ordinary computers, phones or cloud servers. Regulators call these software as a medical device, or in newer documents simply medical device software that stands alone, and IDx-DR's analysis program is an example.

What turns code into a device is not anything in the code. It is the intended use: what the maker says the software is for, in its labeling and claims. A program that draws a graph of blood glucose readings for a person's own interest is not a device in the US; the same graph with a rule that tells the patient how much insulin to take is. The US definition, the EU's rules and India's 2026 guidance draw the boundary differently in detail, and Chapter 26 sets those differences out. They agree on the principle: the claim makes the device, and the claim also fixes what the evidence must show.

A sentence of labeling can make code a device

Two companies can ship identical image-analysis code, one to help researchers count lesions in a study and one to tell clinicians which patients to refer. Only the second makes a medical claim, and only the second needs a medical device's evidence. Changing a sentence of labeling can therefore change the regulatory status of software that has not changed at all.

Inform, drive, diagnose or treat

Once software is a device, the evidence it needs scales with two questions. The first is how serious the patient's situation is: whether a wrong answer could lead to death or irreversible harm, to a serious but treatable deterioration, or to something minor. The second is how much the software's output decides: whether it treats or diagnoses directly, drives the next step of care, or only informs a clinician who still decides. International regulators set out these two axes in a framework for software as a medical device in 2014, and India's 2026 guidance uses the same grid to assign its classes.

IDx-DR sits high on the second axis. Its output is the screening decision itself; no clinician interprets the image. Diabetic retinopathy threatens sight but rarely kills, and a person told to rescreen in a year will usually be seen again, so the situation is serious rather than critical. A stroke-triage program that alerts a specialist to a suspected blocked artery, by contrast, informs a clinician who still reads the scan, but in a situation where minutes count. The two products sit in different cells of the same grid, and Chapter 16 returns to the difference between a program that decides and one that advises.

Solved example: what each function is trusted to decide

Take four functions. A pump's bolus calculator computes an insulin dose from a glucose reading and the carbohydrates entered; it sets the treatment itself, and an error of a few units can cause dangerous hypoglycemia, so it sits at the top of both axes. A triage program that flags suspected large-vessel occlusion on a CT angiogram informs a stroke team in a critical situation: lower on the second axis, highest on the first. A logbook app that stores and displays glucose readings without interpreting them generally falls outside the US device definition, because displaying device data is excluded by statute. A step counter for fitness is a wellness product, not a device. The same phone could run all four; the evidence each needs ranges from none to a clinical trial.

02 — Two questions that set the stakes
HOW MUCH THE SOFTWARE'S OUTPUT DECIDES → INFORMS CLINICAL MANAGEMENT DRIVES CLINICAL MANAGEMENT TREATS OR DIAGNOSES HOW SERIOUS THE PATIENT'S SITUATION IS CRITICAL SERIOUS NON-SERIOUS more evidence → pump bolus calculator stroke triage alert (CT angiogram) IDx-DR screening decision the output is the decision OUTSIDE THE DEVICE DEFINITION glucose logbook that only stores and displays readings fitness step counter grid after the international software framework of 2014, also used in India's 2026 guidance; classes in Chapter 26
SpecificationCodeFocal detail
Plate 02 — The evidence a software device needs rises with two things: how serious the patient's situation is and how much the output decides. IDx-DR's output is the screening decision itself, in a serious but not critical situation; a logbook and a step counter fall outside the device definition.

The claim is the unit of evidence

Because the claim defines the device, it also defines what can go wrong. The FDA's decision summary for IDx-DR lists three risks to health: a false positive leading to an unnecessary referral, a false negative delaying evaluation and treatment, and an operator failing to capture images good enough to analyze. Every control it required answers one of those three: clinical testing, software verification and validation built on a hazard analysis, training and human-factors testing for operators, labeling, and a protocol stating which changes could significantly affect its safety or effectiveness. Chapter 21 shows how that last requirement anticipated, by six years, the plans for changing AI models that the FDA finalized in 2024.

The list is short because the claim is narrow. IDx-DR is indicated for adults with diabetes who have not been diagnosed with retinopathy, used by health care providers, with images from the Topcon NW400. It does not look for glaucoma, and at the time of authorization it was not labeled for patients younger than 22 or for pregnant patients. Each limit removes a population or a condition from what the evidence has to cover. The rest of this guide is about the evidence inside those limits, and about what happens at their edges, where real use tends to wander.

Write the intended use before the first line of code

A useful intended-use statement names the decision the software supports, the patients it applies to, the users, the setting and the inputs it accepts, including the devices that produce them. Draft it before architecture begins, and treat each change to it as a design change: it fixes the class of the device, the hazards to analyze, the population a clinical study must sample, and the edges where use outside the claim is most likely.

+ What this chapter established
  • IDx-DR is a software device: a client and a server-side analysis that turn photographs from a named camera into one of three screening answers.
  • Software becomes a medical device through its intended use, not through anything in its code, so a change of claim can change its status.
  • The evidence a software device needs rises with how serious the patient's situation is and with how much the software's output decides.
  • The claim defines the risks and the controls; narrow indications shrink what the evidence must cover, and real use tends to test the edges.
02 — Faults that do not wear out

Built in on the day it shipped.

+ The questionIf code never wears out, where do its failures come from?

Four conditions at once

In May 2024 Hamilton Medical began correcting the software of its HAMILTON-C6 ventilators, and the FDA classed the action as the most serious type of recall. The fault needed four things to happen together. A clinician pressed the key that briefly raises the oxygen supplied and disconnected the breathing tube to suction the patient's airway. During that disconnection a sensor error occurred, for example because the tubing of the flow sensor was kinked. The ventilator entered its sensor-fail mode. And the patient was reconnected while that mode was still active. In that sequence, and only in it, ventilation might not restart. The FDA's notice lists one injury and one death.

Each of the four events is ordinary on an intensive-care unit. The combination is rare, and it lay inside three versions of the software, 1.1.4 to 1.1.6, in every ventilator that ran them, until version 1.2.3 removed it. No part had worn and no component had drifted. The ventilators that could fail this way were exactly as they had been on the day the software was installed.

03 — Four ordinary events, one failure
software versions 1.1.4 to 1.1.6: the fault is present in every unit, waiting FIXED IN 1.2.3 1 O2 enrichment key pressed; tube disconnected for suctioning 2 sensor error (e.g., kinked flow-sensor tubing) 3 ventilator enters sensor-fail mode 4 patient reconnected while sensor-fail mode is active each one ordinary on an ICU ventilation may not restart ALL FOUR TOGETHER no part wore out TIME FDA recall Z-2020-2024, Class 1; reported: 1 injury, 1 death
CodeFaultFocal detail
Plate 03 — The HAMILTON-C6 fault needed four ordinary events in one sequence. It sat in every ventilator running three software versions until a fourth version removed it: a systematic fault waiting for its input, not a part wearing out.

A fault waits for its input

The FDA's guidance on software validation, issued in 2002 and still in force apart from one section replaced in 2025, puts the difference plainly: "Unlike hardware, software is not a physical entity and does not wear out." A hardware part fails for physical reasons that accumulate: a bearing wears, a capacitor dries, a solder joint cracks under thousands of heating cycles. Its failures are random in the engineering sense: each unit fails at its own moment, and a rate per hour of operation describes the population well.

Software fails only when an input reaches a fault, a flaw in its logic or data, that was there from the start. Engineers call these failures systematic: they occur every time the same conditions recur, in every copy of the same version. The time to failure depends on when the triggering input arrives, which depends on how the device is used, not on how long it has been running. A rarely used function can carry a fault for years before anyone meets it, and a busy hospital can meet it on the first day.

Field years mostly test the common paths

A version that has run in the field for years without a reported failure has shown that the inputs it met so far did not reach a fault. A new hospital, a new workflow or a change in another device can supply the input that the earlier users never did. Field history is useful evidence about the common paths and weak evidence about the rare ones.

Why reliability arithmetic does not transfer

The distinction changes how safety is argued. For hardware, two independent parts that each fail once in many thousand hours rarely fail in the same hour, so redundancy multiplies safety. For software, two copies of the same program fed the same input fail together, so copying the code adds nothing against its own faults.

Solved example: two motors, two copies of the code

Suppose a pump motor fails at random at a rate of 1 in 100,000 per hour, and a second, independent motor is fitted as a backup. The chance that both fail in the same hour is 1 in 100,000 multiplied by 1 in 100,000: 1 in 10,000,000,000. Now suppose the pump's control program has a fault that a particular input triggers once in 100,000 hours of use, and a second processor runs an identical copy as a backup. The input that reaches the first copy reaches the second, so the chance that both fail in the same hour is still 1 in 100,000. Independence, which made the hardware redundancy work, is absent; protection against a software fault has to come from a design that is different or from a check that does not depend on the same code, which is the subject of Chapter 5.

What do such faults look like? A NIST analysis of the 383 software-related device recalls the FDA recorded from 1983 to 1997 could assign a type to 342 of them. Logic faults, a wrong condition or an unhandled case, made up 43 percent; calculation faults made up 24 percent; the rest were spread across data, requirements, timing, interfaces and other types. None is a matter of wear. Each is a decision written into the program that was wrong for some input. The same analysis found that the recalled devices had caused no reported deaths or serious injuries, so the figures describe how software fails, not how often it harms.

04 — Hardware wears out; software changes
HARDWARE PART schematic FAILURE RATE TIME IN SERVICE early failures random failures wear-out SOFTWARE VERSION HISTORY schematic FAILURE RATE (FAULTS MET IN USE) TIME IN SERVICE v1 v2 v3 flat between releases: no wear a change can lower it… …or raise it 79% of software recalls in 1992–1998 followed a change
CodeFaultFocal detail
Plate 04 — A hardware part's failure rate changes with age; software's does not, because nothing wears. It changes only when the code changes, and in the FDA's count most software recalls followed a change made after release.

A change is a new chance to fail

If faults are built in, the moment of building matters, and building does not stop at the first release. In the FDA's count from the 1990s, 192 of the 242 software-related recalls, 79 percent, came from defects introduced by changes made after the software was first distributed. A fix or a new feature reaches parts of the program that were already verified, and anything it disturbs there becomes a new fault.

In March 2024 Tandem Diabetes Care recalled version 2.7 of its t:connect app for iPhone, which works with the company's t:slim X2 insulin pump. The FDA's notice describes the mechanism: the app could crash and be relaunched by the phone's operating system again and again, and the repeated relaunches caused excessive Bluetooth communication with the pump, which drained its battery until it shut down early and insulin delivery stopped. The recall covered 85,863 installed apps, and by mid-April 2024 the company had received 224 reports of injury. The FDA recorded the root cause as a software design change. The pump had not changed at all; a new version of the program that talked to it had.

The two halves of the central question start here. A program that was right for one set of inputs can be wrong for another that arrives later, and a program that was right can become wrong when it is changed. Chapter 3 asks why testing alone cannot close either gap, and Part II describes what the software lifecycle does instead.

Treat a field failure as a fault in every copy

When a software failure is reported, assume the fault is present in every unit running that version and in every version that shares the code. Reconstruct the exact sequence of inputs, write it as a test that fails on the faulty version, and keep that test in the regression suite permanently, so that no later change can bring the fault back unnoticed.

+ What this chapter established
  • Software faults are present from the moment the code is written; a failure occurs when an input reaches one, in every copy of that version.
  • Software failures are systematic, so the time to failure depends on use, and years without failure say little about rare paths.
  • Identical copies of a program fail together, so redundancy that works for hardware does not protect against a software fault.
  • Changes are a major source of new faults: 79 percent of the software recalls the FDA counted in 1992 to 1998 followed changes after release.
03 — Why testing cannot be complete

More paths than seconds.

+ The questionWhy can't a team simply test every case before release?

Thirty decisions, a billion paths

A routine with 30 independent yes-or-no decisions has 1,073,741,824 distinct paths through it, because each decision doubles the number of paths that came before. A device program has thousands of decisions, many of them inside loops that can run any number of times, and each path can be taken with different data values and at different moments relative to other events. The count of distinct behaviors is not large in the way a big inventory is large. It is beyond counting.

Suppose a program had only 100 independent decisions and a test rig could run a billion paths a second. Running every path once would take about 40 trillion years, nearly 3,000 times the age of the universe. Faster computers do not change the conclusion, because each extra decision doubles the work. The FDA's software validation guidance states the consequence without arithmetic: "Except for the simplest of programs, software cannot be exhaustively tested," and "path coverage is generally not achievable."

05 — Each decision doubles the paths
1 decision: 2 paths 2: 4 3: 8 4: 16 16 paths schematic 30 decisions → 1,073,741,824 paths 100 decisions → about 1.3 × 10³⁰ paths at 1,000,000,000 tests a second → about 40 trillion years age of the universe about 13.8 billion years nearly 3,000 times the age of the universe each path can also run with different data and timing
CodeReadoutFocal detail
Plate 05 — Every independent decision doubles the number of paths through a program. With 100 decisions, testing each path once at a billion tests a second would take about 40 trillion years, so complete testing is impossible for any real device.

Presence, not absence

The computer scientist Edsger Dijkstra drew the conclusion that shaped software engineering. In his 1972 Turing Award lecture he said that program testing "can be a very effective way to show the presence of bugs, but is hopelessly inadequate for showing their absence." A test that passes shows that one path, with one set of values, gave the expected answer. It says nothing directly about the paths not taken.

That does not make testing useless. It makes testing a sample, and like any sample its value depends on how it was drawn. A test campaign chosen at random from the space of possible inputs would almost never reach the rare combinations that matter, such as the four simultaneous conditions of the HAMILTON-C6 recall in Chapter 2. A campaign chosen from what the software is required to do, and from the ways it could cause harm, reaches far more of them. The FDA's guidance says testing alone "cannot fully verify that software is complete and correct," which is why Part II describes the rest of the evidence a lifecycle produces.

Coverage counts the lines that ran

A coverage tool reports which statements or branches of the code a test campaign executed. A campaign can execute every statement and still miss a wrong condition, because executing a line with one value does not test it with others. The FDA's guidance calls statement coverage "insufficient to provide confidence" on its own, and even full branch coverage leaves most paths unexercised.

How many failure-free runs prove a rate?

Statistics gives a precise answer to a narrower question: if a program runs many times without failing, how high could its failure rate still be? The rule of three answers it. When no failures are seen in a number of independent trials, the one-sided 95 percent upper confidence bound for the failure probability is about three divided by that number.

Solved example: the cost of a zero

A team runs a function 300 times under realistic conditions and sees no failure. The rule of three says the true failure probability could still be as high as 3 in 300, or 1 percent, at 95 percent confidence. To bring the bound down to 1 in 1,000, the team needs about 3,000 failure-free runs; to 1 in 1,000,000, about 3,000,000. For a rate of 1 failure in 1,000,000,000 hours, a target used for life-critical aviation software, it needs about 3,000,000,000 failure-free hours of realistic operation, which is about 342,000 years. NASA researchers Ricky Butler and George Finelli made the same point in 1993: their table puts the test time needed to show a failure probability of 1 in 1,000,000,000 over a ten-hour mission, on one system, at about 1.1 million years, and even 10,000 copies tested in parallel would need 114 years.

The bound applies only when the test runs resemble real use, so the rare inputs of the field appear in the test at the rate they appear in practice. That is the second weakness of statistical testing for safety: the inputs that cause harm are, by their nature, the ones nobody thought to make common in the test.

06 — What a run of zeros can prove
100 1,000 10,000 1,000,000 10⁹ 10¹⁰ 1 in 10 1 in 100 1 in 1,000 1 in 10⁶ 1 in 10⁹ 1 in 10¹⁰ UPPER 95% BOUND ON FAILURE RATE (LOG SCALE) FAILURE-FREE TRIALS OR HOURS (LOG SCALE) 300 runs → 1 in 100 3,000 → 1 in 1,000 3,000,000 → 1 in 1,000,000 3 × 10⁹ hours → 1 in 10⁹ per hour about 342,000 years of operation unreachable by testing alone Butler & Finelli 1993: about 1.1 million years for 1 in 10⁹ per 10-hour mission (1 in 10¹⁰ per hour) rule of three: with zero failures in n independent realistic trials, the 95% upper bound is about 3/n
ReadoutFocal detail
Plate 06 — With no failures in n realistic trials, the failure rate is below about 3 in n at 95 percent confidence. That is useful for modest claims and useless for the rates life-critical software needs, which would take hundreds of thousands of years to show.

Choosing the inputs that matter

If testing cannot be complete, the question becomes which inputs to try. Three ideas do most of the work. Inputs can be grouped into classes the software should treat alike, with tests at the boundaries between classes, where off-by-one mistakes live. Combinations of settings can be covered systematically rather than exhaustively. And the hazard analysis of Chapter 4 can name the specific sequences, like the ventilator's four conditions, that would cause the most harm, so that each becomes a deliberate test.

The second idea has unusually direct evidence from medical devices. When NIST researchers Dolores Wallace and Richard Kuhn studied FDA recall reports from 1983 to 1997, 109 described their failures in enough detail to see what triggered them. In all but 3 of those 109 (97 percent; the authors give 98), testing every pair of parameter settings would have revealed the failure; the other 3 needed more than two conditions to occur together. Testing all pairs of settings grows slowly with the number of settings, so it is affordable where testing all combinations is not. The HAMILTON-C6 case shows the limit: its failure needed four conditions, so a pairwise campaign would not have been sure to find it, and only an analysis of what happens when a sensor fails during suctioning would have pointed to that sequence.

Testing, then, is necessary and never sufficient. The confidence a regulator accepts comes from testing chosen by requirements and risk, applied at several levels, repeated after every change, and backed by a process that makes faults less likely to be written in the first place. That process is the subject of the next part.

Write the test strategy from the hazards, not from the code

For each hazard the software could contribute to, list the inputs and sequences that could lead to it, including combinations of user actions, sensor states and timing. Make those deliberate tests; cover the remaining settings with boundary values and all-pairs combinations; and record which hazards rely on testing alone and which also rely on design controls, so that a reviewer can see where the testing argument is thin.

+ What this chapter established
  • Paths through software multiply with every decision, so even modest programs have far more behaviors than any campaign can test.
  • Testing can show that faults exist but cannot show their absence; coverage measures what ran, not whether it was right.
  • By the rule of three, proving a very low failure rate by failure-free testing would take an impossible amount of realistic operation.
  • Useful testing is a deliberate sample: input classes and boundaries, all pairs of settings, and sequences named by the hazard analysis.

+ Part II · Confidence from process

The software lifecycle.

IEC 62304, the international standard for the life cycle of medical device software, was published on 9 May 2006 and amended in June 2015. It prescribes activities and the records each leaves behind, not an order in which to do them, and the FDA has recognized the amended edition in full since January 2019. A second edition, widened to health software in general, reached its second committee draft in August 2026 and is not expected to be final before 2028. Because testing cannot be complete, the confidence a regulator accepts comes largely from how the software was made. The six chapters of this part follow that process: requirements derived from risk, an architecture that isolates what is dangerous, an account of the code taken from others, testing at several levels, and the ways agile teams and automated pipelines meet the same obligations.

04 — Requirements and risk

What it must do, and must never do.

+ The questionWhat has to be written down before a line of code is written, and why does risk decide how much?

A requirement is a promise that can be checked

A software requirement is a statement of what the software must do, or must never do, written so that a test, an inspection or an analysis can show whether it holds. "The display shall be easy to read" is not a requirement in this sense; "each dose value shall be shown with its unit, in characters at least 5 mm high, and shall remain on screen until the user confirms it" is. The difference is that the second can fail a test.

Requirements come from the intended use of Chapter 1. A claim to detect more than mild diabetic retinopathy from two photographs per eye implies requirements about which images are accepted, what happens when an image is too dark or blurred, what the three possible answers are and how each is shown, how results are tied to the right patient, and how long the analysis may take. Some come from users and clinicians, some from standards, some from the hazards the software could cause, and some from the rest of the system it lives in. Written down and reviewed, they become the yardstick for everything that follows: design is checked against them, tests are written from them, and a change is judged by which of them it touches.

Risk decides how much

Not every requirement carries the same weight. Risk management for medical devices follows ISO 14971, whose 2019 edition was confirmed in 2025. It traces a chain: a hazard, a potential source of harm; a sequence of events that leads to a hazardous situation, in which someone is exposed to it; and the harm that may follow. Each chain is judged by the severity of the harm and the probability that the harm occurs, and each unacceptable risk gets a control. The immunoassay guide in this library follows the chain for an analyzer, from a fault to a wrong result to a harm; software adds a twist to it.

The twist is probability. Chapter 2 showed that software failures are systematic, so there is no failure rate per hour to put into the chain. International regulators, in a 2025 document on software-specific risk, suggest that it can help to start from a probability of software failure of one and, where possible, to estimate the probability of harm from other factors, judging the risk by severity and by the controls that act if the failure happens. A risk control may itself be software, a check in another part of the program; or it may be hardware, procedure or labeling outside the software altogether.

07 — Every hazard to its evidence
HAZARD RISK CONTROL REQUIREMENT TEST AND RESULT wrong dose displayed unit shown with every dose; confirm before start SRS-041 dose and unit on one line, at least 5 mm high SRS-042 confirmation required TC-118 pass TC-119 pass result shown for the wrong patient patient ID checked against order before display SRS-077 ID match or block display TC-203 FAIL → fixed, re-run: pass start anywhere, reach everything design elements implementing each requirement (not shown) illustrative identifiers; real files hold hundreds of links
SpecificationCodeFaultReadoutFocal detail
Plate 07 — Traceability links each hazard to its control, each control to requirements, and each requirement to tests and results. With the links in place, a reviewer can start from a hazard, a failed test or a change and find everything connected to it.

Three classes of software

IEC 62304 turns this into three software safety classes. In the amended edition, software is class A if it cannot contribute to a hazardous situation, or if any it contributes to does not lead to unacceptable risk once risk controls outside the software are taken into account. It is class B if, after those external controls, it could still contribute to a hazardous situation whose possible harm is a non-serious injury, and class C if that harm could be death or serious injury. Until a team has classified its software, the standard treats it as class C.

The class sets how much process the standard demands. All three classes need a development plan, requirements, testing of the whole software against those requirements, risk management, configuration management, problem resolution and controlled release. In outline, class B adds a documented architecture and verification of software units and their integration; class C adds a detailed design of each unit and more demanding criteria for verifying them. The class also applies to parts of a program, so Chapter 5 shows how an architecture that separates dangerous functions can put most of the code in a lower class.

08 — Classifying software under IEC 62304
SOFTWARE SYSTEM class C assumed until classified can it contribute to a hazardous situation? unacceptable risk after risk controls outside the software? yes only external controls count CLASS A core process no no worst possible harm? yes non-serious injury death or serious injury CLASS B + architecture, unit and integration verification CLASS C + detailed design of each unit, stricter unit criteria per IEC 62304:2006+A1:2015; system testing at every class; activities summarized FDA DOCUMENTATION LEVEL (2023) assessed before risk controls could a failure present a probable risk of death or serious injury? yes no ENHANCED BASIC solved example: the pump with an external volume limit is class B and needs Enhanced
SpecificationCodeFocal detail
Plate 08 — IEC 62304 classes software by the worst harm that remains after risk controls outside the software, and the class sets the activities required. The FDA's documentation level asks a similar question before any controls, so the same software can be class B and still need Enhanced documentation.
Solved example: one program, two classifications

An infusion pump's software computes a delivery rate from a prescription. A wrong rate could kill, so before any controls the worst credible harm is death. Suppose the pump also has an independent hardware circuit, outside the software system, that stops the motor if the delivered volume exceeds a limit the clinician has set, and the team's analysis shows that with this circuit in place the worst remaining harm from a software error is a non-serious injury. Under IEC 62304 as amended, the software may then be class B, because the external control counts. The FDA's 2023 guidance on premarket software documentation instead asks whether a failure "could present a hazardous situation with a probable risk of death or serious injury" before any risk controls are considered. The same pump software would need the FDA's Enhanced documentation level. The guidance itself explains why a declaration of conformity to the whole of IEC 62304 is not needed: the two schemes classify differently.

Only controls outside the software lower its class

A check that runs inside the same program, on the same processor and from the same code base, can fail together with the function it guards, as the redundant copies of Chapter 2 did. IEC 62304 credits risk controls external to the software system when it sets the class. A team that lowers a class by pointing to a software check inside the system has misread the standard and has also left the risk in place.

Tracing the threads

The records that make this auditable are links. Each hazard links to the requirements that control it; each requirement links to the design elements that implement it and to the tests that verify it; each test links to its result. With the links in place, a reviewer can start anywhere: from a hazard, find whether its control was implemented and tested; from a failed test, find which hazard is now uncontrolled; from a change, find every requirement and test it touches. Without them, a stack of documents is only a stack.

The FDA's decision summary for IDx-DR shows the top of such a chain in public. It lists three risks to health, false positives, false negatives and operators failing to capture usable images, and maps each to its mitigations: clinical performance testing, software verification and validation built on a hazard analysis, a change protocol, labeling, operator training and human-factors testing. Below that summary, invisible to the public, sit the requirements and tests the reviewers read. Chapter 7 describes the testing; Chapter 27 lists what the FDA, the EU and India ask to see of it.

Classify early, and record why

Classify the software, and each software item, before architecture is fixed, and write down the hazards, the external controls credited and the worst remaining harm behind each class. Record the FDA documentation level separately, assessed before risk controls. Revisit both whenever a requirement, a control or the intended use changes, because a class that rested on a hardware limit fails silently if that limit is removed.

+ What this chapter established
  • A software requirement states what the software must or must never do in a form that a test or inspection can check.
  • ISO 14971 traces hazards to harms; because software failures are systematic, regulators suggest assuming a software failure will occur.
  • IEC 62304 classes software A, B or C by the worst harm after controls outside the software, and the class sets the process required.
  • Traceability links hazards, requirements, design and tests, so a reviewer can follow any risk to its control and its evidence.
05 — Architecture

Keeping the dangerous part small.

+ The questionHow does the way software is divided into parts change what must be proven about each?

A timer that does not trust the program

A watchdog is a timer, often a separate circuit, that the software must reset at regular intervals. If the program hangs, stuck in a loop or waiting for something that never arrives, it stops resetting the timer, the timer runs out, and the watchdog forces the device into a safe state: it stops a motor, raises an alarm or restarts the processor. The FDA's guidance on infusion pumps, finalized in 2014, asks makers to describe their watchdog timer, lists watchdog tests among the safety mechanisms a pump should have, and lists a watchdog failure among the hardware causes of system failure.

The watchdog embodies the main idea of software architecture for safety. A function that must not fail is protected by something that does not share its weaknesses. The watchdog does not run the dose calculation and does not need to understand it; it only needs proof of life at regular intervals, and its own logic is small enough to verify thoroughly. The same guidance adds a caution that applies to every safety mechanism: the safety case must also cover the hazards that the mechanism itself could start, such as, to take an example of our own, a watchdog that resets a pump in the middle of an infusion.

A program that runs on time can still be wrong

A program that runs normally while computing a wrong dose keeps resetting its watchdog on time, so the watchdog never acts. Protection against a wrong output needs a different mechanism: an independent check of the result, a limit enforced outside the software, or a second calculation by different means. Naming which failure each safety mechanism catches, and which it cannot see, is part of the design.

09 — What a watchdog can and cannot see
A HANG software resets the timer program hangs timeout watchdog timer safe state: motor stopped, alarm A WRONG ANSWER timeout watchdog timer computed dose wrong watchdog never acts needs an independent check of the result TIME schematic
CodeFaultReadoutFocal detail
Plate 09 — A watchdog forces a safe state when software stops responding, but a program that runs on time while computing the wrong answer keeps resetting it. Each safety mechanism catches some failures and is blind to others, and the design has to say which.

Items, units and boundaries

IEC 62304 describes a program as a software system divided into software items, any identifiable parts, which are divided in turn until they reach software units, the items not divided any further. The division is not a filing scheme. It decides where faults can travel. If one item can overwrite another's memory, starve it of processor time or hand it corrupted data without detection, a fault in the first is a fault in the second, and both must be treated as one.

The standard calls the remedy segregation: any mechanism that prevents one software item from negatively affecting another. Segregation can be physical, with the critical function on its own processor; it can use the memory protection of an operating system, so that one process cannot write into another; or it can rest on checks at the boundary, so that data crossing it is validated. Each mechanism has to be shown to work, because a claimed boundary that a fault can cross is worse than none: it lowers the rigor applied to the code on the far side without lowering the risk.

Segregation is what lets the class of Chapter 4 apply to parts of a program rather than the whole. If the items that could contribute to serious harm are separated from the rest, only they need the class C activities, and the user interface, the logging and the reporting code can be developed at class B or A. Architecture is therefore a decision about evidence as much as about structure.

Solved example: what segregation buys

Suppose a pump's software has 200,000 lines, of which the dose calculation and the delivery control, the only parts whose failure could cause death or serious injury, take 8,000. Without demonstrated segregation, the whole program is class C, and the detailed design and stricter unit-verification criteria of that class apply to all 200,000 lines. If the 8,000 lines run on a separate processor that the rest cannot write to, and every command crossing the boundary is checked against the prescription, the class C work applies to 8,000 lines, 4 percent of the program. The remaining 96 percent still needs its own class, its own tests and evidence that the boundary holds.

10 — Segregation puts the strictest work where it matters
SOFTWARE SYSTEM: 200,000 LINES logging and reports CLASS A network communications CLASS B user interface CLASS B operating system and libraries (SOUP) PER ITS ROLE SEPARATE PROCESSOR dose calculation and delivery control: 8,000 lines CLASS C CHECK boundary: separate processor; every command checked against the prescription the boundary must be shown to hold class C work: 8,000 of 200,000 lines = 4%
SpecificationCodeReadoutFocal detail
Plate 10 — If the code that could cause serious harm is separated by a boundary shown to hold, only that code needs class C rigor. Here segregation confines the strictest work to 8,000 of 200,000 lines, while the rest keeps the class and tests that fit its own risk.

Where the parts live

The architecture also has to say where each part runs and what it depends on. Much of the code in a modern device was written by someone else: the operating system, the network stack, a database, a library that reads images, a cloud service. Each is an item the maker did not write and cannot fully inspect, and the architecture has to place it, bound it and say what the device does when it misbehaves. Chapter 6 is about those items.

The thread's device shows the split in its public record. IDx-DR has three software items with their own version numbers: the client on the camera's computer, a service on the server that passes images and results between client and analysis, and the analysis that grades the images. The 510(k) cleared in June 2021 moved the client from version 2.0.1, as its summary lists the authorized software, to 3.2.0, the analysis from 2.0.1 to 2.1.1 and the service from 1.0.0 to 1.1.2; the next, cleared in June 2022, moved them to 3.5.0, 2.3.0 and 1.2.0. Because the parts are separate, a maker can change one and argue, with evidence, that the others are untouched. The split also shapes what can go wrong: an analysis that runs on a server depends on the network, and the 2021 version added messages that tell an operator whether a submission was lost or is simply not ready yet.

Chapter 21 returns to the 2022 change, which replaced one model inside the analysis, the classifier that judges image quality, and changed which patients received an answer. An architecture lets a change be confined to one item. It does not guarantee that the change's effects stay there, which is why every change needs regression tests at the level of the whole system.

Draw the architecture with its classes and boundaries

Draw every software item, including those from third parties, with its safety class, where it runs, and the mechanism that separates it from its neighbors. For each boundary that lowers a class, record the evidence that the mechanism works, such as memory-protection settings or the separate processor, and include a test that tries to break it. For each safety mechanism, record which failures it detects, which it cannot, and what hazards it could create itself.

+ What this chapter established
  • A watchdog shows the core idea: protect a critical function with a mechanism that does not share its weaknesses.
  • IEC 62304 divides software into items and units; segregation is any mechanism that stops one item from harming another.
  • Demonstrated segregation lets the strictest class apply only to the code that could cause serious harm, not to the whole program.
  • IDx-DR's three separately versioned items show how architecture confines a change, though system-level tests must confirm its effects.
06 — Software you didn't write

SOUP, OTS and the bill of materials.

+ The questionMost of the code in a modern device was written by someone else. How can a maker answer for it?

When the operating system expires

On 8 April 2014 Microsoft ended extended support for Windows XP, twelve years and three months after the system's lifecycle began. After that date Microsoft stopped routine security fixes for it. Medical devices built on Windows inherited that calendar, and many outlived it: an imaging scanner, a laboratory analyzer or a radiology workstation is expensive, is expected to last a decade or more, and was validated with one exact version of its operating system. Replacing that system is a design change that needs new testing and, for some devices, a new review.

The cost of waiting became visible on 12 May 2017, when the WannaCry ransomware spread across networks running unpatched or unsupported versions of Windows. The UK's National Audit Office found that at least 81 of the 236 trusts of the National Health Service in England, 34 percent, were affected, and identified 6,912 cancelled appointments, with more than 19,000 estimated. Bayer confirmed two reports from US customers of infected devices of its own, and Siemens Healthineers warned that some of its products might be affected. In 2020 a security vendor's survey of its customers' networks reported that 83 percent of the medical imaging devices it saw ran operating systems that were no longer supported, a share that had risen sharply since 2018 and that the vendor attributed to the end of support for Windows 7.

11 — When the operating system outlives its support
WINDOWS XP SUPPORTED April 2009: mainstream support ends 8 Apr 2014: extended support ends; no more security fixes 2002 2004 2006 2008 2010 2012 2016 2018 2020 2022 2024 12 May 2017: WannaCry; at least 81 of 236 NHS trusts in England affected 2020 vendor survey: 83% of imaging devices on unsupported operating systems (vendor's figure) example device: launched 2009, sold to 2012, used 10 years 8 years without security fixes support ends before the device does
CodeThreatFaultFocal detail
Plate 11 — Windows XP's support ended in April 2014, but devices built on it stayed in service for years. A device sold in 2012 and used for ten years runs eight years without security fixes, which is why makers must now state each component's end-of-support date.
Solved example: support ends before the device does

Suppose a device is launched in 2009 on Windows XP, sold until 2012, and used for ten years from the date of sale. Extended support ended in April 2014, five years after launch. A unit sold in 2012 stays in service until 2022, eight years after its operating system stopped receiving security fixes. Any vulnerability published in those eight years stays open on that unit unless the maker can fix it some other way. The FDA's 2026 cybersecurity guidance addresses this by asking makers to state, for each software component, its level of support and its end-of-support date.

SOUP and OTS

Two terms name the code a maker did not write, and they overlap without being the same. IEC 62304 calls it SOUP, software of unknown provenance: a software item that is already developed and generally available and was not developed to be part of the device, or a software item developed earlier for which adequate records of its development are not available. The FDA calls it OTS, off-the-shelf software: a generally available software component used by a device maker that cannot claim complete control of its life cycle.

The emphasis differs. SOUP is about provenance: a maker's own old code, written before anyone kept records, is SOUP even though nobody bought it. OTS is about control: a commercial operating system is OTS because its maker, not the device maker, decides what goes into it and when. Most third-party components are both. A compiler or a test tool that never ships in the device is not SOUP, because it is not part of the device; it is validated for its use under the quality-system rules of Chapter 9.

IEC 62304 asks the same things of every SOUP item. The maker states what functional and performance requirements the device places on it, what hardware and software it needs, and its exact identity and version. It evaluates the anomalies the component's own maker has published, to see whether any could cause a hazard, and controls its configuration so that the version tested is the version shipped. The FDA's guidance on off-the-shelf software, last revised in August 2023, asks for much the same and adds a plan for maintaining support. At its Enhanced documentation level, it also asks for assurance about how the component's developer works and what the device maker will do if that developer changes or stops its support.

How much of the code is borrowed

The borrowed share is usually most of the program. A vendor that audits software for company acquisitions reported in 2026 that 98 percent of the 947 codebases it examined contained open-source code, with a mean of 1,180 open-source components per application, and that 87 percent contained at least one component with a known vulnerability. The sample is software being sold in a transaction, not medical devices, but device software is built from the same libraries. An operating system, a network stack, a database, an image library, a cryptography package and a user-interface toolkit can each pull in dozens of further components of their own.

Models and data are now part of the list. A device can include a model pretrained by another company on images the device maker has never seen, or a public dataset used for tuning. Neither is code in the usual sense, but each is a component whose origin, version and known limits matter as much as a library's. The CycloneDX format for bills of materials added a machine-learning section in 2023 so that models and datasets can be listed with the code.

12 — What a software bill of materials lists
DEVICE SOFTWARE V3.2 maker's own code operating system TLS library database image decoder compression library a dependency of a dependency pretrained model public tuning dataset SBOM RECORD: IMAGE DECODER NTIA 2021 MINIMUM supplier open-source project component name image decoder version x.y.z unique identifier package URL or CPE dependency relationship → compression library author of SBOM data device maker timestamp date and time FDA GUIDANCE ADDS support level actively maintained end-of-support date yyyy-mm-dd CISA 2026 ADDS hash SHA-256 digest license license identifier tool name, generation context when fixes stop VEX STATEMENT FOR A PUBLISHED VULNERABILITY NOT AFFECTED AFFECTED FIXED UNDER INVESTIGATION stated separately from the SBOM: can the vulnerability be exploited in this device?
CodeDataThreatFocal detail
Plate 12 — An SBOM lists every component, including dependencies of dependencies and now models and datasets, with fields that identify each exactly. The FDA also asks for each component's support level and end-of-support date; whether a vulnerability matters is stated separately in a VEX.

The bill of materials

A software bill of materials, or SBOM, is the machine-readable list of every component in a piece of software, with enough detail to identify each exactly. The US National Telecommunications and Information Administration set out seven minimum fields in July 2021: the supplier, the component's name, its version, other unique identifiers, its dependency relationships, the author of the SBOM data and a timestamp. The US cybersecurity agency CISA published updated minimum elements in mid-2026, adding a cryptographic hash of each component, its license, the tool that generated the SBOM and the context in which it was generated, and extending the scope explicitly to AI software and software delivered as a service. Two formats dominate: SPDX, published as the international standard ISO/IEC 5962 in 2021, and CycloneDX, an Ecma standard since 2024.

An SBOM lists ingredients; exposure is a separate question

A component with a published vulnerability may be present but unreachable: the device may never call the affected function or never expose it to a network. A component with no published vulnerability may still be flawed. The SBOM answers which components are in the device; whether a given vulnerability can be exploited in that device is a separate statement, which the industry calls a VEX: not affected, affected, fixed, or under investigation.

For devices the SBOM is no longer optional in the US. Since 29 March 2023 the law has required every premarket submission for a "cyber device", one with software that can connect to the internet, a condition the FDA reads broadly, to include one, including commercial, open-source and off-the-shelf components, and the FDA's guidance asks for each component's support level and end-of-support date alongside it. Chapter 24 follows the SBOM after release, where it does its real work: when a new vulnerability is published in a component, the list tells a maker, and a hospital, which devices contain it, far faster than a search of old build records.

Keep a register of every borrowed component

For each SOUP or OTS item, record its exact version, its supplier, the requirements the device places on it, the published anomalies reviewed and the conclusion for each, its support level and end-of-support date, and a plan for replacing it. Pin versions so that what was tested is what ships, generate the SBOM from the build rather than by hand, and review the anomaly lists again before each release.

+ What this chapter established
  • Windows XP's end of support in 2014 left devices running unpatched for years, as WannaCry showed in 2017.
  • IEC 62304's SOUP is code of unknown provenance; the FDA's OTS is code whose life cycle the maker cannot control; most components are both.
  • For each borrowed component a maker must state requirements, exact version and evaluated anomalies, and keep it under configuration control.
  • An SBOM lists every component, now including models and datasets; whether a vulnerability is exploitable is a separate VEX statement.
07 — Verification

Evidence at every level.

+ The questionWhich tests, at which level, give evidence a regulator accepts, and when is testing enough?

A sequence of steps nobody combined

In February 2023 Elekta began correcting its Monaco radiotherapy treatment planning system after finding that one sequence of user steps could make it display an inaccurate dose. A planner who added contours to a plan and then re-optimized it, without forcing a density value outside the patient's external outline, could see a dose distribution that did not match what the plan would deliver. The correction covered 2,020 installed systems running four builds of version 5.11, and the FDA recorded the root cause as software design. Each step on its own, drawing a contour, assigning densities, optimizing a plan, displaying a dose, was an ordinary use of the software. The fault lay in the combination.

Failures of this kind are why device software is tested at more than one level. A test of a single function can show that the function meets its own specification. Only tests that put functions together, in the sequences users actually follow, can show how they behave in combination, and only tests of the whole system in its intended environment can show whether the device does what its users need.

Levels of evidence

Verification is the FDA's word for objective evidence that the outputs of a stage of development meet the requirements set for that stage; in short, that the software was built as specified. It happens at several levels. Unit testing checks each smallest piece in isolation, often automatically, against its detailed design. Integration testing checks that units and items work together, including the boundaries Chapter 5 asked an architecture to defend. System testing checks the complete software, usually on the real hardware, against the software requirements. Around the tests sit other verification activities that find faults without running the code at all: reviews of requirements and design, inspection of the code by a second engineer, and static analysis tools that search the source for whole classes of mistakes, such as reading memory that was never written.

IEC 62304 asks for system testing at every class, unit and integration verification at class B and above, and records of each. The FDA's 2023 guidance on premarket software documentation asks, at the Basic level, for a summary of unit, integration and system testing and the protocol and report of the system-level tests; at the Enhanced level it asks for the unit and integration test protocols and reports as well. Both also ask for a list of the unresolved anomalies, the known defects the software will ship with, each with an assessment of its effect on safety and effectiveness. A device can be released with known defects, provided each has been judged and none leaves an unacceptable risk.

13 — Verification at three levels, validation above them
user needs and intended use validation: usability and clinical evidence validates software requirements system tests verifies architecture (items) integration tests verifies detailed design (units) unit tests CODE where combinations fail (Monaco, 2023) verification: built as specified validation: the right specification IEC 62304 covers verification; device validation is outside its scope find faults without running the code REVIEWS STATIC ANALYSIS CODE INSPECTION FDA 2023, Basic: summary + system test protocol and report; Enhanced adds unit and integration protocols and reports the V is a map of evidence, not an order of work
SpecificationCodeReadoutFocal detail
Plate 13 — Each level of specification is checked by its own level of testing, and validation asks whether the specification itself meets users' needs. Faults that live in combinations of functions, like the Monaco dose display, surface at integration and system level.

How much testing is enough

Chapter 3 showed that the answer cannot be "all of it". Coverage tools measure part of the answer by counting what a test campaign executed, and the FDA's 2002 guidance describes a ladder of coverage measures. Statement coverage, every line run at least once, it calls insufficient to provide confidence. Decision coverage runs each branch both ways. Condition coverage makes each condition inside a decision both true and false. Multiple-condition coverage tries every combination of the conditions in each decision. Full path coverage it calls generally not achievable. The FDA does not prescribe a level; the maker chooses one in proportion to risk and justifies it.

Solved example: one decision, three coverage measures

An infusion pump stops when "(line pressure is high AND the occlusion alarm is enabled) OR the door is open". Decision coverage needs two tests: one in which the pump stops and one in which it does not. Condition coverage can also be met with two tests, one with all three conditions true and one with all three false, which never shows what happens when pressure is high but the alarm is disabled. Multiple-condition coverage needs every combination of the three conditions, two to the power three, or 8 tests. For one decision the difference between 2 tests and 8 is small; across a program with thousands of decisions it is the difference between a campaign that can be run every night and one that cannot, which is why the strictest measures are kept for the code whose failure could cause serious harm.

Coverage measures say what ran, so they are paired with measures of what was checked: every requirement traced to at least one test that would fail if the requirement were broken, every hazard-related sequence from the risk analysis tested on purpose, and combinations of settings covered by the all-pairs method of Chapter 3. When a campaign meets those targets, passes, and leaves only judged anomalies, the team has the evidence a reviewer looks for. It has not proved the software correct, and the documentation says so by listing what was not tested and why.

Regression after every change

The FDA's count of the 1990s, in which 79 percent of software recalls followed a change, makes regression testing the most important habit in maintenance. Each change is followed by re-running the tests that cover what it touched and, for anything beyond a trivial change, the whole system-level suite, so that behavior verified before the change is verified again after it. A suite that runs automatically makes this cheap enough to do every time, which is the subject of Chapter 9.

The thread's device shows a regression test on a model. When Digital Diagnostics replaced the classifier that judges image quality in IDx-DR version 2.3, cleared in 2022, it re-ran the new version on the images from the 2017 trial and compared the answers with the original version's. Chapter 21 shows what that comparison found, and what it could and could not prove.

14 — Every change re-runs what it could disturb
v1.0 v1.1 v1.2 v2.0 change change change regression suite: each square a test that passes fixed before release a change broke an old behavior: caught found in the build, not in the field the suite keeps every test written for every past fault (Chapter 2) schematic FDA count 1992–1998: 192 of 242 software recalls (79%) followed changes after release
CodeFaultReadoutFocal detail
Plate 14 — Regression testing re-runs earlier tests after every change, and the suite grows with each fault ever found. It is the main defense against the most common source of software recalls: changes made after release.

Verification and validation

Verification asks whether the software meets its specification; validation asks whether the specification was right. The FDA defines software validation as confirmation, by examination and objective evidence, that software specifications conform to user needs and intended uses. For a software device, validation reaches beyond testing code: it includes usability studies with representative users and, where the claim requires it, clinical evidence that the outputs are right for patients. IEC 62304 deliberately stops short of it; its scope excludes the validation and final release of the device, which belong to the maker's design controls.

A passing system test validates nothing about patients

A system test checks the software against requirements the team wrote. If a requirement was wrong, for example an image-quality threshold set for a camera the clinics do not use, every test can pass while the device fails its users. Validation needs evidence from outside the specification: real users, real conditions and, for a diagnostic claim, a comparison with an independent reference.

IDx-DR's 2018 evidence shows both. The trial of the opening scene validated the claim: novice operators with four hours of training, real patients, a reference standard. Alongside it, a smaller study checked repeatability. Twenty-four participants were each imaged ten times by three operators on two cameras of the same model; the system analyzed 235 of the image sets, gave identical answers for 23 of the 24 participants across every repeat, and agreed with itself 99.6 percent of the time. Repeatability is not accuracy, but a screening program whose answer changed with the operator would have failed before accuracy was even asked.

Write the verification plan as a coverage argument

For each software item, state which levels of testing apply, which coverage measure is targeted and why it fits the item's class, how requirements and hazard sequences map to tests, and which test suites are re-run after which kinds of change. Keep the unresolved-anomaly list with each release and record, for every entry, why it leaves no unacceptable risk.

+ What this chapter established
  • Device software is verified at the unit, integration and system levels, with reviews and static analysis catching faults that tests miss.
  • Coverage measures how much of the code ran; the level chosen should rise with the harm a failure could cause, and is paired with requirement and hazard coverage.
  • A release may carry known defects only if each is listed and judged; regression testing after every change guards against the most common source of recalls.
  • Verification shows the software meets its specification; validation, including usability and clinical evidence, shows the specification meets patients' needs.
08 — Agile under a standard

Sprints with a paper trail.

+ The questionDoes IEC 62304 force a team into phases, or can it work in two-week sprints?

A standard without a sequence

IEC 62304 lists activities: planning, requirements analysis, architectural design, detailed design, implementation and unit verification, integration, system testing and release, with maintenance, risk management, configuration management and problem resolution running alongside. Read in that order, the list looks like a waterfall, in which each phase is finished and signed before the next begins. For years many device teams worked that way and assumed the standard required it. The standard asks for the activities to be done and recorded and for the maker's development plan to say how; it does not say in what order, or how many times.

Most software outside regulated industries is now built iteratively. A team works in short cycles of one to four weeks, each ending with working, tested software; requirements are refined as users see results; and the design grows with the product. The question for a device maker is whether that way of working can produce the evidence of Chapters 4 to 7: requirements traced to hazards and tests, a documented architecture, verification at each level, and a controlled release. A technical report from AAMI, the US standards association for medical technology, answers it. TIR45, first published in 2012 and revised in 2023, describes how agile practices map onto IEC 62304, and the FDA has recognized both editions in full, the second in May 2025.

Four layers of time

TIR45 describes the work in four nested layers. The project or product layer runs for months to a year or more and holds the high-level requirements and the overall architecture. A release runs for one to several months and ends with software that could be delivered. An increment, often called a sprint, runs for one to four weeks and ends with integrated, tested software. A story, one small piece of user-visible function, takes one to a few days.

Each IEC 62304 activity is placed in one or more layers. Planning happens at every layer, at its own scale. High-level requirements and the coarse architecture belong to the project; detailed requirements, detailed design and implementation with its unit tests belong to each story; integration and system testing happen at story, increment and release level; and the formal release activity sits at the project layer, carried out each time a release ships. A user story with its acceptance criteria becomes a design input, and the end of an increment or a release becomes a point for design review. The development plan states which of those reviews are formal and recorded.

15 — Four layers of time in an agile project
AAMI TIR45 LAYER IEC 62304 ACTIVITIES PLACED IN IT PROJECT OR PRODUCT months to a year or more 5.1 planning 5.2 high-level requirements 5.3 architecture 5.8 release RELEASE one to several months 5.1 planning 5.6 integration 5.7 system testing INCREMENT (SPRINT) one to four weeks 5.1 planning 5.6 integration 5.7 system testing STORY one to a few days 5.1 planning 5.2 detailed requirements 5.3 architecture 5.4 detailed design 5.5 implementation and unit tests 5.6 integration 5.7 system testing each release that ships after AAMI TIR45 (2012; 2023 edition, FDA-recognized 2025); clause numbers from IEC 62304
SpecificationCodeFocal detail
Plate 15 — AAMI TIR45 places each IEC 62304 activity in one or more of four nested layers of an agile project. Detailed design and unit testing happen story by story, integration and system testing in every increment, and the formal release activity at the project layer, each time a release ships.
Solved example: what a year of sprints leaves behind

Suppose a team plans a 12-month project in two-week increments, with a release to users every three months. That is 26 increments and 4 releases. If the team completes about ten stories per increment, it finishes about 260 stories in the year, and under a definition of done that includes them, each leaves updated requirements, design notes, unit tests, trace links and, where it touches a hazard, an updated risk analysis. The four release ends carry the formal system test, the anomaly review and the release approval. Instead of one large documentation effort at the end, the evidence is produced in about 260 small pieces, each reviewed while the work is fresh, and each release is assembled from records that already exist.

A sprint review is a meeting; a release is a regulated event

Software can move from a developer's machine to a test environment many times a day, and an increment can be demonstrated every two weeks. A release to patients is different: IEC 62304 requires that verification be complete, that known residual anomalies be documented and evaluated, that the released version and how it was built be recorded, and that the software be archived and reliably delivered. Agile speeds up everything before that point. It does not shrink the point itself.

Done means verified and recorded

The practice that makes agile work under a standard is a strict definition of done. A story is not finished when its code runs; it is finished when its requirement is written and reviewed, its design and code are reviewed, its unit and integration tests pass, its trace links are in place, the risk analysis reflects any new hazard or control, and any change to SOUP is recorded. If any of these is missing, the story goes back into the work, exactly as it would if a test failed.

Two further habits keep the records honest. The backlog of stories is not the requirements specification, but it feeds it: as stories are accepted, their requirements enter a controlled baseline that is approved at each release, so that the release is checked against a stable specification rather than a moving list. And traceability is kept live in the tools the team already uses, linking stories, requirements, code changes and tests as the work happens, so that the trace matrix of Chapter 4 is a report generated from the links rather than a document written afterwards.

16 — Done means verified and recorded
DEFINITION OF DONE requirement written and reviewed design and code reviewed unit and integration tests pass trace links in place risk analysis updated SOUP changes recorded STORY show dose unit with value missing item: back to the work DONE the trace matrix becomes a report RECORDS PRODUCED PER INCREMENT each increment: about 10 stories waterfall: documents written at the end increments 1–26 over one year illustrative: 26 increments, about 260 stories
SpecificationCodeFaultReadoutFocal detail
Plate 16 — Under a strict definition of done, each story leaves its requirement, reviews, tests, trace links and risk updates before it counts as finished. The evidence then accumulates increment by increment instead of being written after the work.

What changed in 2023

The 2023 edition of TIR45 kept the core and updated the setting. According to AAMI, it adapts documentation to digital practice, with streamlined approval processes and signature requirements, and brings cybersecurity, risk management, design validation and usability engineering into agile work rather than treating them as separate phases. Its references were brought up to date with ISO 13485:2016, the amended IEC 62304 and ISO 14971:2019. The FDA accepts declarations of conformity to the 2012 edition only until 2 July 2028.

The same shift appears elsewhere. The second edition of ISPE's GAMP 5, the guide that regulated companies use to validate the computerized systems they buy and build, published in 2022, states that its life cycle is not inherently linear and supports iterative and incremental methods; Chapter 19 returns to it. The draft second edition of IEC 62304 itself, according to the German electrical standards association VDE in July 2026, places agile development in an informative annex.

Put the agile model in the development plan

Write the layers, their lengths and the activities placed in each into the software development plan, and state which reviews are formal. Make the definition of done include requirements, design, tests, trace links and risk updates for every story, and approve a requirements baseline at each release. An auditor who opens the plan should find the team's real way of working described there, not a waterfall it does not follow.

+ What this chapter established
  • IEC 62304 requires activities and records, not an order of phases, so iterative development can meet it if the plan says how.
  • AAMI TIR45, recognized by the FDA, maps the standard's activities onto project, release, increment and story layers.
  • A strict definition of done, a requirements baseline at each release and live traceability let evidence accumulate as the work is done.
  • Increments can be frequent, but a release to users stays a controlled event with verification, anomaly review, archiving and approval.
09 — Pipelines, releases and fixes

Every build a record.

+ The questionWhen code is built and tested automatically on every change, what turns that pipeline into evidence, and where must a person still decide?

Deploying many times a day

In the 2024 survey of software delivery run by DORA, a research program at Google Cloud, the best-performing fifth of respondents, 19 percent, deployed changes on demand, often several times a day. A change took less than a day to go from code to production, 5 percent of their deployments failed, and they recovered from a failed deployment in less than an hour. The survey covers technology workers in general, not medical device makers, and the gap is deliberate: a device maker cannot put a new version in front of patients several times a day, and should not.

What the best general software teams do have is a pipeline: an automated sequence that takes every change a developer submits, builds the software from source, runs the tests, checks the code with analysis tools, and packages the result, with no step done by hand. Device makers increasingly use the same machinery, not to release faster to patients but to make every internal build a complete, repeatable record. The pipeline turns the regression testing of Chapter 7 from an occasional campaign into something that happens on every change.

Solved example: a failure rate meets a deployment rate

At the elite team's change failure rate of 5 percent, three deployments a day produce 0.15 failed deployments a day, about one a week counting every day. In a test environment that is a cheap lesson: the failure is found within hours and the change is rolled back. For released device software the same arithmetic would mean a field failure each week. The device maker's answer is to run the frequent cycle internally, where failures are expected and contained, and to keep the release to patients as a separate, verified and approved event, as Chapter 8 described.

What makes a pipeline evidence

A pipeline produces evidence when its outputs can be trusted and traced. Configuration management, which IEC 62304 requires at every class, means that every item that goes into a build, source code, SOUP versions, build scripts, the compiler and its settings, is identified and controlled, so that the same inputs produce the same software. A reproducible build goes one step further: rebuilding a past release from its recorded inputs gives a bit-for-bit identical result, so the version tested and the version shipped can be shown to be the same.

Each stage then leaves a record tied to the exact build it checked. Unit and integration test results are stored with the commit that produced them and linked to the requirements they verify. Static analysis reports, coverage reports and the software bill of materials of Chapter 6 are generated by the pipeline rather than assembled by hand, so they describe the build that exists rather than the one someone remembers. The full system tests, which may take hours on real hardware, run nightly or before each candidate release. When a release is proposed, the evidence for it is not collected; it is retrieved.

17 — A pipeline that leaves a record at every step
commit change ID reproducible build build hash static analysis analysis report unit and integration tests test results + trace SBOM generated SBOM nightly system tests on hardware system test report release candidate archived archive tool fail: fix and resubmit people decide: anomalies judged, release approved the pipeline checks; a person approves RELEASE TO USERS runs on every change; the release is a separate, approved event
SpecificationCodeReadoutFocal detail
Plate 17 — A device pipeline builds and tests every change and leaves a record tied to each build, from test results to the SBOM. Judging the remaining defects and approving the release stay with named people.

Two decisions stay with people. Someone must judge each unresolved anomaly and decide that none leaves an unacceptable risk, and someone authorized must approve the release itself. The pipeline can check that every required record exists and every test passed. Whether the remaining defects are acceptable for patients is a judgment, and IEC 62304 and the regulators expect a named person to make it.

Trusting the tools

The pipeline's tools are software too: the build system, the test runner, the coverage tool, the static analyzer, the SBOM generator, the issue tracker that holds the trace links. None of them ships in the device, so none is SOUP, but each is quality-system software that must be validated for its intended use, whether bought, open-source or written in-house. A test runner that reports a pass for a test that never ran, or a coverage tool that miscounts, would manufacture false evidence about a device that might be dangerous.

ISO 13485:2016, the quality-management standard that the FDA's rules have incorporated since February 2026, requires a maker to validate software used in its quality system and in production before first use and after changes, in proportion to the risk of its use. ISO's technical report 80002-2, published in 2017, gives guidance for doing so. The FDA's guidance on computer software assurance, finalized on 24 September 2025 and reissued on 3 February 2026 under the title Computer Software Assurance for Production and Quality Management System Software, sets out a risk-based approach. A function is high process risk if its failure "may result in a quality problem that foreseeably compromises safety"; such functions get documented, scripted testing, while others can be assured by lighter, unscripted methods such as exploratory testing. The guidance replaces one section of the FDA's 2002 software validation guidance and explicitly does not cover the device software itself, which remains under design controls.

The pipeline's own failures become false evidence

A test step that silently skips tests after a configuration change, a runner that treats a crash as a pass, or an SBOM generator that misses dependencies loaded at run time all produce records that look complete. Seeding known faults and confirming that the pipeline catches them is the most direct check, and it should be repeated whenever a tool or its configuration changes.

18 — Assurance in proportion to risk
software used in production or the quality system device software itself: not covered (design controls, Chapters 4–7) identify its intended use could its failure cause a quality problem that foreseeably compromises safety? YES: HIGH PROCESS RISK scripted testing (robust), documented NO: NOT HIGH PROCESS RISK unscripted testing: exploratory, scenario, error guessing establish the record EXAMPLES · test runner recording verification results · SBOM generator · build system · team chat · meeting scheduler risk to the evidence
SpecificationCodeReadoutFocal detail
Plate 18 — Computer software assurance scales the effort of validating a tool to the risk its failure poses: scripted testing for tools whose failure could compromise safety, lighter methods for the rest. The device software itself stays under design controls.

After release: fixes and maintenance

Release starts the longest part of a device's software life. IEC 62304 requires a maintenance plan and a problem-resolution process: field reports and internal findings become problem reports, each is investigated and classified for its effect on safety, changes go through the same controlled process as development, and trends across problems are analyzed. Chapter 2's lessons apply with full force here. In the FDA's count from the 1990s, 79 percent of software recalls followed changes made after release, and the Tandem app recall of 2024 was a change to software that already worked.

Hosted software changes the mechanics. When the analysis runs on a maker's server, as IDx-DR's does, a fix can reach every user at once and can be rolled back at once, which is an advantage for safety and a risk for control: a change can also reach every user at once. A switch that turns on a function for some users, often called a feature flag, is a change to the released device and goes through change control, however small the switch; whether it needs a regulator's review depends on its effect on claims and risk (Chapter 28). And much code in the field predates the processes now expected of it. The 2015 amendment to IEC 62304 added a route for this legacy software, legally marketed but without enough evidence that it was built to the standard: a risk analysis using field experience, a gap analysis against the standard's requirements, a plan to close the gaps that matter, and a documented rationale for continued use. Whether a given change needs a regulator's approval before release is a separate question, which Chapter 28 answers for the US, the EU and India.

Validate the pipeline tools in proportion to the evidence they produce

List the tools in the pipeline and, for each, the records it produces and what a silent failure would do to the device's evidence. Give the tools that produce verification results, coverage, builds and SBOMs scripted testing with seeded faults; pin their versions and archive the build environment with each release; and use lighter, unscripted assurance for tools whose failure could not compromise a release.

+ What this chapter established
  • A pipeline builds and tests every change automatically; device makers use it to make each build a complete record, not to release faster to patients.
  • Configuration management, reproducible builds and records tied to each build let a release's evidence be retrieved rather than assembled.
  • Pipeline tools are off-the-shelf software that must be validated in proportion to risk, the approach of the FDA's computer software assurance guidance.
  • After release, maintenance and problem resolution follow the same controls, and legacy software needs a gap analysis and a rationale for continued use.

+ Part III · Learning from data

How an AI model is built and scored.

In 2016 a team at Google trained a network to grade photographs of the retina for diabetic retinopathy without writing a single rule about what the disease looks like: no instruction to look for small hemorrhages, no threshold for the size of a lesion. Every rule the network applied had been fitted to 128,175 images that ophthalmologists had already graded. A program built that way can still be tested, but only by comparing its answers with right answers on data it has not seen, and every step of that comparison involves choices that decide what the result means. The four chapters of this part follow those choices: what a model is, how the data that builds it is kept apart from the data that judges it, what the right answers are and who decided them, and how a score becomes a decision with a measurable error.

10 — A model is a fitted function

Rules nobody wrote.

+ The questionWhat is a machine-learning model, and how is building one different from writing code?

Twenty-five million adjustable numbers

ResNet-50, one of the most widely used networks for analyzing images, has 25,557,032 adjustable numbers in the reference implementation published with the PyTorch software library. Each is a parameter: most are weights that multiply values inside the network before they are passed on, and the rest are offsets added along the way. A photograph goes in as a grid of pixel intensities; it passes through some fifty layers, each of which combines the values from the layer before using its weights; and a score comes out, for example the network's estimate that the image shows a particular disease. Change the weights and the same photograph produces a different score.

A machine-learning model is that whole arrangement: a fixed structure, chosen by engineers, with its parameters set from data rather than by hand. The structure is code like any other, written, reviewed and tested. The parameters are not written at all. They are found by training: the model scores a batch of examples whose right answers are known, a measure of how wrong it was, called the loss, is computed, and every parameter is nudged slightly in the direction that would have made the loss smaller. Repeated over millions of batches, the nudges settle into values that make the model's scores match the known answers well. The Google system of 2016, an ensemble of ten Inception-v3 networks, each of which its authors describe as having about 22 million parameters, was trained this way on 128,175 retinal images, each graded three to seven times by a panel of 54 US ophthalmologists and senior residents.

19 — A fixed structure with parameters set by data
retinal photo PIXELS IN STRUCTURE: WRITTEN AND REVIEWED LIKE ANY CODE score: 0.83 threshold REFER nobody wrote these numbers PARAMETERS: SET BY TRAINING, NOT BY HAND ResNet-50: 25,557,032 parameters (PyTorch reference implementation) compare with known answer → loss → nudge every parameter LABELED EXAMPLES e.g., 128,175 graded images, Google 2016 schematic
CodeDataReadoutFocal detail
Plate 19 — A machine-learning model is a fixed structure, written and reviewed like other code, whose parameters are found by training: scoring labeled examples, measuring the error and nudging every parameter. ResNet-50 has about 25.6 million of them, none written by hand.

Fitting, and fitting too well

Training finds parameters that fit the examples. Whether they fit anything else is a separate question, and the reason is visible in the numbers above: the network had more parameters than there were images to fit. A model with that much freedom can, in principle, memorize its training set, giving the right answer for every image it has seen and an arbitrary one for any image it has not. Engineers limit this by the choice of structure, by starting from parameters already trained on more than a million ordinary photographs, by penalties that keep weights small and by stopping training before the fit becomes too close. None of these guarantees that the model will work on new images; only a test on new images can show it.

Solved example: a perfect fit that predicts worse

Take ten points along a straight line, y equals 2x for x from 0 to 9, each moved up or down by a small fixed amount, between about −1.4 and +1.3.4. A straight line, two parameters, fitted to the ten points misses them by 0.93 on average (root mean square). A polynomial with ten parameters can pass through every point exactly, so its error on the training points is zero. Now test both on nine new points taken halfway between the old ones, from the same line with fresh noise. The straight line misses them by 0.78 on average; the polynomial misses them by 1.67, more than twice as much, because between the training points it bends to follow the noise. The model with the better training score is the worse model.

The omics guide in this library meets the same problem in another form: a molecular signature discovered from many more measurements than patients fits its discovery data almost perfectly and fails on the next cohort, and only a locked signature tested on independent samples shows whether it is real. An image network differs in scale, not in principle. Its training score is a description of the past; its value is a prediction about data it has not seen.

20 — A perfect fit that predicts worse
0 1 2 3 4 5 6 7 8 9 0 5 10 15 20 y x straight line (2 parameters): error 0.93 on training points, 0.78 on new points polynomial (10 parameters): error 0 on training points, 1.67 on new points training points new points error = root-mean-square miss fits the noise, misses the next point
CodeDataFaultFocal detail
Plate 20 — A ten-parameter polynomial fits ten noisy points perfectly and predicts new points between them worse than a straight line. A model's score on its training data describes the past; only data it has not seen shows whether it generalizes.

The data is part of the design

Because the parameters come from the examples, whatever regularities the examples contain are candidates for the model to learn. Some are the disease; others are accidents of how the data was collected: which hospital an image came from, which camera took it, what text was stamped in its corner. A model has no way to tell a sign of disease from a sign of the clinic where diseased patients were photographed, unless the data makes the difference visible. Chapter 11 shows a network that learned exactly that.

This is the sense in which the invariant of this guide extends to AI. A conventional program's faults are written into its code; a model's faults can also be written into its data, by what the data contains, what it lacks and how its labels were assigned. The device is the code, the parameters and the data that set them, together with the processing that turns a raw image into the grid of numbers the model expects. Changing any of them changes the device.

A model's behavior comes from examples nobody can fully list

Reviewers can read the code of a model's structure, but they cannot read its parameters as rules, and no list of the features a network relies on is complete. Evidence about what a model does therefore comes from its behavior on carefully chosen data, and the quality of that data, not the elegance of the code, sets the limit on what can be claimed.

Locked and learning models

A model whose parameters are fixed after training is locked: the same input always gives the same output, and it can be verified like any other software. IDx-DR's analysis was locked before the 2017 trial began, as the paper states. A model that keeps adjusting its parameters from new data in the field is continuously learning: its behavior on Tuesday may differ from Monday's, and the evidence gathered before release describes a model that no longer exists.

In 2019 the FDA wrote that the AI devices it had cleared or approved had typically "only included algorithms that are 'locked' prior to marketing." Updates happen, but as new versions: retrained, tested and released through the change controls of Chapter 21. When one of the earliest authorized plans for changing an AI device, for a program that guides cardiac ultrasound, was modified by a 510(k) in 2020, the summary stated that all algorithm modifications would be "trained, tuned, and locked prior to release" and excluded continuously learning algorithms. A draft EU manufacturing guideline of 2025 on AI in medicines production goes further, saying that dynamic models should not be used in critical applications at all. The question of how a model may change after authorization is the second half of this guide's central question, and Part V answers it.

Treat the trained parameters as a released item

Put the model's structure, its trained parameters, the preprocessing code and an identifier of the exact training data under configuration management as one released item, with a version that changes whenever any of them does. Record the data's origin, size and labeling method alongside, so that a later reviewer can tell which data produced which version of the model.

+ What this chapter established
  • A machine-learning model is a fixed structure whose millions of parameters are set by training on examples with known answers.
  • A model can fit its training data closely and still predict poorly; only a test on data it has not seen shows whether it generalizes.
  • The training data is part of the design, so a model can learn accidents of data collection as readily as signs of disease.
  • Locked models give the same output for the same input; most authorized AI devices are locked and change only through new, tested versions.
11 — Train, tune, test

Keeping the judge away from the builder.

+ The questionWhy must the data that judges a model be kept apart from the data that built it, and how does it leak back in?

A model that recognized the hospital

In 2018 a group at the Icahn School of Medicine at Mount Sinai in New York trained a network to detect pneumonia on chest radiographs, using images from two hospital systems, with a third held out for external testing, 158,323 in all, and then examined what it had learned. Pneumonia was far more common in one source than in another: 34.2 percent of the images from Mount Sinai Hospital showed it, against 1.2 percent of those from the NIH Clinical Center. A model trained on both sources together scored an area under the curve of 0.931 on test images from the same two sources, an excellent result. Scored on each hospital's test images separately, it reached only 0.805 for Mount Sinai and 0.733 for the NIH.

The gap had a simple cause. A second network could tell which hospital system a radiograph came from in 99.95 percent of the NIH test images and 99.98 percent of Mount Sinai's, helped by cues such as the laterality markers placed on the film, including a metal token. Within Mount Sinai, the portable scanners used on wards and an inverted color scheme on emergency-department images let a network identify even the department. Ranking images by nothing more than the pneumonia rate of their hospital, ignoring every pixel of lung, gave an area under the curve of 0.861 on the pooled test set, most of the way to the model's 0.931 on the same images. Part of what looked like pneumonia detection was hospital detection. The authors reported that the cues "only became apparent to us after manual image review."

21 — A pneumonia model that recognized the hospital
0.5 0.6 0.7 0.8 0.9 1.0 AREA UNDER THE CURVE 0.5 = chance jointly trained model, pooled test set (Mount Sinai + NIH) 0.931 same model, Mount Sinai test images only 0.805 same model, NIH test images only 0.733 external test, Indiana University 0.815 ranking by hospital pneumonia rate alone, no image content 0.861 knowing only the hospital beats either site PNEUMONIA PREVALENCE MOUNT SINAI 34.2% NIH 1.2% a second network identified the hospital system in 99.95% (NIH) and 99.98% (Mount Sinai) of test images 158,323 radiographs from three hospital systems; split by patient 70/10/20 Zech et al., PLOS Medicine 2018
DataFaultReadoutFocal detail
Plate 21 — A network trained on two hospitals scored 0.931 on their pooled test images but only 0.805 and 0.733 on each separately. Ranking images by their hospital's pneumonia rate alone scored 0.861: part of what looked like pneumonia detection was hospital detection.

Three kinds of data

A model is built and judged with three separate sets of data. The training set is what the parameters are fitted to. The tuning set, which machine-learning engineers usually call the validation set, is used to make choices the training itself does not make: how long to train, how large a model to use, which version of several to keep. The test set is held back until those choices are finished and is used once, to estimate how the finished model will perform. The regulatory word validation, which Chapter 7 used for proof that a device meets users' needs, means something different, and a submission that uses both senses has to say which it means.

The separation matters because every set that influences the model stops being a fair judge of it. A model scored on its training data is scored on examples it was fitted to; a model scored on the data used to pick it is scored on the data that selected it for doing well there. Only data that played no part in building the model estimates how it will do on the next patient. The international principles of good machine-learning practice, finalized by the International Medical Device Regulators Forum in January 2025, state it as their fourth principle: training datasets are independent of test sets.

How the judge leaks into the builder

Independence fails in ordinary ways. The same patient can appear on both sides of a split: two eyes, two visits, several slices of one scan, each a separate image but all carrying the same anatomy and often the same disease. A model that has seen a patient's left eye in training has partly seen the right eye in testing. The Mount Sinai group split its data by patient, 70 percent of patients for training, 10 for tuning and 20 for testing. The FDA wrote the same concern into the rules for retinal diagnostic software in 2018, requiring that analysis of performance make no unjustified assumption that repeated samples from one patient are independent.

Other leaks are quieter. Images can be normalized using statistics computed from the whole dataset, test images included. Duplicates and near-duplicates can sit in both sets. A source can be both common in training and over-represented in testing, so that its cues are rewarded twice, as the hospital cues were. And a team can look at the test results, adjust the model, and look again, until the test set has become a second tuning set. Each of these makes the reported score higher than the score on the next patient will be.

Each look at the test set spends some of its value

A test set used once gives an unbiased estimate for the model that was finished before the look. Every later decision informed by that score, a changed threshold, a retrained model, a different preprocessing step, makes the next score on the same set more flattering. Keeping a sealed test set for the final estimate, and a fresh one for each major version, keeps the estimate honest.

22 — Split by patient, not by image
LEAKY: SPLIT BY IMAGE TRAIN TEST the model has partly seen the test SEALED: SPLIT BY PATIENT TRAIN 70% TUNE 10% TEST 20% each patient's images stay in one set OTHER LEAKS duplicates normalization using all data same site's cues in both sets tuning on the test set
DataFaultFocal detail
Plate 22 — Splitting by image lets one patient's eyes or visits appear in both training and test sets, so the test partly measures memory. Splitting by patient, as the pneumonia study did with 70, 10 and 20 percent, keeps the judge independent of the builder.

How large a test set

A test set must also be large enough that the score means something. A sensitivity is a proportion of diseased patients the model detects, and its uncertainty shrinks with the number of diseased patients in the test set, not with the total. A test set with plenty of healthy patients and few sick ones gives a precise specificity and a vague sensitivity.

Solved example: diseased patients needed for a precise sensitivity

A team expects a sensitivity of about 87 percent and wants its 95 percent confidence interval to extend no more than 5 percentage points either side. The usual approximation for a proportion gives the number of diseased patients needed: 1.96 squared, times 0.87, times 0.13, divided by 0.05 squared, which is about 174. If the disease is present in 24 percent of the people tested, about 725 participants are needed to include 174 with the disease. The IDx-DR trial was planned for at least 149 participants with the disease and 682 without, to give at least 85 percent power for its prespecified statistical tests, and it analyzed 198 and 621.

The calculation sets a floor, not a target. A test set drawn from one hospital, one camera or one population can be large and still unrepresentative, which is the subject of Part IV. Size makes the estimate precise; only the right composition makes it the estimate that matters.

Seal the test set before the model is built

Split the data by patient and, where possible, by site, before any training begins. Store the test set where the development team cannot use it for training or tuning, record who accessed it and when, and use it once for the final estimate of each version. Size it from the number of positive cases the claimed sensitivity needs, and record its composition so that a reviewer can compare it with the intended population.

+ What this chapter established
  • A pneumonia network scored far better on pooled data than on each hospital, because it had learned to recognize the hospital.
  • Training, tuning and test sets must be separate; only data that played no part in building a model estimates its performance.
  • Splitting by image instead of patient, shared preprocessing, duplicates and repeated looks at the test set all inflate the reported score.
  • Precision depends on the number of diseased patients in the test set; about 174 are needed to pin an 87 percent sensitivity within 5 points.
12 — Ground truth

What the right answer is, and who decided.

+ The questionWhen a model is scored, what is it scored against?

The answer key

For every participant in the IDx-DR trial, the right answer was decided at the Fundus Photograph Reading Center of the University of Wisconsin, from images the program never saw. Certified photographers took four overlapping stereo pairs of each retina, covering a far wider field than the program's two pictures per eye, and a scan of the macula by optical coherence tomography with at least 121 cross-sections. Three experienced readers graded the photographs on the severity scale of the Early Treatment Diabetic Retinopathy Study, without seeing the program's output or the other imaging, and the majority grade stood. A participant counted as having more than mild diabetic retinopathy if the worse eye reached level 35 or higher on that scale or showed macular edema.

That protocol is the trial's reference standard, often called ground truth: the best available judgment of the true state, against which the device is scored. The trial's headline numbers, the 87.4 percent sensitivity, the 89.5 percent specificity and the predictive values of Chapter 13, are statements of agreement with the photographic grade; adding the macular scans, in a secondary analysis, gave 85.9 and 90.7 percent, figures adjusted for the trial's enrichment. If the reference were wrong for some participants, the numbers would be wrong in ways the trial could not detect.

23 — What the device sees, and what the reference sees
THE DEVICE 2 images per eye, 45 degrees, Topcon NW400 IDx-DR analysis detected / not detected / insufficient readers never see the device's answer THE REFERENCE STANDARD (READING CENTER) OCT macular cube, at least 121 cross-sections 4 widefield stereo pairs more and different information than the device 3 masked readers secondary analysis only majority photo grade: ETDRS 35 or higher, or macular edema = more than mild DR agreement = sensitivity and specificity
SpecificationCodeDataReadoutFocal detail
Plate 23 — IDx-DR saw two photographs per eye; the reference standard used four widefield stereo pairs and a macular scan, graded by three readers who never saw the program's answer. The headline figures are agreement with the photographic grade; the macular scan fed a secondary analysis.

Experts disagree

Grading retinal photographs is skilled work, and skilled graders differ. A 2018 study by researchers at Google, with retina specialists from two clinical practices, measured how much. Three fellowship-trained retina specialists graded 1,813 photographs independently and then resolved every disagreement in adjudication sessions, producing a consensus grade for each image. Measured against that consensus, the three specialists working alone detected moderate or worse retinopathy in 74.6, 74.4 and 82.1 percent of the images that had it, while correctly clearing about 99 percent of those that did not. Three board-certified general ophthalmologists did similarly. A single expert, in other words, missed about one case in four that the adjudicated panel found.

The same study showed how much the reference matters to a model. Agreement between each grader and the consensus, measured on a weighted scale where 1 means perfect, ranged from 0.80 to 0.91 for the specialists. Tuning the network, an improved version of the Google system of Chapter 10, on a small set of adjudicated grades improved it substantially, and the improved network scored an area under the curve of 0.986 against the adjudicated reference. How good a model looks depends on which human answer it is compared with, and a model trained or tested against a single grader inherits that grader's misses.

Matching one expert can mean matching that expert's mistakes

If a model is trained on one grader's labels and tested against the same grader, a high score shows that the model has learned to agree with that person, including the cases that person gets wrong. Independent, adjudicated or multimodal reference standards cost more because they are the only way to measure accuracy rather than imitation.

24 — Experts against an adjudicated consensus
SENSITIVITY, MODERATE OR WORSE RETINOPATHY SPECIFICITY 60% 70% 80% 90% 100% retina specialist 1 74.6 99.1 retina specialist 2 74.4 99.3 retina specialist 3 82.1 99.3 ophthalmologist 1 75.2 99.1 ophthalmologist 2 74.9 97.9 ophthalmologist 3 76.4 97.5 specialists' majority 88.1 99.4 ophthalmologists' majority 83.8 98.1 algorithm 97.1 92.3 a single expert missed about 1 case in 4 Krause et al., Ophthalmology 2018 (authors include Google staff); 1,813 images reference: three retina specialists after adjudication
SpecificationCodeFocal detail
Plate 24 — Measured against an adjudicated consensus, single retina specialists and ophthalmologists detected only 74 to 82 percent of moderate or worse retinopathy, while keeping specificity near 98 to 99 percent. A model scored against one grader inherits that grader's misses.

The reference caps what can be measured

An imperfect reference makes a perfect device look imperfect. When the device is right and the reference is wrong, the disagreement is counted as the device's error; when both are wrong in the same direction, the error is invisible.

Solved example: a perfect model scored against an imperfect reference

Take 1,000 people, of whom 240 truly have the disease, the 24 percent prevalence of the IDx-DR trial. Suppose the reference standard misses 5 percent of true disease, labeling 12 of the 240 as healthy, and wrongly labels 1 percent of healthy people as diseased, about 8 of the 760. Now score a model that is always right. The reference calls 236 people diseased, 228 truly diseased and 8 not; the model flags the 228, so its apparent sensitivity is 228 out of 236, 96.6 percent. The reference calls 764 people healthy, including the 12 it missed; the model clears 752 of them, so its apparent specificity is 98.4 percent. A flawless model loses more than three points of sensitivity to the reference's own errors.

The direction of the bias depends on how the reference and the device err. If both struggle with the same hard cases, small lesions or poor images, they may agree on the wrong answer and inflate the score. This is one reason the IDx-DR trial used a reference built from more and different information than the device received: wider photographs and stereo views, with a scan of the macula for a secondary analysis, read by people who never saw the device's answer. Disagreements then reflect the device's limits rather than shared blind spots.

A reference fit for the claim

Regulators' own wording has moved on this point. The good machine-learning practice principles published by the FDA, Health Canada and the UK's MHRA in 2021 asked that reference datasets be "based upon best available methods"; the international version finalized in January 2025 asks instead that reference standards be "fit-for-purpose". One reading of the change is that the best method may be impossible or unethical for every patient, and that the right reference depends on the claim. A stroke-triage program authorized in 2018 was scored against neuroradiologists' readings, with an additional neuroradiologist breaking disagreements. A dermatology dataset assembled at Stanford for testing skin-lesion models used biopsy results, so every label was confirmed by a pathologist. IDx-DR used a reading-center protocol because its claim was a screening decision about retinopathy severity, which is exactly what that protocol grades.

Whatever the choice, the reference must be defined before the test, applied the same way to every case, independent of the device's output, and described in the labeling, so that a reader of the results knows what the device was compared with. Chapter 17 shows how to find that description in a public summary and what it tells a reader about the limits of the numbers.

Define the reference before collecting the test data

Write down the reference standard in the study protocol: who grades, with what information, using which scale, how disagreements are resolved, and how graders are kept from seeing the device's output. Measure and report the agreement among graders, plan adjudication for disagreements, and prefer a reference that draws on more or different information than the device uses.

+ What this chapter established
  • IDx-DR was scored against a reading-center reference built from wider stereo photographs and macular scans, read by three masked graders.
  • Single experts disagree with an adjudicated consensus; retina specialists working alone missed about a quarter of moderate or worse cases.
  • An imperfect reference lowers a perfect model's apparent accuracy and can hide errors that model and reference share.
  • A reference standard must be fit for the claim, fixed before testing, independent of the device and described in the labeling.
13 — Measuring performance

A threshold turns a score into a decision.

+ The questionA model outputs a score; who chooses where yes begins, and what does moving that point trade?

Four counts

Of the 819 participants whose results could be analyzed in the IDx-DR trial, 198 had more than mild retinopathy by the reference standard and 621 did not. The program flagged 173 of the 198 and missed 25; it cleared 556 of the 621 and flagged 65 who did not have the disease. Every performance figure for the trial comes from those four counts. Sensitivity, the share of diseased participants flagged, is 173 out of 198, 87.4 percent. Specificity, the share of disease-free participants cleared, is 556 out of 621, 89.5 percent. The positive predictive value, the share of flagged participants who had the disease, is 173 out of 238, 72.7 percent; the negative predictive value, the share of cleared participants who did not, is 556 out of 581, 95.7 percent.

Each figure is an estimate from a sample and carries uncertainty. The exact 95 percent confidence interval for the observed sensitivity runs from about 81.9 to 91.7 percent; for specificity, from about 86.9 to 91.8 percent. The paper's headline figures, 87.2 and 90.7 percent, are slightly different because the trial deliberately recruited extra participants with poorly controlled diabetes, and the authors corrected for that enrichment with a statistical model. The trial's statistical test asked whether the lower end of each interval cleared a floor set in advance, 75 percent for sensitivity and 77.5 percent for specificity, while the alternatives of 85 and 82.5 percent, used to size the trial, are what the FDA's summary calls the prespecified thresholds; the estimates exceeded both. A study has to clear its floor with room to spare, and a small study, with wide intervals, cannot.

25 — Four counts behind every figure
IDx-DR trial, n = 819 REFERENCE: DISEASE (198) REFERENCE: NO DISEASE (621) PROGRAM: DETECTED (238) PROGRAM: NOT DETECTED (581) TRUE POSITIVES 173 FALSE POSITIVES 65 FALSE NEGATIVES 25 TRUE NEGATIVES 556 sensitivity 173/198 = 87.4% exact 95% CI 81.9–91.7 specificity 556/621 = 89.5% exact 95% CI 86.9–91.8 PPV 173/238 = 72.7% NPV 556/581 = 95.7% depends on how common the disease is prevalence in the analyzed set 198/819 = 24%; predictive values hold only at that prevalence
SpecificationCodeFaultReadoutFocal detail
Plate 25 — Every figure from the IDx-DR trial comes from four counts: 173 true positives, 65 false positives, 25 false negatives and 556 true negatives. Sensitivity and specificity describe the device; the predictive values hold only at the trial's 24 percent prevalence.

A threshold turns a score into a decision

Most models do not output a decision. They output a score, and a threshold turns it into one: above it, refer; below it, do not. Moving the threshold trades one error for the other. A lower threshold flags more patients, catching more disease and raising more false alarms; a higher one does the reverse. The receiver operating characteristic curve, or ROC curve, plots that trade across every possible threshold, sensitivity against the false-positive rate, and the area under the curve, between 0.5 for a coin toss and 1 for a perfect separation, summarizes how well the score separates the two groups regardless of where the threshold is set.

The Google network of Chapter 10 shows the choice in practice. On one validation set of 8,788 gradable images, its area under the curve was 0.991, and its authors reported two operating points on the same curve: a high-specificity point with 90.3 percent sensitivity and 98.1 percent specificity, and a high-sensitivity point with 97.5 percent sensitivity and 93.4 percent specificity. The model was the same in both cases. Which point to ship depends on what a false alarm and a missed case each cost in the setting of use, a clinical judgment rather than a property of the model, and it must be fixed before the test, because a threshold tuned on the test set inflates the result as Chapter 11 warned.

A single summary number hides the operating point

An area under the curve describes the whole curve, including thresholds nobody will use. Two models with the same area can have very different sensitivity at the specificity a clinic needs. A claim is only interpretable with the threshold, the sensitivity and specificity at it, their confidence intervals, the prevalence in the test population and the share of cases the device declined.

26 — One model, two operating points
0 0.05 0.10 0.15 0.20 0.25 0.30 0.80 0.85 0.90 0.95 1.00 SENSITIVITY FALSE-POSITIVE RATE (1 − SPECIFICITY) high-specificity threshold: 90.3% / 98.1% high-sensitivity threshold: 97.5% / 93.4% sensitivity / specificity AUC 0.991 on 8,788 gradable images (EyePACS-1) axes zoomed: x 0 to 0.30, y 0.80 to 1.00 schematic between the marked points moving the threshold same model; the threshold is a clinical choice
ReadoutFocal detail
Plate 26 — One model gives different sensitivity and specificity depending on where its threshold is set. The Google network of 2016 reported two points on the same curve; which to use depends on what a missed case and a false alarm cost in the setting of use.

Prevalence changes what a positive means

Sensitivity and specificity describe the device; predictive values describe what its answers mean in a particular population, and they change with how common the disease is. The PCR and immunoassay guides in this library work through the same arithmetic for laboratory tests; it applies unchanged to software.

Solved example: the same program in two populations

At IDx-DR's observed sensitivity of 87.4 percent and specificity of 89.5 percent, take a population in which 24 percent have the disease, as among the trial's analyzed participants. Of every 1,000 people, 240 are diseased and the program flags 210 of them; 760 are healthy and it flags 80 of them by mistake. A positive answer is right 210 times out of 290, about 72 percent, close to the trial's 72.7 percent. Now take a population in which 7 percent have the disease: 70 diseased, of whom 61 are flagged, and 930 healthy, of whom 98 are flagged. A positive answer is now right only 61 times out of 159, about 38 percent, while a negative answer is right about 99 percent of the time. The program is unchanged; the meaning of its answers is not.

Low prevalence and a weak score together produce a heavy burden of false alarms. In 2021 researchers at Michigan Medicine tested a sepsis prediction model built into a widely used electronic health record on 38,455 hospitalizations, 7 percent of which involved sepsis. The model's area under the curve was 0.63, against 0.76 to 0.83 reported in the developer's own documentation. At the alert threshold the hospital used, it flagged 18 percent of hospitalizations, caught 33 percent of sepsis cases and missed 67 percent, and only 12 percent of the hospitalizations it flagged involved sepsis: alerting once per patient, clinicians would evaluate about eight patients to find one with sepsis. The published counts reproduce every rate: 843 of 6,971 flagged hospitalizations had sepsis.

Answers the program declines to give

A diagnostic program can also decline to answer, and how those cases are counted changes the figures. IDx-DR returned "insufficient quality" for some participants even after dilation, and its labeling directs that such patients be referred. The trial's headline sensitivity and specificity exclude them: the program gave a usable answer for 96.1 percent of participants whose reference images could be graded. A complete account therefore reports how often the device answers, as well as its sensitivity and specificity when it does. Chapter 17 recomputes IDx-DR's sensitivity with the declined cases counted, and the immunoassay guide shows the same question for a blood test with a zone that gives no answer.

Fix the operating point from the clinical costs, before the test

Decide with clinicians what a missed case and a false alarm each cost in the intended setting, choose the threshold on the tuning data accordingly, and freeze it before the test set is opened. Report sensitivity, specificity and their intervals at that threshold, the prevalence of the test population, predictive values for the populations where the device will be used, and the rate of declined answers.

+ What this chapter established
  • Sensitivity, specificity and predictive values all come from four counts; IDx-DR's trial gave 87.4 percent sensitivity and 89.5 percent specificity.
  • A threshold turns a model's score into a decision, trading missed cases against false alarms along the ROC curve.
  • Predictive values change with prevalence: the same program's positive answers are right about 72 percent of the time at 24 percent prevalence and 38 percent at 7.
  • A full account also reports how often the device declines to answer, and the operating point behind any summary number.

+ Part IV · Patients the model never met

Performance where it is used.

In two studies of children and young adults with diabetes, for whom IDx-DR is not labeled, its specificity was about 79 percent, against about 90 percent in the adults of the 2017 trial: roughly twice as many false referrals among patients without the disease. The comparison is not like for like, since the youth studies used their own reference readings rather than the adult trial's reading-center protocol, and the company's founder was an author of the later one. But the patients were different, and so was the result. A model's accuracy is a measurement made on particular people, images, references and conditions, and it moves when any of them moves. The four chapters of this part follow that movement: what happens when a model meets a different setting, how it can be accurate on average and wrong for a group, what changes when a clinician and a model decide together, and how to read a published performance claim for what it does and does not cover.

14 — Dataset shift

Tested on yesterday's patients.

+ The questionWhy does a model that passed its test do worse in another clinic, on another camera, or a year later?

Eleven clinics in Thailand

Thailand had, by one press account, about 4.5 million people with diabetes and around 200 retinal specialists to examine their eyes when Google and the Thai Ministry of Public Health began deploying a deep-learning screening system in primary care. Over eight months, researchers made regular visits to 11 clinics in the provinces of Pathum Thani and Chiang Mai, watching nurses photograph patients' eyes and interviewing them about the system. The study, presented at a human-computer interaction conference in 2020, did not measure grading accuracy. It found a model that often would not grade at all.

The system had been built to refuse images below a quality threshold, a sensible safety choice in the laboratory. In the clinics, lighting differed from room to room and images came out blurred or with dark areas, so the system marked them ungradable. A press account of the study reported that more than a fifth of images were rejected. The protocol at first sent patients with rejected images to a specialist. The researchers changed it so that eye specialists reviewed the ungradable images together with the patient's records, instead of referring everyone automatically. The authors, most of them Google staff, described the core finding as a tension between the model's requirements for image quality and the images an under-resourced setting could produce.

Google's system in a prospective cohort

A separate study in Thailand measured accuracy directly. Between December 2018 and March 2020, nine primary-care sites in the national screening program used the system in real time, with regional retina specialists over-reading every image as a safety measure. Of 7,940 people screened, 7,651 were analyzed and 31.5 percent were referred; the accuracy figures come from a smaller subset whose size the abstract does not give. For vision-threatening retinopathy, the system's sensitivity was 91.4 percent against 84.8 percent for the specialists, and its specificity was 95.4 percent against their 95.5 percent, according to the 2022 paper, whose authors include Google staff and whose funding came from Google and a Thai hospital.

Taken together, the two studies show one common shape of real-world shift. In the cohort study the model's accuracy matched or beat the specialists' over-reads. The losses came outside the grading model, in the quality gate and the workflow around it: lighting, image quality and the referrals that followed a rejected image. A test on curated images measured one link of a longer chain, and the links it did not measure decided much of what patients experienced.

27 — Where the losses were in Thailand
patient arrives nurse photographs the eye image quality check model grades the image result and referral LOSSES OUTSIDE THE GRADING MODEL the lab threshold met real lighting many images marked ungradable in real clinic lighting (human-centered study, 11 clinics, 2020) a press account: more than a fifth rejected prospective cohort: sensitivity 91.4% vs 84.8% for retina specialists; specificity 95.4% vs 95.5% (prospective cohort, 9 sites, 7,651 analyzed, 2022) protocol at first referred every rejected image; changed so specialists review them with records authors of both studies include Google staff
CodeFaultReadoutFocal detail
Plate 27 — In a prospective Thai cohort the model matched or beat retina specialists in accuracy, but real clinic lighting produced many ungradable images and every rejection became a referral. The losses lay in the links a laboratory test did not measure.

Kinds of shift

A model is fitted and tested on one distribution of cases: patients, images and labels in particular proportions. Dataset shift is any difference between that distribution and the one the model meets in use, and it comes in recognizable kinds. Input shift, also called covariate shift, changes what the model sees: a different camera, different lighting, a different scanner protocol, an older or younger population. Prevalence shift changes how common the disease is, which Chapter 13 showed changes what a positive answer means. Concept shift changes what counts as the right answer: a new grading guideline, a different reference. And acquisition shift hides inside the processing chain: a software update on a scanner, a new image format, a change in compression, invisible to a person looking at the image and possibly visible to a model.

In the same study as the pneumonia network of Chapter 11, a network trained only on Mount Sinai's images met a shift between hospital systems: its area under the curve fell from 0.802 on Mount Sinai's test data to 0.717 on the NIH's. Time produces shifts too. Populations change, treatments change the appearance of disease, and clinical practice changes which patients are tested at all. A model tested once describes the cases it was tested on, at the time it was tested.

Solved example: a lower specificity in a population outside the label

In the adult trial, IDx-DR cleared 89.5 percent of participants without the disease, so among every 1,000 such adults it would refer about 105 unnecessarily. In the youth studies its specificity was about 79 percent, so among every 1,000 young people without the disease it would refer about 210 unnecessarily, twice as many. Sensitivity in one of those studies was reported as about 86 percent, close to the adult figure. A screening program that used the device outside its labeled age range would see its specialist clinics fill with false referrals, a cost that no adult study could have shown.

28 — Four ways the data can move
INPUT SHIFT schematic in use tested on IMAGE BRIGHTNESS new camera, lighting, age group IDx-DR in youth: specificity about 79% vs 90% PREVALENCE SHIFT tested on 24% in use 7% SHARE OF PATIENTS WITH THE DISEASE the same model, different predictive values (Chapter 13) CONCEPT SHIFT tested on in use THE SAME IMAGES, ORDERED BY SEVERITY new grading guideline or reference ACQUISITION SHIFT scanner software version format JPEG → PNG model JPEG vs PNG, scanner update: invisible to people, maybe not to a model the quietest shift
SpecificationCodeDataFaultReadoutFocal detail
Plate 28 — A model meets a different distribution of cases when its inputs, the prevalence of disease, the definition of the right answer or the acquisition chain change. Acquisition shifts, such as a new image format, are invisible to people and may change what a model sees.

External and prospective tests

Two kinds of test reach further than a held-out split. An external test uses data from sites, devices or periods that contributed nothing to development, and it shows how the model travels. A prospective test runs the device forward in its intended setting, on patients enrolled for the purpose, with the workflow and the operators it will really have. The IDx-DR trial was prospective in primary care, which is why its figures carry more weight than a retrospective test on archived images would. The good machine-learning practice principles ask for testing that demonstrates performance under clinically relevant conditions, and both kinds answer that request.

One external hospital is a sample of one

A model that holds its accuracy at a second hospital has passed one more test, not the test of every hospital. Sites differ in equipment, protocols, populations and workflow at once, and a single external result cannot separate those effects. Several external sites, chosen to span the intended use, say far more than one large one.

Neither removes the problem, because the next site is always new. What an external and prospective test can do is show that performance survives the kinds of difference the intended use contains, and name the conditions it was tested under, so that a user can see when a site falls outside them. Chapter 19 describes the check a hospital can run before it relies on a device, and Chapter 20 the monitoring that continues afterwards.

Describe the conditions of the test as part of the claim

Record, for every test set, the sites, devices and their software versions, image formats, acquisition protocols, operators, population characteristics and dates. State them in the labeling as the conditions under which performance was shown, and give users a way to compare their own setting with them before use.

+ What this chapter established
  • In Thai clinics a retinal model rejected many images taken in real lighting, and the workflow around those rejections decided much of what patients experienced.
  • In a prospective Thai cohort, the same kind of system matched or beat retina specialists' accuracy, so much of the loss lay outside the model.
  • Dataset shift changes inputs, prevalence, the right answer or the acquisition chain; IDx-DR's specificity fell from about 90 to 79 percent outside its labeled age range.
  • External and prospective tests show how far performance travels, and the conditions they cover belong in the claim.
15 — Subgroups

Accurate on average, wrong for some.

+ The questionCan a model be accurate overall and still fail one group of patients, and how would anyone find out?

A test set built to compare skin tones

Three published models for telling malignant skin lesions from benign ones had reported areas under the curve between 0.88 and 0.94 on their own test sets. In 2022 a group at Stanford tested them on a new collection, the Diverse Dermatology Images set: 656 images from 570 patients seen at Stanford clinics between 2010 and 2020, every diagnosis confirmed by a biopsy report, with images of the lightest and darkest skin tones matched by diagnosis, age, sex and date. On the whole set, the three models scored between 0.56 and 0.67. Split by skin tone, the drop was steeper for darker skin. In the figures of the group's preprint, the best of the three scored 0.72 on the lightest skin types and 0.57 on the darkest; another scored 0.61 and 0.50, no better than chance.

The published paper adds two findings that sharpen the lesson. Dermatologists reading the same images, scored against the biopsy results, were also less accurate on darker skin, which matters because dermatologists' judgments supply the labels for most training sets. And fine-tuning two of the models on images from the new set closed the gap between light and dark skin, which suggests that the original failure came largely from what the models had been trained on.

29 — Dermatology models on light and dark skin
0.5 0.6 0.7 0.8 0.9 1.0 0.4 ROC AREA UNDER THE CURVE SKIN TYPE model as published, on DDI after fine-tuning on DDI ModelDerm originally reported 0.93–0.94 DeepDerm 0.88 HAM10000 0.92 0.64 I–II 0.55 V–VI 0.61 I–II 0.50 V–VI 0.72 I–II 0.74 V–VI 0.72 I–II 0.57 V–VI 0.74 I–II 0.77 V–VI API only, not fine-tuned chance chance level on the darkest skin 656 biopsy-confirmed images, 570 patients; preprint figures (arXiv 2203.08807)
CodeDataFocal detail
Plate 29 — On a balanced, biopsy-confirmed test set, three dermatology models scored far below their reported figures, and lowest on the darkest skin types, one at chance. Fine-tuning on images from the new set closed the gap.

Averages hide groups

A single accuracy figure for a whole test set is an average over its members, weighted by how many of each kind it contains. If one group is small, the average can be high while that group's performance is poor, and nothing in the headline reveals it. A 2021 study of chest-radiograph classifiers trained on three large public datasets, together more than 700,000 images, measured how often each called a sick patient's film normal. The rate was higher for female patients, for patients under 20, for Black and Hispanic patients and for patients insured through Medicaid, and higher still where these characteristics combined, for example in Hispanic women compared with white women.

The remedy is to look. The IDx-DR trial reported its population, 28.6 percent of participants African American, 16.1 percent Hispanic, 63.4 percent White and 1.6 percent Asian, with a median age of 59, and analyzed sensitivity by age, sex, race, ethnicity, blood-sugar control, lens status and site. None had a significant effect on sensitivity; specificity was somewhat higher in participants over 65. Such an analysis cannot prove that no group does worse, but it can show that no large difference was detected among the groups present in meaningful numbers, and it tells a reader which groups were too small to say.

A wide subgroup interval is an open question

A subgroup result of 90 percent from a few dozen patients is compatible with performance far below the overall figure. Reporting the point estimate alone invites the conclusion that the group is served well. The interval, and the number of patients behind it, show whether the evidence is strong enough to conclude anything.

How many patients a subgroup needs

The precision of a subgroup estimate depends on the number of patients in that subgroup, which is usually a fraction of the whole test set. A study sized for a precise overall sensitivity may give a vague one for every group within it.

Solved example: the same 90 percent from 50 patients and from 500

A model detects disease in 45 of 50 diseased patients in a subgroup: 90 percent. The 95 percent confidence interval, by the Wilson method that the FDA's guidance on diagnostic statistics uses in its examples, runs from 78.6 to 95.7 percent. With 450 detected out of 500, the estimate is still 90 percent, but the interval narrows to 87.1 to 92.3 percent. With 9 out of 10, it spans 59.6 to 98.2 percent and says almost nothing. To show that a subgroup's sensitivity is within a few points of the overall figure, the subgroup needs hundreds of diseased patients of its own, which is why subgroups are planned when the study is designed, not discovered afterwards.

30 — The same 90 percent from 10, 50 and 500 patients
55% 60% 65% 70% 75% 80% 85% 90% 95% 100% SENSITIVITY Wilson score intervals, as in the FDA's 2007 diagnostic statistics guidance 9 of 10 59.6 98.2 45 of 50 78.6 95.7 450 of 500 87.1 92.3 90% 87.4%: overall sensitivity in the IDx-DR trial (for comparison) compatible with 79%
SpecificationReadoutFocal detail
Plate 30 — A subgroup's estimate is only as precise as the number of patients in it. The same 90 percent spans about 60 to 98 percent from 10 patients, 79 to 96 percent from 50 and 87 to 92 percent from 500.

Looking at many subgroups brings its own trap. A study that compares twenty subgroups will usually find one or two that differ by chance alone, so subgroups are named in the protocol before the data are seen, and a difference found afterwards is treated as a question for the next study rather than a conclusion. When a planned subgroup does fall short, a maker has three honest responses: add data from that group and retrain, as fine-tuning on the dermatology images did; narrow the claim, as IDx-DR's labeling does by covering only the adults its trial enrolled; or tell users plainly where performance is unproven. Each is a change to the device or its labeling, with the evidence that change requires.

Representative data, by rule

Regulators have written representativeness into their expectations. The good machine-learning practice principles of 2021 asked that clinical study participants and data sets be representative of the intended patient population, and the international version of 2025 asks that clinical evaluation use datasets representative of that population. When the FDA created the regulation for stroke-triage software in 2018, its special controls required results showing effective triage across relevant subgroups. India's 2026 guidance on medical device software goes further for AI: it asks makers to disclose the demographic and geographic composition of training, validation and test data, to justify models trained or validated outside India, and to assess bias, generalizability and robustness across relevant Indian sub-populations. The EU's AI Act adds data-governance duties for high-risk systems, which Chapter 29 places on its dated timeline.

Representative does not mean proportional. A subgroup that is rare in the population but at high risk may need to be over-sampled so that its performance can be measured at all, and the study then corrects for the over-sampling when it reports overall figures, as the IDx-DR trial did for its enrichment. What a regulator asks to see is that the groups the intended use contains were present, in numbers that support a conclusion, and that their results are reported.

Plan the subgroups when the study is designed

List the subgroups the intended population contains, including skin tone, sex, age, ethnicity, disease severity, device model and site, and size the study so that each group whose performance matters has enough diseased and healthy patients for a useful interval. Report every planned subgroup with its interval, and say plainly which groups were too small to assess.

+ What this chapter established
  • Dermatology models that scored 0.88 to 0.94 on their own test sets fell to 0.50–0.57, at or near chance, on the darkest skin in a balanced, biopsy-confirmed set.
  • An overall figure can hide poor performance in a small group; chest-radiograph classifiers missed disease more often in several under-served groups.
  • A subgroup estimate is only as precise as the number of patients in it: 90 percent from 50 patients spans about 79 to 96 percent.
  • Regulators in the US, the EU and India now expect data representative of the intended population and results reported by subgroup.
16 — Humans in the loop

The reader and the machine.

+ The questionWhen AI assists a clinician rather than deciding alone, what has to be measured: the model, the person, or the pair?

When the suggestion is wrong

An assistant that is right most of the time can still make an expert worse. In a study published in 2023, 27 radiologists read 40 test mammograms, each shown with a suggested category on the standard breast-imaging scale that the readers were told came from an AI system. The suggestions had in fact been set by the investigators, and 12 of the 40 were deliberately wrong. When the suggestion was right, readers of every level of experience rated about 80 percent of mammograms correctly. When it was wrong, inexperienced readers rated 19.8 percent correctly, moderately experienced readers 24.8 percent and very experienced readers 45.5 percent.

The pattern has a name: automation bias, the tendency to accept a machine's output in place of one's own judgment, which studies suggest is strongest when the person is uncertain. It turns the model's errors into the clinician's errors, and it is strongest among the people an assistant is often meant to help most, those with less experience.

31 — When the suggestion is wrong
INEXPERIENCED MODERATELY EXPERIENCED VERY EXPERIENCED 0% 20% 40% 60% 80% 100% SHARE OF MAMMOGRAMS RATED CORRECTLY correct suggestion wrong suggestion 79.7% 19.8% the model's error becomes the reader's 81.3% 24.8% 82.3% 45.5% Dratsch et al., Radiology 2023: 27 radiologists, 40 test mammograms the 'AI' suggestions were set by the investigators; 12 of 40 deliberately wrong
FaultReadoutFocal detail
Plate 31 — When a suggestion presented as AI was wrong, radiologists' accuracy fell from about 80 percent to 20 percent for inexperienced readers and 46 percent for very experienced ones. Automation bias turns a model's errors into the clinician's.
Solved example: how a model's errors pass through a reader

In the study's design, the suggestion was right in 28 of 40 cases, 70 percent. A reader who is right 79.7 percent of the time with a correct suggestion and 19.8 percent with a wrong one ends at 0.7 times 79.7 plus 0.3 times 19.8, about 62 percent overall. A very experienced reader, right 82.3 percent and 45.5 percent of the time, ends at about 71 percent. The model's error rate, multiplied by how far each reader follows a wrong suggestion, sets most of the gap between them. The study's 30 percent rate of wrong suggestions was chosen to measure the effect, not to mimic any product; a better model would make wrong suggestions rarer, and each one may be harder to spot.

Three roles for a model

Software that analyzes clinical data plays one of three roles, and each changes what has to be measured. An autonomous device gives the answer itself: IDx-DR's output is the screening decision, and no clinician interprets the images first, so the device's own accuracy is what reaches the patient. An assistive device gives a reader extra information, such as marks on suspected lesions, and the reader decides; the performance that matters is the reader's with the device compared with the reader's without it. A triage or notification device works alongside the usual workflow: the stroke-triage program the FDA authorized in February 2018 was described as a "notification-only, parallel workflow tool" that alerts a specialist to a suspected blocked artery while "trained radiologists read all images per standard of care, regardless of the performance" of the device.

The role also decides where automation bias can do harm. An autonomous device has no human to bias, but no human to catch its errors either. A parallel triage tool is designed so that its misses do not remove the usual read, although its alerts may still change what gets read first. An assistive device places its output directly in front of the reader at the moment of decision, which is where the mammography study showed the risk. The FDA's revised guidance on clinical decision support, issued in January 2026, places the same concern in its criteria: software that a clinician cannot independently review, or that supports a time-critical decision where the clinician is likely to rely on it without review, is a device software function.

32 — Three roles for a model
AUTONOMOUS IDx-DR model refer / rescreen patient the model's error reaches the patient directly ASSISTIVE detection aid model marks the reader also sees the image reader reader's decision where automation bias acts measure the reader with and without the aid PARALLEL TRIAGE stroke alert, 2018 model alert to specialist radiologist reads every image per standard of care report a miss does not remove the usual read
CodeDataReadoutFocal detail
Plate 32 — An autonomous device's accuracy reaches the patient directly; an assistive device's value is the reader's performance with it; a parallel triage tool adds alerts without removing the usual read. Each role changes what must be measured.

Measuring the pair

For an assistive device, the evidence comes from a reader study. The FDA's guidance on computer-assisted detection in radiology, last revised in September 2022, describes the usual design: many readers each read many cases, with and without the device, and a "fully-crossed" design, in which every reader reads every case in both conditions, "offers the greatest statistical power for a given number of cases." Reading sessions are separated "by at least four weeks to avoid memory bias," the primary measure is usually the area under the ROC curve for readers with the aid against readers without it, and the number of readers must be justified rather than fixed by rule. Case sets may be enriched with diseased cases, the guidance says, but enrichment can change how readers behave and biases predictive values, the problem of Chapter 13 in another form.

The good machine-learning practice principles name the target directly. The 2021 version asked for focus on the performance of the human-AI team; the international version of 2025 asks that a device be assessed with a focus on human-AI interaction in the intended use environment. A model's standalone figures remain necessary, because they show what it does on its own, but for an assistive device they do not show what patients receive.

The pair can perform worse than either partner

A strong model and a capable reader can combine badly if the reader follows the model's errors and overrides its correct calls. Only a study of readers working with the device, on cases that include the model's failures at a realistic rate, shows whether the pair beats the reader alone. Standalone accuracy cannot answer that question.

Telling users what they need to know

Clinicians can resist a wrong suggestion only if they know when the model is likely to be wrong. In June 2024 the FDA, Health Canada and the UK's MHRA published guiding principles on transparency for machine-learning devices, asking that makers share the device's intended use, performance, limitations and, where it can be done, the logic behind its outputs, with the right users at the right time and in a form designed for them. For IDx-DR, the labeling states the patients, the camera and the conditions under which performance was shown, and tells users that an "insufficient quality" result after dilation may itself be a sign of disease that needs referral. Chapter 17 reads that kind of information as an outside reader would, and asks what it leaves out.

Design the display and measure the team

Decide whether the device is autonomous, assistive or parallel, and design the display for that role: show the model's confidence or the evidence behind a suggestion where it helps review, and avoid placing a suggestion where it anchors the reader before an independent look. Measure the reader with and without the device in a study whose cases include the model's errors at a realistic rate, and train users on the failure modes the study finds.

+ What this chapter established
  • In a 2023 mammography study, wrong suggestions labeled as AI cut readers' accuracy from about 80 percent to between 20 and 46 percent.
  • Autonomous, assistive and parallel triage devices place the model differently, and each role changes what has to be measured.
  • Reader studies, ideally fully crossed with a washout of at least four weeks, measure the clinician with and without the device.
  • Transparency about intended use, performance and limits gives users what they need to recognize a wrong output.
17 — Reading a performance claim

What a decision summary says, and leaves out.

+ The questionWhat should a reader look for in a software device's published performance, and what will it not tell them?

A public document with the evidence in it

The FDA publishes a decision summary for every device it grants through De Novo, and a shorter summary for most devices it clears through a 510(k). IDx-DR's is free to download. It gives the indications for use and the limitations, describes the software and its architecture, states the special controls that every later device of the same type must meet, summarizes the software documentation and the clinical study, lists the risks and their mitigations, and ends with the agency's judgment that the probable benefits outweigh the probable risks. For an outside reader, a hospital evaluating a product or an engineer studying a competitor, it is the most complete regulatory account of the evidence.

It also rewards careful reading. Its heading gives "DATE OF DE NOVO: January 12, 2018", which is the date the request was received; the FDA's database shows that the request was granted on 11 April 2018. Its table of results gives an observed sensitivity of 87.4 percent with a 95 percent confidence interval of 81.9 to 92.9 percent, but an exact interval computed from the trial's counts, 173 of 198, runs from 81.9 to 91.7 percent, which is what the device's 2022 510(k) summary prints. Such slips do not change the conclusion. They do show that a summary is a document written by people, to be read with a calculator at hand.

Recompute the figures a summary prints

Rebuilding a table's percentages and intervals from its counts takes minutes and catches transcription errors, mismatched denominators and intervals of the wrong kind. When the counts are not given, that absence is itself worth noting, because a reader then cannot check how the cases the device declined were treated.

33 — Reading a decision summary
De Novo decision summary DEN180001 (IDx-DR) DATE OF DE NOVO: January 12, 2018 this is the receipt date; granted 11 April 2018 (FDA database) INDICATIONS FOR USE AND LIMITATIONS who, where, which camera (Topcon NW400) SPECIAL CONTROLS software V&V on a hazard analysis, cybersecurity, training, change protocol, labeling CLINICAL STUDY endpoint observed 95% CI sensitivity 87.4% 81.9–92.9 sensitivity 87.4%, printed interval 81.9–92.9; exact interval from 173/198: 81.9–91.7 recompute from the counts BENEFIT-RISK probable benefits outweigh probable risks
SpecificationFaultReadoutFocal detail
Plate 33 — A decision summary is the most complete regulatory account of a device's evidence, and it repays careful reading: IDx-DR's gives its receipt date as the date of the De Novo and prints a sensitivity interval that the trial's counts do not reproduce.

Questions to put to any claim

A performance claim can be read against a short list. Who was tested: the intended population, how participants were recruited, and whether the study was enriched. Where and by whom: the setting, the operators and their training. Against what: the reference standard and who applied it, masked or not. Which endpoints, and whether they were fixed before the study. How many: the counts behind each percentage, and the intervals. For which groups: subgroup results and the groups too small to assess. On what inputs: the camera or scanner, its software, the image format, the version of the device's own software. What happened to the cases it declined. And who ran the study: its funding, and whether the authors work for the maker.

IDx-DR's record answers most of these well. The population was adults with diabetes not previously diagnosed with retinopathy, partly enriched with participants whose diabetes was poorly controlled. The setting was ten primary-care practices with operators who had never imaged an eye, and the reference was a masked reading-center protocol. The endpoints were set in advance with input from the FDA, the counts are published, results by subgroup are reported, and the camera and software version are named. The paper also discloses that its first author was a shareholder, director and employee of IDx, the funder, which by the company's account he founded; another author held shares, and the statistician was paid by the company. What the record cannot say is how the device performs outside those conditions, which is why Chapter 14's youth studies and Chapter 20's monitoring matter.

Counting the cases it declined

The trial's headline sensitivity and specificity cover the 819 participants for whom both the reference standard and the program gave a usable result. Another 33 participants had gradable reference images but received "insufficient quality" from the program, and 10 of those 33 had the disease. How they are counted changes the result.

Solved example: three ways to count the declined cases

Leaving the 33 out, as the headline figures do, the program flagged 173 of 198 diseased participants: 87.4 percent. Counting the 10 diseased participants among the 33 as misses, it flagged 173 of 208: 83.2 percent, with an exact 95 percent interval of 77.4 to 88.0 percent, still above the trial's floor of 75 percent. Counting them as referrals, as the labeling directs for a persistent insufficient-quality result, the program sent 183 of 208 diseased participants onward: 88.0 percent. For the 23 declined participants without the disease, the same choice moves specificity from 89.5 percent to 86.3 percent if they are counted as false referrals. A full claim states which convention it uses and gives the others, as the immunoassay guide in this library recommends for a blood test with an indeterminate zone.

34 — Three ways to count the declined cases
900 enrolled 892 completed 852 with a reference grade 819 analyzable: 198 with disease, 621 without 40 ungradable by the reading center set aside 33 'insufficient quality' from the program 10 with disease, 23 without how these are counted changes the answer A declined cases left out (headline) 173/198 = 87.4% B declined diseased cases counted as misses 173/208 = 83.2% exact 95% CI 77.4–88.0 C declined cases counted as referrals (as labeling directs) 183/208 = 88.0% 70% 75% 80% 85% 90% 95% SENSITIVITY trial floor: 75% A B C specificity: 89.5% with declined cases left out; 86.3% if the 23 without disease count as false referrals
SpecificationFaultReadoutFocal detail
Plate 34 — IDx-DR declined 33 participants with gradable reference images, 10 of whom had the disease. Its sensitivity is 87.4, 83.2 or 88.0 percent depending on whether those cases are left out, counted as misses or counted as referrals.

What a summary leaves out

Some things are absent by design. A decision summary rarely describes the training data, the model's structure or how its threshold was chosen; it reports the clinical evidence for the locked device, not how the device was built. It describes performance at authorization, not in use: IDx-DR's summary cannot report how the program has done in the clinics that bought it. And it describes the device as authorized, for the inputs named in it; it says nothing about cameras or populations outside the label.

Other gaps come from the difference between a regulator's words and a seller's. The FDA's press release of 2018 called IDx-DR the first device to give "a screening decision without the need for a clinician to also interpret the image or results"; the maker's materials speak of a diagnosis and of a device "De Novo-cleared", though a De Novo is granted, not cleared. A developer's own performance figure can differ widely from an independent one: Chapter 13's sepsis model was reported by its developer at an area under the curve of 0.76 to 0.83 and measured by an outside team at 0.63. The rule that follows is simple. A claim is read in the regulator's record and the peer-reviewed paper, with the maker's own statements labeled as such, and an independent evaluation is worth more than any of them.

Write a one-page claim sheet before relying on a device

Before buying, deploying or competing with a software device, fill in a single page: intended population, setting, operators, reference standard, prespecified endpoints, counts and intervals, subgroups, inputs and versions, the handling of declined cases, the study's funding and authors, and any independent evaluation. Mark each answer with its source, and treat any blank line as a question for the maker or a reason for a local test.

+ What this chapter established
  • A decision summary is the most complete regulatory account of a device's evidence, and it should be read with a calculator, as its dates and intervals show.
  • A claim is read against who was tested, where, against what reference, with which endpoints, counts, subgroups, inputs and sponsors.
  • How declined cases are counted changes the result: IDx-DR's sensitivity is 87.4, 83.2 or 88.0 percent depending on the convention.
  • Summaries leave out training data and real-world performance, and a maker's wording and figures need labeling and independent checks.

+ Part V · Into the hospital, and after

Deploying, watching and changing a model.

On 10 June 2021 the FDA cleared a new version of IDx-DR that could receive images in DICOM, the format hospitals use for medical images, put its instructions on screen step by step, let clinics configure file names and keep images locally, and told operators during the exam whether an image was good enough to send. None of these changes touched what the program claimed to detect, and the clearance relied on no new clinical data. They were about fitting into clinics. A year later a second clearance replaced one of the models inside the analysis. The four chapters of this part follow a model from authorization into use: connecting it to a hospital's systems, checking it on the hospital's own patients, watching it once it runs, and changing it without losing the evidence. By the end of the part, the central question has its answer.

18 — Not plug-and-play

Fitting the software into a hospital.

+ The questionAn authorized AI device arrives at a hospital. What has to be connected, configured and agreed before it is used on the first patient?

The label on every image

Every image a hospital's scanners and cameras produce in the DICOM format carries a header: a list of tagged fields describing where the image came from. The field tagged 0008,0060 records the modality, such as CT or ophthalmic photography; 0008,0070 the manufacturer of the equipment; 0008,1090 the model name; 0018,1020 the software versions running on it; 0008,1030 a description of the study; and 0010,0020 the patient's identifier. The DICOM standard, revised several times a year, defines thousands of such fields; a 2026 release was current in October 2026.

For an AI device the header is the first line of defense. A device authorized for images from one camera can read the manufacturer and model fields and refuse anything else, and it can record the software version of the equipment that produced each image, so that a later change in performance can be traced to an update on the scanner. Many fields are filled in by local protocols or typed by staff, however, and the same examination can be described differently in two hospitals. A device that routes images by free-text descriptions will meet descriptions its makers never saw.

35 — What a DICOM header tells a device
image (pixel data) schematic DICOM HEADER 0008,0060 Modality e.g. CT, OP (ophthalmic photography) 0008,0070 Manufacturer equipment maker 0008,1090 Manufacturer's model name camera or scanner model 0018,1020 Software versions equipment software version 0008,1030 Study description free text: differs between hospitals 0010,0020 Patient ID local identifier typed by staff or set by local protocol: routing on it meets descriptions the makers never saw the device can check what it was tested on AI DEVICE: INPUT CHECK specified inputs (e.g. Topcon NW400) inside specification: analyze outside: refuse, with a message the user can act on logged with each result: trace later performance changes to an equipment update DICOM defines thousands of such fields; a 2026 release current in October 2026
SpecificationCodeDataFaultReadoutFocal detail
Plate 35 — The header of every DICOM image says which equipment and software produced it. A device can use it to refuse inputs outside its specification and to trace changes in performance to equipment updates, but free-text fields such as the study description vary from hospital to hospital.

Where the software sits

An AI device rarely stands alone. An imaging model receives studies from the picture archiving and communication system, the PACS, after the scanner sends them there; it returns its results to the PACS, to the radiologist's worklist, to the electronic health record or to a specialist's phone; and it may run on a server in the hospital or in the maker's cloud. Each of those connections is an interface with its own standard. DICOM carries images and image-based results. HL7 version 2, which its standards body says is used by 95 percent of US healthcare organizations, carries orders and reports as messages. FHIR, HL7's newer standard built from web resources that can be addressed individually, has been at release 5 since March 2023; release 6 was in its second normative ballot in July 2026.

Standards for the AI step itself are younger. The IHE initiative, which writes profiles that tell vendors how to combine standards for a task, has published one for AI results, which defines how analysis results are encoded in DICOM objects and displayed, and one for AI workflow, which defines how a request for analysis is sent, managed and performed. In October 2026 both were still at trial implementation, the stage before final text. A hospital connecting several AI products therefore often builds part of each integration itself, or buys a platform that does it.

Inputs inside and outside the specification

The authorization defines the inputs. IDx-DR is indicated for use with the Topcon NW400, and its 2021 summary sets a minimum image resolution of 22 pixels per degree. Inside those limits the trial's evidence applies; outside them it does not, whatever the images look like. A device has two choices when an input falls outside its specification: refuse it, with a message the user can act on, or analyze it anyway and risk an answer the evidence never covered. IDx-DR refuses, with its "insufficient quality" output and, since 2021, with feedback during the exam so that the operator can retake the image.

Inputs also drift after installation. A hospital changes a scanner protocol to reduce dose, a vendor updates the scanner's reconstruction software, the PACS begins compressing images to save storage, or a new interface converts images to a different format on the way to the model. Each change can be invisible to a person reading the images and plain to a model, the acquisition shift of Chapter 14. Chapter 19 describes a hospital study in which a change of image format, more than any change in the patients, took a model from excellent to useless.

A change on the hospital's side can change the device's inputs

A new scanner protocol, a reconstruction update, compression in the archive or a format conversion in an interface engine can move a model's inputs outside the conditions it was tested under, without any change to the device itself. Agreeing in advance who notifies whom of such changes is part of installing the device.

Solved example: where the minutes go in stroke triage

The stroke-triage program authorized in 2018 was compared with usual care, retrospectively, in 44 cases with a confirmed blocked artery that it had flagged correctly: the specialist was notified a median of 5.6 minutes after the CT angiogram, against 51.5 minutes when notification came through the radiology report. In a 2023 randomized trial of the same company's software at four Houston stroke centers, involving 243 treated patients, the time from arrival to the start of clot removal fell by an estimated 11.2 minutes from a median of 100, about 11 percent. The two studies measured different intervals in different hospitals, so the numbers cannot be subtracted, but together they show the shape of the problem: the model shortens one link, and transfer of the images, the alert reaching the right person, the team assembling and the procedure room being ready decide how much of the gain reaches the patient. The trial found no significant difference in patients' functional outcome, an exploratory measure.

36 — Where the minutes go in stroke triage
CT angiogram images transferred AI analysis alert reaches specialist team assembles procedure room ready clot removal starts the link the software shortens the rest of the chain decides how much of the gain reaches the patient 2018 COMPARISON 44 cases with a confirmed blocked artery time from CT angiogram to specialist notified (median) through the radiology report: 51.5 min with the program: 5.6 min 0 10 20 30 40 50 60 min 2023 RANDOMIZED TRIAL 4 Houston stroke centers, 243 treated patients arrival to start of clot removal (median) usual care: 100 min with the software: estimated 11.2 min shorter 0 10 20 30 40 50 60 70 80 90 100 110 min functional outcome: no significant difference (exploratory) different intervals in different hospitals: the numbers cannot be subtracted
CodeDataReadoutFocal detail
Plate 36 — A triage model shortens one link in the chain from scan to treatment. In a 2018 comparison notification came at a median of 5.6 minutes against 51.5; in a 2023 trial the time to clot removal fell by an estimated 11.2 minutes from a median of 100. Transfer, the team and the procedure room decide how much of the gain reaches the patient.

Who answers for what

The maker answers for the device within its specified environment, and the EU's medical device regulation makes it state that environment: manufacturers "shall set out minimum requirements concerning hardware, IT networks characteristics and IT security measures" needed to run the software as intended. The hospital answers for the network, the interfaces, the other systems and the way the device is used. Between them sits the integration, which neither controls alone.

The standard for that middle ground is IEC 80001-1, whose 2021 edition sets requirements for organizations applying risk management before, during and after connecting a medical device or health software to their IT infrastructure, covering safety, effectiveness and security. It is already marked for revision, and ISO intends to replace it with a new standard in the 81001 series. An earlier technical report in the same family describes responsibility agreements: written statements of which party, the hospital, the IT supplier or the device maker, does what across the life of the connection. The questions such an agreement must settle are concrete: who tests the interfaces, who approves changes on each side, who is told when a scanner or the archive is updated, and what clinicians do when the AI is unavailable.

Treat installation as a project with its own risk file

Before the first patient, list every interface the device depends on and test each with real local data, and confirm that the local images and their headers fall inside the device's specification. Measure the time from acquisition to result against the clinical need. Define what clinicians do when the device is down or declines an input, and sign a responsibility agreement that names who notifies whom of changes on either side.

+ What this chapter established
  • DICOM headers identify the equipment and software behind every image, letting a device refuse inputs outside its specification and trace changes.
  • AI devices depend on interfaces to PACS, health records and cloud services; the IHE profiles for AI were still trial implementations in October 2026.
  • A model's evidence covers only inputs inside its specification, and hospital-side changes can move inputs outside it without touching the device.
  • The maker states the required environment, the hospital manages the connection under IEC 80001-1, and a responsibility agreement divides the rest.
19 — Validating it where it runs

Site acceptance, local performance and computerized system validation.

+ The questionThe maker validated the software; why does the site have to validate it again, and how much is enough?

Two validations, two questions

A device that arrives at a hospital has been validated once already, by its maker, and validated again would seem redundant. It is not, because the two validations answer different questions. The maker's validation shows that the device meets its intended use in its specified environment, on the patients and inputs of its studies. The hospital's question is whether that evidence carries over to this hospital: its scanners and their settings, its interfaces, its patients and its way of working. Only the hospital has the data to answer it.

Professional societies now say so plainly. In a joint statement published in January 2024, the radiology societies of the United States, Canada, Europe, Australia and New Zealand called testing on local data, with local systems and workflows, "essential", and recommended that each site "perform a statistically rigorous evaluation of performance on their own local data." Where that is not feasible, they advised comparing local data with the vendor's test data and, where the two differ, proceeding "with great caution, if at all."

A trial nobody saw

In August 2020 a team at the Hospital for Sick Children in Toronto began running a deep-learning model on live clinical data with its outputs hidden from clinicians, an approach the team calls a silent trial. The model predicted, from renal ultrasound images, which children with swelling of the kidney would need surgery. On a random 20 percent held out from its development data, 1,643 kidneys from 294 patients in all, its area under the curve had been 0.90. In the silent trial, on 523 kidneys from 150 patients seen between August and December 2020, it was 0.50: no better than chance.

The team traced the collapse to three differences. The live patients were younger, the share of obstructed right kidneys was higher, and the images reached the model in a different form: the development data had been processed JPEG files, while the live data arrived as unprocessed PNG files that looked different to the model even after the same preprocessing. Adjusting for age and side barely helped, raising the area to 0.51. Reprocessing the live images to match the original pipeline raised it to 0.84 to 0.85. After retraining on the original and silent-trial data together, a second silent trial on 711 kidneys from 202 patients gave 0.91 to 0.92. The decisive fault was not clinical at all. It lay in the integration chain of Chapter 18, and only running the model on the hospital's own data, before anyone acted on it, revealed it.

37 — A model that fell to chance and came back
0.5 0.6 0.7 0.8 0.9 1.0 AREA UNDER THE CURVE chance 0.90 random 20% held out (development data) of 1,643 kidneys, 294 patients 0.50 silent trial, Aug–Dec 2020 523 kidneys, 150 patients 0.51 adjusted for age and side 0.84–0.85 live images reprocessed to match the original pipeline 0.91–0.92 retrained; second silent trial 711 kidneys, 202 patients the main fault was in the image format, not the patients development: processed JPEG files ≠ live: unprocessed PNG files looked different to the model after the same preprocessing Hospital for Sick Children, Toronto; outputs hidden from clinicians
DataFaultReadoutFocal detail
Plate 37 — Run silently on live data in Toronto, a model that had scored 0.90 fell to 0.50, no better than chance. Adjusting for the patients barely helped; reprocessing the live images to match the development pipeline restored most of it, and retraining brought it to 0.91 to 0.92.

Acceptance and a local check

A site's own validation has two layers. The first is acceptance: confirming that the device is installed as specified, that each interface passes the right data, that results reach the right screen for the right patient, and that the fallback works when the device is unavailable. The second is a local performance check: running the device on a sample of the site's own cases, silently or retrospectively, and comparing its outputs with a local reference. The multisociety statement suggests focusing that review on the predictive values clinicians will experience, identifying the cases where the device would change care, and categorizing its false positives and negatives before deciding whether to deploy. It recommends repeating the review whenever the AI software or the equipment used with it changes.

Solved example: how many positive cases a local check needs

A hospital wants evidence that a device's sensitivity on its patients is above 75 percent, and expects it to be about 87 percent, as in the device's studies. The standard calculation for comparing a proportion with a fixed value, at one-sided 5 percent significance and 80 percent power, gives about 69 diseased cases. If 10 percent of the hospital's screened patients have the disease, that means about 690 consecutive cases. With 60 detections out of 69, 87 percent, the exact one-sided lower confidence bound is about 78 percent, above the target. A check run on a few dozen convenient cases cannot reach that conclusion, which is why the societies call for statistical rigor rather than a demonstration.

The results of the check become the baseline for the monitoring of Chapter 20: the distribution of inputs, the rate of declined cases, and the local sensitivity, specificity and predictive values against which later drift is measured.

Computerized system validation

Some sites must go further, because the rules they work under require it. Laboratories, blood services, organizations running clinical trials and pharmaceutical manufacturers operate under good-practice rules, collectively called GxP, that require any computerized system affecting product quality, patient safety or data integrity to be validated by its user. The method most of them follow is ISPE's GAMP 5, whose second edition was published in July 2022. It is risk-based: the effort scales with the system's impact, complexity and novelty, and its authors describe it as placing patient safety and product quality ahead of compliance for its own sake and supporting iterative and agile methods.

GAMP sorts software into categories: infrastructure software such as operating systems and databases, and then a continuum from products used without configuration through configured products to custom software. The more a system is configured or customized for the site, the more the site must verify itself; for a standard product, the user leans on the supplier's evidence after assessing the supplier. The work follows a chain. User requirements say what the site needs, and a risk assessment shows where failure would matter. The system is then specified and configured, and verified, traditionally as installation, operational and performance qualification and increasingly as risk-based scripted and unscripted testing. A traceability matrix links requirements to tests, release is controlled, and the system is reviewed periodically for as long as it is used. Data integrity runs through all of it: records that are attributable, complete and protected, with audit trails, as the FDA's rule on electronic records and signatures, 21 CFR Part 11, has required since 1997.

38 — Computerized system validation, GAMP 5
infrastructure software (operating systems, databases) products used without configuration configured products site configuration custom software lean on the supplier's evidence, after assessing the supplier more configuration and customization → more the site verifies itself outside the continuum user requirements risk assessment specification and configuration verification: installation, operational and performance qualification, or risk-based scripted and unscripted testing traceability matrix release periodic review effort scales with impact, complexity and novelty data integrity: attributable, complete, protected records with audit trails (21 CFR Part 11, since 1997) GAMP 5, second edition (ISPE, July 2022); GAMP guide to AI (July 2025)
SpecificationCodeDataFocal detail
Plate 38 — GAMP 5 scales a site's validation to risk: the more a system is configured or customized, the more the site verifies itself, and the chain from user requirements to periodic review rests on records that are attributable, complete and protected.
A vendor's validation package is a starting point for the site

Supplier documentation lets a site avoid repeating tests that do not depend on its environment. It cannot cover the site's interfaces, configuration, data or patients, and a package accepted without assessing the supplier, or without local performance data, leaves the site's own question unanswered.

AI has added to the method. The 2022 edition of GAMP 5 included an appendix on machine learning with a life cycle for an ML subsystem, and in July 2025 ISPE published a separate GAMP guide to AI, covering data governance, model testing, monitoring and change management for AI-enabled systems used in GxP work. In July 2025 the European Commission also put out for consultation a new annex on AI to the EU's good manufacturing practice rules. It would admit only static, deterministic models in critical manufacturing uses, would not apply to generative AI and large language models, and would require acceptance criteria at least as high as the performance of the process the model replaces. As of October 2026 the annex was still a draft. It governs medicines manufacturing, not medical devices, but its test-data rules read like a summary of Chapters 11 to 13.

The builder's standard and the user's

IEC 62304 and GAMP 5 describe the same software from opposite sides. IEC 62304 is written for the organization that builds device software, and its records show how the device was made. GAMP 5 is written for the organization that uses a computerized system, and its records show that the system does what that user needs, in that user's configuration. A laboratory working under good-practice rules and installing an AI-enabled analyzer, or a contract research organization using an AI reading tool in a trial, needs both: the maker's evidence, and its own validation built on it. Makers can shorten the second by supplying a validation package with their product: requirements and traceability, test scripts the site can re-run, and installation and operational qualification documents.

Plan the site validation before the contract is signed

Write the site's requirements and acceptance criteria, including the local performance check and its sample size, before purchase, and ask the supplier for its validation package and its specified environment. Where GxP rules apply, scope the work with GAMP 5's categories and risk assessment; in every case, run the device silently or retrospectively on local cases before clinicians rely on it, and keep the results as the monitoring baseline.

+ What this chapter established
  • The maker's validation covers the intended use in a specified environment; only the site can show that the evidence carries over to its systems and patients.
  • A silent trial in Toronto took a model from 0.90 to 0.50, to 0.85 after fixing an image-format mismatch in the integration and to 0.92 after retraining.
  • A local check needs enough diseased cases to support a conclusion, about 69 to show a sensitivity above 75 percent, and becomes the monitoring baseline.
  • Where GxP rules apply, sites validate AI systems with GAMP 5's risk-based method, now extended to AI, building on the supplier's evidence.
20 — Watching a model in the field

When the true answer arrives late.

+ The questionHow can a maker, or a hospital, tell a model is still working when the right answer arrives months later, or never?

Truth arrives late

A model can be wrong for months before anyone could know. When a screening program tells a patient to come back in a year, nobody checks whether that answer was right until the patient returns, if they do. When it refers a patient, the specialist's examination confirms or overturns the answer weeks later, in another clinic's records. For most outputs of most diagnostic models, the reference standard of Chapter 12 is never applied at all in routine use. The direct measure of performance, agreement with the right answer, is available late, partially and only for some patients.

The early months of the COVID-19 pandemic showed how fast the ground can move meanwhile. A study of 24 hospitals in four US health systems found that daily alerts from a widely used sepsis prediction model rose by 43 percent while the hospitals' patient census fell by 35 percent, ahead of the surge in COVID admissions. The patients had changed, and the model's inputs with them. At Michigan Medicine, one of the four systems, the flood of alerts led to the alerts being paused, according to the university's account of the study. Nobody needed to wait for confirmed diagnoses to see that something had shifted: the volume of the model's own output showed it.

What can be watched at once

Monitoring therefore starts with what is available immediately. The inputs can be checked as they arrive: the equipment and software versions in the image headers of Chapter 18, image quality measures, patient age and other characteristics the device records. The outputs can be counted: the share of positive results, the share of declined cases, the distribution of scores. None of these says directly whether the model is right. Each can show, within days, that the conditions it was validated under no longer hold, which is the trigger to look harder.

Solved example: the positive rate as an early warning

Suppose a screening device with 87.4 percent sensitivity and 89.5 percent specificity runs at a site where 10 percent of patients have the disease. It should return a positive result for about 18.2 percent of patients: 8.7 points from the diseased and 9.5 from false alarms among the healthy. At 200 screens a week, that is about 36 positives a week, with a typical week-to-week spread of about 5. If a camera fault lowers specificity to 80 percent, the expected positive rate rises to about 26.7 percent, about 53 positives a week, more than three spreads above the baseline. A simple chart of weekly positives would flag the change within a few weeks, months before referral outcomes could confirm that the extra positives were false.

39 — The positive rate as an early warning
0 10 20 30 40 50 60 70 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 WEEK POSITIVE RESULTS PER WEEK (200 SCREENS) expected 36 (18.2%) after a camera fault: expected 53 (26.7%) baseline + 3 spreads = 51 camera fault (specificity 89.5% → 80%) flagged within a few weeks schematic: expected levels from the worked example; real weeks scatter around them DRIFT THE COUNTS CAN MISS discrimination (AUC) maintained calibration predicted risk observed rate time → predicted risk overstates schematic, after a 9-year study of kidney-injury models
FaultReadoutFocal detail
Plate 39 — At 200 screens a week, a fall in specificity from 89.5 to 80 percent lifts expected positives from 36 to 53, more than three spreads above baseline, so a weekly chart flags it long before referral outcomes arrive. Calibration can drift while the area under the curve holds.

Some drift hides from counts. A 2017 study at US Veterans Affairs hospitals followed seven models predicting acute kidney injury, regression and machine-learning models alike, for nine years after their development data. All kept their ability to rank patients by risk; their discrimination was maintained. Their calibration drifted: every model came to overstate the risk as the rate of kidney injury changed, and the overprediction grew steadily for the regression models while staying small for the random forest and the neural network. A model used with a fixed threshold on predicted risk would have alerted more and more often for the same patients, with a stable area under the curve throughout.

A stable area under the curve can hide a drifting threshold

Discrimination measures whether a model ranks sick patients above healthy ones; calibration measures whether its scores mean what they say. Monitoring that tracks only the area under the curve can miss a drift that moves every decision made at a fixed threshold, which is the drift that changes alert rates and referrals.

Closing the loop with outcomes

The direct measures follow. Referred patients' specialist findings can be collected and matched to the device's outputs, giving a running estimate of the positive predictive value. A random sample of negative results can be sent for a reference reading, the only way to estimate missed cases before patients return with them. Complaints, adverse event reports and service records feed the maker's postmarket surveillance. The multisociety radiology statement of 2024 recommends re-evaluating each AI tool on updated local data "at specified intervals, but at least annually," and after any new version, noting that "naturally occurring data drift will cause AI model performance to degrade over time," and it asks institutions to define in advance how a problem will be escalated and resolved.

40 — Two clocks of a monitoring plan
WATCHED CONTINUOUSLY: DAYS inputs: equipment and software versions, image quality, patient age outputs: positive rate, declined rate, score distribution CLOSED ON A SCHEDULE: WEEKS TO MONTHS reference reading of a random sample of negatives specialist findings matched to positives: predictive value complaints, adverse events, service records → maker's postmarket surveillance baseline from the site validation (Chapter 19) limit crossed? investigate restrict use suspend change the device (Chapter 21) named reviewer at the site named reviewer at the maker the only early estimate of missed cases radiology societies (2024): re-evaluate on local data at least annually and after any new version
SpecificationDataReadoutFocal detail
Plate 40 — Monitoring runs on two clocks: inputs and output rates show within days that conditions have shifted, while reference reads of sampled negatives and matched outcomes measure accuracy over months. Named people on both sides act when a limit is crossed.

Who watches

Responsibility is shared and not yet settled. The maker has postmarket obligations everywhere a device is sold, and the good machine-learning practice principles end with the expectation that deployed models are monitored for performance and that the risks of retraining are managed. The hospital sees the inputs, the workflow and the outcomes, and only it can link the device's outputs to its own patients' later records. In September 2025 the FDA opened a public discussion on how the real-world performance of AI-enabled devices should be measured and evaluated, asking about metrics, monitoring methods, data sources and the triggers for a response; it described the document as proposing no policy, and no follow-up had been published by October 2026. The EU's AI Act will add monitoring duties for hospitals that deploy high-risk AI, on the timetable set out in Chapter 29.

What a monitoring plan needs is clear even where the division of labor is not: a baseline from the site validation of Chapter 19, measures of inputs and outputs watched continuously, a reference sample and outcome matching on a schedule, thresholds that trigger investigation, and named people on each side who act on them. When monitoring shows that performance has moved, the response is a change to the device, its labeling or its use, and Chapter 21 sets out how a model can be changed without losing the evidence behind it.

Write the monitoring plan before go-live

Define the input checks, the output and declined-case rates to chart, their baseline from the local validation and the limits that trigger investigation. Schedule a reference reading of a random sample of negatives and the matching of positives to specialist findings, assign who on the maker's and the site's side reviews each, and state what happens when a limit is crossed: investigation, restricted use, suspension or retraining.

+ What this chapter established
  • For most outputs of a diagnostic model the right answer arrives late or never, so accuracy cannot be monitored directly in real time.
  • Inputs, output rates and declined-case rates show within days when conditions have shifted, as sepsis alerts did early in the pandemic.
  • Calibration can drift while discrimination holds, so monitoring must watch decisions at the threshold, not only the area under the curve.
  • Reference reads of sampled negatives, outcome matching and agreed triggers close the loop; maker and site share the work.
21 — Change control for a model

Agreeing in advance what may change.

+ The questionHow can a model be retrained after authorization without a new review every time, and what keeps the retrained model as good as the one that was authorized?

From a discussion paper to a final guidance

On 2 April 2019 the FDA published a discussion paper, explicitly not a draft guidance, proposing a new way to regulate changes to machine-learning software. Its premise was that such software improves by retraining, and that requiring a new premarket review for every improvement would either freeze models at their first version or require a new submission for each improvement. Its proposal was that a maker could describe in advance a "region of potential changes", which it called the pre-specifications, together with an algorithm change protocol: the methods it would use to make, test and control those changes. If the plan was authorized with the device, changes inside it could be made without a new submission.

The idea took five years to become final guidance. The FDA, Health Canada and the UK's MHRA published five guiding principles in October 2023: a plan should be focused and bounded, risk-based, evidence-based and transparent, and should take a total product lifecycle perspective. The FDA's final guidance on predetermined change control plans for AI-enabled device software followed on 4 December 2024 and was reissued on 18 August 2025. A broader draft guidance on such plans for all devices, issued in August 2024, was still a draft in October 2026.

What a plan contains

The final guidance asks for three things. A description of modifications specifies the planned changes and the characteristics and performance the modified device will have: for example, retraining on new data from the same kinds of sites, or adjusting a threshold within stated limits. A modification protocol describes how each change will be developed, validated and implemented: how new data will be collected and kept separate, how the model will be retrained, how its performance will be evaluated, the acceptance criteria it must meet, and how users will be told. An impact assessment documents the benefits and risks of carrying out the plan, and how those risks will be controlled.

The plan has firm limits. Modifications must stay within the device's intended use, and generally within its indications; a change of what the device is for needs a new submission. Changing the plan itself generally needs a new submission too. The device's labeling should say that it contains machine learning and has an authorized plan, and the public summary of the authorization should describe the plan, so that users and other makers can see what may change without further review. The plan can be authorized through a 510(k), a De Novo or a premarket approval.

41 — What a change control plan holds
INTENDED USE (AND GENERALLY THE INDICATIONS) AUTHORIZED PLAN v2 v3 v4 retrain on new data from the same kinds of sites adjust threshold within stated limits v5 change not in the plan → new submission changing the plan itself → generally a new submission NEW USE new intended use → new submission THE PLAN'S THREE PARTS 1 DESCRIPTION OF MODIFICATIONS what may change, and the performance the modified device will have 2 MODIFICATION PROTOCOL data collection and separation, retraining, evaluation on a sequestered test set acceptance criteria fixed in advance, overall and by subgroup how users are told fixed before any modification is made 3 IMPACT ASSESSMENT benefits, risks and their controls 2019 discussion paper 2023 guiding principles (FDA, Health Canada, MHRA) Dec 2024 final guidance (reissued Aug 2025)
SpecificationCodeFocal detail
Plate 41 — A change control plan lets a model change without a new review only inside its bounds and only through tests fixed in advance. A change outside the plan, a new intended use or a change to the plan itself goes back to the regulator.
The acceptance criteria carry the plan's weight

A plan that lists permitted changes but sets loose or self-adjusting acceptance criteria gives the regulator nothing to rely on. The criteria that matter, in this guide's view, are fixed before any modification is made: performance on a sequestered test set at least as good as the authorized version, overall and in each named subgroup, with the same reference standard. A modified model that misses them does not ship.

The thread's change

IDx-DR's own record shows a change made before the guidance existed, and what evidence it used. The 2018 De Novo already required, as a special control, "a protocol... that describes the level of change in device technical specifications that could significantly affect the safety or effectiveness of the device," an early, device-specific form of the same idea. In June 2022 the FDA cleared version 2.3, whose changes included a new classifier for judging image quality inside the analysis. To show the change did no harm, the maker re-ran the new version on the images from the 2017 trial, 850 of which could be evaluated, and compared its answers with the original version's.

Solved example: what a re-run on the original images showed

Against the reading-center reference, version 2.3 flagged 171 of 195 diseased participants it analyzed, 87.7 percent, against version 2.0's 173 of 198, 87.4 percent; it cleared 553 of 614 participants without disease, 90.1 percent, against 556 of 621, 89.5 percent. On the participants each version analyzed, both figures rose slightly. Counted over the same 850 participants, however, version 2.3 detected 171 diseased participants against 173 and cleared 553 without disease against 556: the rise came from declining more people. The new quality classifier changed which participants received an answer: after all resubmissions, version 2.3 gave a usable result for 809 of 850 participants, 95.2 percent, against version 2.0's 819, 96.4 percent; the 96.1 percent of the opening scene counted 852 participants with a reference grade. Ten more participants out of 850 were declined, 1.2 points. In a program screening 10,000 people a year, that would mean, at the trial's case mix, about 120 more patients sent for re-imaging or referral for image quality alone. The comparison was a regression test on the original patients; it showed that the new version did not do worse on them, and it could say nothing about patients the trial never included.

42 — Version 2.3 re-run on the 2017 images
SENSITIVITY v2.0 173/198 · 87.4 % v2.3 171/195 · 87.7 % SPECIFICITY v2.0 556/621 · 89.5 % v2.3 553/614 · 90.1 % USABLE RESULT AFTER RESUBMISSIONS v2.0 819/850 · 96.4 % v2.3 809/850 · 95.2 % rose on those analyzed 10 more of 850 declined: 1.2 points; about 120 more per 10,000 screened fewer answers explain the rise 80 85 90 95 100 PERCENT (AXIS STARTS AT 80) reference: reading center; 850 evaluable participants from the 2017 trial; 510(k) cleared June 2022
FaultReadoutFocal detail
Plate 42 — Re-run on the 2017 trial's images, version 2.3 looked slightly more sensitive and specific than version 2.0 on the participants it analyzed, but its new image-quality check declined ten more of 850; over the same 850 it made five fewer correct calls. A regression test on the original patients shows both effects; it says nothing about patients the trial never included.

Plans in use

The plans the FDA has authorized are now numerous enough to describe. One of the earliest was in the De Novo authorization of Caption Guidance, a program that guides ultrasound users to acquire cardiac images, in February 2020; when the plan was modified by a 510(k) that September, its summary stated that all algorithm modifications would be "trained, tuned, and locked prior to release", and continuously learning algorithms were excluded. A peer-reviewed study published in July 2026, using FDA data to April 2026, counted 170 devices across all review panels cleared with a change control plan, 34 of them radiology AI devices, 22 of those cleared in 2025; it found that the public summaries described neither ongoing performance monitoring nor predefined triggers for retraining. The insulin-dosing app UpDoc, cleared in December 2025, lists five categories of permitted change, from new insulin formulations to alternative ways of entering data, none of which alters the cleared dosing logic.

A plan changes what the maker may do without a new review. It does not change what a hospital needs to know. Each new version alters the device the site validated in Chapter 19, and a site that relies on its baseline needs to be told what changed, to re-run its local check when the change could move performance, and to reset its monitoring limits from Chapter 20.

The answer so far

Parts II to V together answer the central question. For conventional software, the evidence that it will be right for the next patient is a lifecycle: requirements derived from risk and traced to tests, an architecture that confines the dangerous code, an account of every borrowed component, verification at each level and regression after every change. For an AI model, it adds a test kept apart from training and drawn from the setting of use, against a reference fit for the claim, reported with its uncertainty and by subgroup, then checked at each site and watched in use. What keeps that evidence true after an update is change control: every change judged against the claims it could move, and, for a model, a plan agreed in advance that bounds the changes and fixes the tests each new version must pass. Part VI adds the threat that does not wait for an update.

Write the change plan with the first submission

Decide before the first submission which changes the model is likely to need: retraining on new data, new input devices, threshold adjustments. Describe them, the data they will use, the sequestered test sets and the acceptance criteria overall and by subgroup, and how users and sites will be told of each new version. Keep the plan narrow enough that each permitted change can be tested against criteria fixed today.

+ What this chapter established
  • The FDA's 2019 discussion paper led to final guidance in December 2024: a predetermined change control plan lets authorized changes proceed without a new submission.
  • A plan holds a description of modifications, a modification protocol with fixed acceptance criteria, and an impact assessment, all within the intended use.
  • IDx-DR's 2022 version re-ran the 2017 images: accuracy on analyzed participants rose slightly because 1.2 points more of them received no answer.
  • Each new version changes the device a site validated, so sites must be told, re-check performance where needed and reset their monitoring.

+ Part VI · Defending the device

Cybersecurity.

On 28 August 2017 Abbott wrote to physicians that a firmware update for several of its pacemakers, formerly made by St. Jude Medical, was ready to close vulnerabilities in the devices' radio link. The update could not be sent remotely. Each patient had to visit a clinic, where the update took about three minutes through the programmer's wand while the pacemaker ran in a backup mode; press reports of the FDA's communication put the number of affected devices in the US at about 465,000. Nothing in those pacemakers had failed. According to a cardiology society's summary of the FDA's communication, the flaws could allow intrusions, some of which could affect how the device operated. The three chapters of this part treat that kind of failure: what threats a device faces, which defenses must be designed in before release, and how weaknesses found after release are judged and fixed.

22 — Threats

An attacker is an input.

+ The questionHow can a device that meets every requirement still be made to harm a patient?

A pump that could not be patched

On 27 June 2019 the FDA warned patients and clinicians that certain Medtronic MiniMed insulin pumps were being recalled because an unauthorized person "could potentially connect wirelessly to a nearby MiniMed insulin pump and change the pump's settings." The pumps, the MiniMed 508 and several models of the Paradigm series, communicated by radio with remote controllers and glucose meters, and the protocol they used lacked proper authentication: a pump could not tell its own patient's devices from an attacker's transmitter within range. Medtronic had identified about 4,000 patients in the US who were potentially using vulnerable pumps.

The remedy was unusual. The FDA's announcement said Medtronic was "unable to adequately update" the pumps "with any software or patch," and the company offered patients replacement pumps with stronger built-in protection. The FDA said it was not aware of any confirmed patient harm. The US cybersecurity agency's advisory on the same day described the weakness as improper access control, exploitable only by an attacker nearby and requiring high skill, with no known public exploit. The recall record for the MiniMed 508, posted in March 2020, was still open in October 2026.

43 — A pump that could not tell who was speaking
RADIO RANGE MINIMED 508 / PARADIGM PUMP REMOTE CONTROLLER GLUCOSE METER commands and readings attacker's transmitter (nearby) change settings, control insulin delivery no proper authentication: both are accepted a requirement nobody wrote: refuse other senders 27 Jun 2019 FDA warning · about 4,000 US patients identified by Medtronic could not be updated with any software or patch → replacement pumps offered CISA advisory: CVSS 7.1 High, adjacent access, high skill, no known public exploit; no confirmed patient harm
ThreatFaultFocal detail
Plate 43 — The recalled MiniMed pumps met their requirements but accepted radio commands without proper authentication, so a nearby attacker could in principle change settings. The pumps could not be patched, and patients were offered replacements.

Safety risk and security risk

The pumps met their requirements. They delivered insulin as programmed, alarmed as specified and passed their verification. What they lacked was a requirement nobody had written in the form that mattered: refuse commands from anyone but the patient's own devices. The fault was built in on the day they shipped, like every software fault in this guide, and the input that reached it was not an unlucky combination of events but a deliberate one.

That difference changes how risk is judged. Safety risk management, Chapter 4's ISO 14971 chain, estimates how likely a sequence of events is. An attacker is not a sequence of events with a probability; an attacker chooses the inputs, searches for the rare combination and repeats it at will. The FDA's 2026 guidance on cybersecurity in premarket submissions therefore describes security risk assessment as non-probabilistic: it judges a vulnerability by its exploitability, how easily an adversary could use it, and by the harm that would follow, and it treats known vulnerabilities as reasonably foreseeable. Where a security weakness could lead to patient harm, the two assessments meet, and a security control becomes a safety control.

Solved example: reading a severity score

The 2019 advisory scored the pump weakness 7.1, High, on version 3 of the Common Vulnerability Scoring System, with the vector AV:A/AC:H/PR:N/UI:N/S:U/C:L/I:H/A:H. Read left to right: the attack vector is adjacent, meaning within radio range rather than over the internet; attack complexity is high; no privileges and no user interaction are needed; the scope is unchanged; the impact on confidentiality is low, but on integrity and availability high, because settings and delivery could be changed or stopped. A 2018 advisory on the same pumps' remote controllers had scored two weaknesses 4.8 and 5.3, Medium, because each affected only one of confidentiality or integrity, and one also needed user interaction. The score summarizes technical severity. It says nothing about how many patients use the device or what a wrong dose of insulin does, which is why a maker's assessment adds the clinical harm on top.

The attack surface

Every way data or commands can enter a device is part of its attack surface: radio and Bluetooth links, Wi-Fi and Ethernet, USB and serial ports, the connection to a cloud service, a service interface used only by technicians. The law that has governed connected devices in the US since March 2023 applies to a "cyber device", which, among other conditions, includes software and can connect to the internet. The FDA reads that broadly, "whether intentionally or unintentionally, through any means": Wi-Fi, cellular, Bluetooth, magnetic inductive links, USB, Ethernet and serial ports all count, and so does a USB connection used briefly for service.

A threat model lists those entry points, the assets behind them, such as the dose settings, the patient's data, the device's software itself, and the adversaries who might want them, and asks for each path what an attacker could do and what would stop them. MITRE and the Medical Device Innovation Consortium, with FDA funding, published a playbook for threat modeling medical devices in 2021, and its guidance expects the threat model to be part of the design, updated as the architecture of Chapter 5 changes. IDx-DR's 2018 special controls already required a cybersecurity vulnerability and management process, because its images and answers cross the internet between the client and the server.

44 — An attack surface and its threat model
DEVICE dose settings patient data device software AUTHENTICATION radio / Bluetooth RESTRICTED ACCESS service and debug interface DISABLED WHEN UNUSED USB / serial port ENCRYPTION Wi-Fi / Ethernet SIGNED UPDATES ONLY cloud connection nearby attacker insider with service access remote attacker even a brief USB service connection counts ENTRY POINT WHAT COULD BE READ, CHANGED OR STOPPED CONTROL radio dose commands authenticate sender USB / serial port device software disabled when unused cloud patient data, results encryption, authentication the FDA reads 'can connect to the internet' broadly: any of these interfaces counts
SpecificationCodeDataThreatFocal detail
Plate 44 — Every interface through which data or commands can enter is part of a device's attack surface. A threat model lists each entry point, what an attacker could read, change or stop through it, and the control that prevents it, and is updated with the architecture.
A low likelihood of attack is a forecast about an adversary

Estimates that an attack is unlikely rest on assumptions about who would try, with what skill and for what gain, and those assumptions age. Research, published exploits and automated tools lower the skill an attack needs. Judging a weakness by how easily it could be used, and by the harm that would follow, ages better than judging it by how likely an attack seems today.

When security becomes safety

Some attacks threaten patients directly, as a changed insulin dose would. Others threaten data, and through it trust and care. In January 2025 the FDA warned that the software of a patient monitor sold as the Contec CMS8000, and under another name, contained a backdoor and could send patient data outside the care setting. It advised hospitals that relied on the monitors' remote monitoring to unplug and stop using them, and others to use them for local monitoring only. The fix released in July 2025 removed the monitors' networking entirely: the cure for an unsafe connection was to have none.

The pump and the monitor show the two hard truths of device security. A defense that was not designed in may be impossible to add later, and the safest response to some weaknesses is to give up a function. Chapter 23 describes the defenses that have to be in the design from the start.

Build the threat model from the architecture

Start from the architecture diagram of Chapter 5 and mark every interface, including service and debug ports and cloud connections. For each, list what could be read, changed or stopped through it, by whom and with what effort, and the control that prevents it. Judge each path by exploitability and clinical harm rather than by an estimate of how likely an attack is, and update the model with every design change.

+ What this chapter established
  • MiniMed pumps recalled in 2019 accepted radio commands without proper authentication and could not be patched, so patients were offered replacements.
  • Security risk is judged by exploitability and harm, not probability, because an attacker chooses the inputs; known vulnerabilities count as foreseeable.
  • Every interface is part of the attack surface; US law and the FDA's broad reading make almost any connectable device a cyber device.
  • A threat model built from the architecture names each path, its harm and its control, and some weaknesses can be removed only with the function.
23 — Secure by design

Built to be defended.

+ The questionWhich defenses have to be designed in before release, because they cannot be added later?

A signature checked at every start

A firmware image is one file: the complete program a device will run, compiled, tested and released under a single version number. Before it leaves the maker, a release system computes a fingerprint of that file, a short number called a hash that changes completely if a single bit of the file changes, and signs the fingerprint with a private key kept on a guarded machine. The signature travels with the file. Inside the device, in memory that cannot be rewritten, sit the matching public key and a small piece of start-up code. Every time the device powers on, that code computes the fingerprint again, checks it against the signature, and runs the firmware only if the two agree.

This chain is called secure boot, and the protected code and key at its start are the root of trust. Each stage checks the next before handing over control, so a program that someone has altered, however slightly, is refused before it runs. NIST's guidelines on firmware resiliency, published in May 2018 for computing platforms generally, describe three duties that such a design serves: protecting the firmware against unauthorized changes, detecting changes that occur anyway, and "recovering from attacks rapidly and securely," usually by returning to a copy known to be good.

45 — Secure boot: a chain of checks
AT THE MAKER firmware image hash (fingerprint) signed with the private key on a guarded release machine image + signature shipped or updated IN THE DEVICE, AT EVERY START ROOT OF TRUST start-up code + public key in memory that cannot be rewritten fixed in hardware before the first unit ships MATCH? checks bootloader signature bootloader MATCH? checks firmware signature firmware runs the application mismatch: refuse to run fall back to known-good copy protect · detect · recover (NIST SP 800-193, 2018) CHECKSUM VS SIGNATURE CRC attacker edits file → recomputes CRC → passes signature attacker edits file → cannot sign without the private key → fails FDA 2026: do not rely on CRCs as security controls
SpecificationCodeThreatFaultReadoutFocal detail
Plate 45 — Secure boot runs only firmware whose signature matches a key held in hardware that cannot be rewritten. A checksum catches accidental corruption, but an attacker can recompute it; only a signature made with a private key proves the file is genuine.

The design depends on something present in the hardware from the first unit shipped: protected memory for the key and the start-up code. A device built without it cannot acquire it through an update, because the update itself would have to be trusted by code that cannot tell a genuine update from a forged one.

Why a checksum is not a lock

Embedded software has long carried checksums, most often a cyclic redundancy check, or CRC, which detects a file corrupted by a faulty memory chip or a noisy cable. The FDA's premarket cybersecurity guidance, revised on 3 February 2026, says plainly that such checks "do not provide integrity or authentication protections in a security environment," and tells makers not to rely on them as security controls. The reason is simple. A CRC is computed by a published method from the file alone, so anyone who changes the file can compute a new one that matches. A signature can be produced only with the private key, so a matching signature proves both that the file is unchanged and that it came from the holder of the key.

The same distinction runs through every defense in this chapter. A safety mechanism protects against chance: a corrupted bit, a failed sensor, an unlikely input. A security control must hold against someone who knows how the mechanism works and is trying to defeat it, the adversary of Chapter 22. Many of the controls look alike from outside. Their design differs in what they assume about the person on the other side.

A lifecycle with security in it

The standard for building security into health software is IEC 81001-5-1, published in December 2021. It was written to sit on top of a device maker's existing process rather than beside it: its scope is based on IEC 62304, its clauses follow IEC 62304's order, and it takes its security activities from IEC 62443-4-1, the equivalent standard for industrial control products. Where IEC 62304 has a software risk management process, it has a security risk management process, and it adds security tasks such as threat modeling, security requirements, security testing and the handling of vulnerabilities after release. A team already following IEC 62304 extends its procedures rather than writing a second set. The standard has been flagged for revision since September 2025.

The FDA calls such a process a secure product development framework, and its 2026 guidance names IEC 81001-5-1 and ANSI/ISA 62443-4-1, the US adoption of IEC 62443-4-1, as examples that can satisfy it. For the risk assessment itself it points to two documents from AAMI: a technical report, TIR57, published in 2016, and the standard ANSI/AAMI SW96, published in 2023 and aligned with the ISO 14971 process of Chapter 4.

46 — IEC 62304 and IEC 81001-5-1 side by side
security activities taken from IEC 62443-4-1 (industrial products) IEC 62304 (SAFETY) IEC 81001-5-1:2021 (SECURITY) TASKS ADDED 5 software development 5 development + security tasks 6 software maintenance 6 maintenance + security updates 7 software risk management 7 security risk management 8 configuration management 8 configuration management 9 problem resolution 9 problem resolution, including vulnerabilities security risk management where IEC 62304 has software risk management threat modeling security requirements security testing handling vulnerabilities after release scope based on IEC 62304; clauses in the same order · revision flagged September 2025 FDA 2026: an example of a secure product development framework
CodeThreatFocal detail
Plate 46 — IEC 81001-5-1 was built to extend an IEC 62304 process rather than sit beside it: its clauses follow the same order, it adds a security risk management process and security tasks, and the FDA accepts it as a secure product development framework.

Eight kinds of control

The guidance's first appendix sorts the defenses a device may need into eight categories: authentication, authorization, cryptography, the integrity of code, data and execution, confidentiality, event detection and logging, resiliency and recovery, and firmware and software updates. The list reads like a summary of what the MiniMed pumps lacked. Their radio protocol did not authenticate the sender, so a pump could not tell its patient's remote from an attacker's transmitter. A 2018 advisory on the pumps' remote controllers had also described a replay attack, in which recorded traffic could be sent again to trigger a bolus when the remote options were enabled.

The guidance answers each of those faults in a sentence. It asks makers to "use cryptographically strong authentication, where the authentication functionality resides on the device," and not to use passwords "that are hardcoded, default, easily guessed, or easily compromised." It asks for anti-replay measures, such as a number used only once in each exchange, for commands that could do harm. It asks that updates be authenticated before they are installed, and that hardware-based security be used where possible. The architecture the maker submits is expected to show the device in four views: the whole system with its connections, the paths by which one compromise could harm many patients, how updates travel from the maker to every unit, and the security of each use case. The testing it lists includes penetration testing, fuzz testing, tests with malformed and unexpected inputs, analysis of the attack surface and of how weaknesses can be chained, and software composition analysis of the compiled binaries.

The defenses that matter most are fixed before the first unit ships

Protected memory for keys, a processor able to check signatures quickly, room for a second firmware image and a radio protocol with authentication are decisions taken early in design. A device shipped without them can sometimes be protected by its surroundings, never fully repaired, and the MiniMed recall shows that the remedy may be a new device.

Designing for the patch

A device that can be updated has a security control the MiniMed pumps lacked: a way to change its own code after release. The Abbott pacemakers of 2017 had one, though only through a programmer's wand in a clinic, and Abbott's letter to physicians shows that the update path itself carries risk. It recommended against replacing the devices to avoid the vulnerability, and asked physicians to weigh the update's risks for each patient, with particular care for patients who depend on their pacemaker for every heartbeat.

Solved example: the risk of a patch

Abbott's letter gave three failure rates from its earlier firmware updates: the device reloading its previous firmware after an incomplete update, 0.161 percent; loss of the programmed settings, 0.023 percent; complete loss of device function, 0.003 percent. Per 100,000 updates, that is about 161 devices left on the old version but working, 23 needing to be reprogrammed, and 3 that stop working. For a patient whose heart does not beat without pacing, the last outcome is the one that matters, which is why the letter advised considering the update for such patients where temporary pacing was available. The most common failure was also the safest, because the device kept its previous firmware and returned to it when the new one did not load completely. That fallback is a design decision, and it turned most failed updates into a device still working on its old version.

The example shows the shape of the decision a maker faces after release. A patch closes a vulnerability that, for these pacemakers, had produced no reports of a compromised device; it opens a small, measurable risk of its own. Good design makes the trade easier: an update path that authenticates the update, writes it to a spare slot, checks it, and keeps the old version until the new one has started correctly. Chapter 24 follows what happens when the vulnerability arrives after release, in code the maker did not write.

Write the security requirements into the first architecture

Decide at the start which hardware the device needs for security: protected key storage, signature checking at start-up, room for a fallback image. Add authentication, anti-replay and authorization to every interface in the threat model, plan how updates will reach every unit and how a failed one recovers, and list the security tests, from fuzzing to penetration testing, in the verification plan.

+ What this chapter established
  • Secure boot checks a signed firmware image against a key held in protected hardware at every start, so altered code never runs.
  • A CRC detects accidents but not attackers; only a signature made with a private key proves a file is genuine.
  • IEC 81001-5-1 adds security risk management and security tasks to an IEC 62304 process, and the FDA accepts it as a development framework.
  • Controls such as authentication and an update path with a fallback must be designed in; a device built without them may have to be replaced.
24 — Vulnerabilities after release

The threat changes while the code does not.

+ The questionOnce a device ships, how are new weaknesses in its code and its components found, judged and fixed?

Eleven flaws in someone else's code

Eleven vulnerabilities, six of them serious enough to let an attacker on the network run code of their own choosing, were disclosed in July 2019 in IPnet, a network stack: the part of an operating system that sends and receives data over a network. The security company that found them, Armis, named them URGENT/11 and reported that the affected code had been part of the VxWorks real-time operating system since version 6.5, and of several other operating systems that had used the same stack. On 1 October 2019 the FDA warned patients, clinicians and makers that medical devices and hospital networks using those operating systems could be affected. Its announcement named six operating systems from six suppliers, said that device makers' notices so far included an imaging system, an infusion pump and an anesthesia machine, and said the agency knew of no adverse events.

No device maker had written IPnet. It reached their products inside operating systems they had chosen and, under Chapter 6's rules, treated as SOUP. The code in those devices had not changed since release. What changed was public knowledge about it, and with it the risk. This is the general pattern of security after release: the threat changes while the code does not, and most of the code at issue came from someone else.

47 — One network stack, many devices
OPERATING SYSTEMS NAMED BY THE FDA DEVICES IN MAKERS' NOTICES IPnet network stack URGENT/11 11 vulnerabilities; 6 allow remote code execution disclosed July 2019 written by none of the device makers discoverer: Armis (security company) VxWorks, since 6.5 Wind River OSE ENEA INTEGRITY Green Hills ThreadX Microsoft ITRON TRON ZebOS IP Infusion imaging system OWN CODE infusion pump OWN CODE anesthesia machine OWN CODE own code unchanged since release hospital networks FDA communication 1 Oct 2019: no adverse events known
CodeThreatFocal detail
Plate 47 — URGENT/11 reached medical devices through a network stack inside operating systems that device makers had chosen. None of them wrote the vulnerable code, and their own code had not changed; what changed was public knowledge about it.

Matching the list to the news

The first task when a vulnerability is published is finding out which products contain the affected component, and in which version. This is the work the software bill of materials of Chapter 6 exists for. A maker that keeps a machine-readable SBOM for every released version can search all of them for the component within hours; a maker that must ask each development team, or inspect old build files, may take weeks. The FDA's 2026 premarket guidance expects makers to monitor sources such as NIST's National Vulnerability Database, to assess the known vulnerabilities in each component, and to give particular weight to those in the US cybersecurity agency's catalog of vulnerabilities already being exploited in real attacks.

Presence is only the first question. A device may contain the vulnerable component without using the vulnerable function, or without exposing it to any interface an attacker can reach. The answer for each vulnerability and each product is recorded as a VEX statement, a Vulnerability Exploitability eXchange, which the US Commerce Department's telecommunications agency, NTIA, described in 2021 as an assertion about the status of a vulnerability in a specific product, with four values: not affected, affected, fixed, or under investigation.

Solved example: from 24 patches to 4

In June 2020, after vulnerabilities in another widely used network stack, made by Treck, were disclosed, the infusion pump maker B. Braun had received 24 patches from Treck for its Outlook 400ES pump and had determined that 20 did not apply, according to a trade report. Four of 24 is 17 percent. Each of the 20 needed a reasoned statement of why the pump was not affected; each of the 4 needed an assessment of exploitability and harm, and a fix. The scale of that work grows with the bill of materials. If a product carries the mean of 1,180 open-source components that a code-audit vendor reported (Chapter 6), and 1 percent of them receive a new advisory in a year, about 12 advisories arrive for that product each year; at B. Braun's ratio, about 2 would need a fix, and every one would need a written conclusion. Without an SBOM, the first step of each, knowing whether the component is present at all, is the slowest.

Judging what a weakness means for patients

A vulnerability that applies is judged on two scales. The first is technical severity, usually expressed with the Common Vulnerability Scoring System that Chapter 22 read for the MiniMed pumps. Its fourth version, published on 1 November 2023, added a supplemental Safety metric that records whether exploitation could cause injury, using the consequence categories of the functional safety standard IEC 61508. The specification is explicit that supplemental metrics do not change the score. A rubric for applying the scoring system to medical devices, written by MITRE under an FDA contract, was qualified by the FDA as a medical device development tool in October 2020.

The second scale is clinical. The FDA's postmarket cybersecurity guidance of December 2016 divides risks into controlled, where the residual risk of patient harm from a vulnerability is acceptable, and uncontrolled, where it is not. For an uncontrolled risk, the agency said it did not intend to enforce its usual reporting rules for corrections if four conditions held. There were no known serious adverse events or deaths; the maker told its customers within 30 days of learning of the risk, with interim measures and a plan; it fixed the risk, validated the fix and distributed it within 60 days; and it took part in an information-sharing organization for the sector. Since March 2023 the law has set the pace for connected devices: known unacceptable vulnerabilities fixed on a "reasonably justified regular cycle," and critical ones that could cause uncontrolled risks "as soon as possible out of cycle."

48 — From a published vulnerability to a fix
vulnerability published national database; catalog of exploited vulnerabilities search the SBOM of every released version hours with an SBOM, weeks without component present? vulnerable function reachable through an interface? yes VEX STATEMENT under investigation not affected affected fixed no no yes a written conclusion for every match technical severity (CVSS) 0 10 clinical harm controlled uncontrolled residual risk: acceptable or not unacceptable, not critical critical fix on the regular cycle out of cycle, as soon as possible (critical) FDA 2016 postmarket guidance, uncontrolled risk: tell customers within 30 days, fix within 60 A TRADE REPORT, JUNE 2020 Treck stack, one infusion pump did not apply: 20 applied: 4 24 patches → 4 to fix (17%)
SpecificationCodeThreatFaultReadoutFocal detail
Plate 48 — An SBOM turns a new vulnerability into a search; a VEX statement records whether each product is actually exposed. Of 24 patches one infusion pump maker received for a third-party stack, 4 applied, and each of the 24 needed a written conclusion.

Coordinated disclosure

Most device vulnerabilities are found by people outside the company: academic researchers, security firms, hospital engineers. Coordinated vulnerability disclosure is the practice by which they report a weakness privately to the maker, the maker acknowledges and investigates it, and the two agree when the details are published, ideally with a fix or mitigation available. A government coordinator often stands between them; the MiniMed advisories of 2018 and 2019 were published by the US government's coordinator, now part of CISA, after researchers reported the flaws and Medtronic analyzed which models shared them. Two international standards describe the maker's side, ISO/IEC 29147 for receiving and publishing reports and ISO/IEC 30111 for handling them internally, and both were flagged for revision in 2025. Since 2023 US law has required makers of connected devices to submit a plan for monitoring and addressing vulnerabilities after release, "including coordinated vulnerability disclosure and related procedures."

The hospital's share

A fix protects no patient until it is installed, and for many devices installation happens in the hospital. Each update has to be scheduled around clinical use, tested against the hospital's own configuration where the integration of Chapter 18 could be affected, and applied unit by unit, sometimes by the maker's service engineers and sometimes by the hospital's clinical engineering staff. Armis, which found URGENT/11, estimated in December 2020 that 97 percent of the operational-technology devices affected had not been patched; its figure covers far more than medical devices, but it shows how slowly fixes reach equipment that is not a personal computer.

A fix protects patients only once it is installed

Publishing a patch moves the risk from the maker's queue to the hospital's. Until a unit is updated, its protection depends on what surrounds it: network segmentation, restricted access, disabled services and monitoring. Hospitals need the maker's SBOM and VEX statements to know which units are exposed, and makers need the hospitals' installed versions to know which patients remain at risk.

The EU's medical device guidance on cybersecurity, issued in 2019 and revised in 2020, describes security as a joint responsibility of makers, integrators, operators and users, while noting that only makers carry the legal duties. It also says a device should not depend on security controls in its environment. In practice, the defense of a device in the field is shared: the maker designs it to be patched, watches for new weaknesses and issues fixes; the hospital keeps an inventory of what it runs, installs fixes and protects what it cannot yet fix.

Make vulnerability handling a routine with deadlines

Keep the SBOM of every released version searchable, monitor vulnerability databases and the catalog of exploited vulnerabilities daily, and record a VEX statement for each match. Judge each applicable vulnerability by exploitability and clinical harm, set a target date in the regular cycle or out of it, publish a disclosure policy with a contact address, and tell hospitals which versions are affected and what to do until they can patch.

+ What this chapter established
  • URGENT/11 put eleven vulnerabilities into devices through an operating system's network stack that no device maker had written.
  • A searchable SBOM finds affected products within hours, and a VEX statement records whether each one is actually exposed.
  • Severity scores measure the weakness; FDA guidance and US law set fix timelines by whether the risk to patients is controlled.
  • Researchers report through coordinated disclosure, and fixes protect patients only when hospitals install them; both sides share the defense.

+ Part VII · The field in 2026

Who builds software devices.

The FDA's public list of AI-enabled medical devices was last updated on 4 September 2026. It is a table rather than a register: each row is one marketing authorization, and the agency says the list "is not a comprehensive resource." The page itself gives no total. Independent analysts counted 1,451 entries decided by the end of 2025 and 1,524 including decisions up to 30 March 2026, and in November 2025 the director of the FDA's device center told an advisory committee that the agency had authorized more than 1,200 AI-enabled devices. The single chapter of this part uses the list and the makers' own records to describe the field as it stood in October 2026: which devices reach patients, who makes them, and where generative AI has arrived. It is the most dated chapter in the guide and will be the first to age.

25 — Who builds them

Software and AI devices in 2026.

+ The questionWhich software and AI devices reach patients in 2026, who makes them, and where does generative AI stand?

Mostly images

About three in four entries on the list are radiology devices. One analyst's count, reported in July 2026, put radiology at 1,104 of the 1,451 entries decided by the end of 2025, 76 percent, and at about the same share of the devices authorized in 2025 alone. The rest come from other specialties, ophthalmology among them, where IDx-DR sits.

The reasons are not hard to see from the earlier chapters. Images are digital from the moment they are made, they carry the standard headers of Chapter 18, and hospitals have kept them for decades in archives alongside the reports that describe them, which gives makers training data with something close to a label attached. A reading task also has a natural reference, the expert reader of Chapter 12, against which a model's answers can be scored.

Regulation has reinforced the pattern. When the FDA granted a De Novo to the stroke-triage program ContaCT on 13 February 2018, it created a new classification for radiological software that flags suspected findings and notifies a specialist. Later devices of the same type could then be cleared through the faster 510(k) route by showing they were substantially equivalent to one already on the market, and later devices were. A first authorization in a category opens the road for the products that follow.

49 — Who is on the FDA's AI list
ENTRIES DECIDED BY THE END OF 2025: 1,451 radiology 1,104 (76%) other specialties 347 (24%) about three in four are radiology RADIOLOGY AUTHORIZATIONS BY COMPANY, THROUGH END OF 2025 0 20 40 60 80 100 120 GE HealthCare 120 Siemens Healthineers 89 Philips 50 Canon 45 United Imaging 38 Aidoc 31 DeepHealth 28 SCANNER AND IMAGING-SYSTEM MAKERS AI SOFTWARE SPECIALISTS an analyst's count reported July 2026; the FDA page gives no total and says the list is not comprehensive GE HealthCare's own figure, July 2025: 100 listed authorizations (maker's figure; different date and rules)
ReadoutFocal detail
Plate 49 — About three in four entries on the FDA's list of AI-enabled devices are radiology devices, and scanner makers lead the counts. The totals are analysts' counts: the FDA publishes the list but not a total.

Who makes them

The same analyst's counts of radiology authorizations by company, through the end of 2025, put four makers of scanners and imaging systems at the top: GE HealthCare with 120, Siemens Healthineers with 89, Philips with 50 and Canon with 45, followed by United Imaging with 38. Aidoc, a company that makes only AI software, came next with 31, and DeepHealth with 28. Much of the equipment makers' AI runs inside or beside their own scanners and reaches hospitals with the hardware. The software specialists sell triage and detection programs that run beside the archive, often through a platform that hosts several makers' tools behind one integration.

A third group rarely appears on the list but carries much of this guide: makers of pumps, pacemakers, ventilators and monitors, whose software drives therapy rather than reading images. Most of their software is conventional code, judged by the lifecycle of Part II and defended as in Part VI. The MiniMed and Abbott cases of Chapters 22 and 23 came from this group, not from AI.

Solved example: what a count of AI devices counts

GE HealthCare said in July 2025 that it had 100 listed authorizations; the analyst's count put it at 120 radiology authorizations by the end of 2025. Both can be right, because they count at different dates and by different rules. Units matter more. Qure.ai, a company based in Mumbai, said in February 2026 that it held 26 FDA-cleared indications across 9 products, about 2.9 indications per product, while its own regulatory page listed 7 cleared products. Aidoc announced in January 2026 a clearance covering 14 acute findings on CT from a single model, 11 of them newly cleared. If that was one submission, it is one entry counted by submission and 14 counted by indication. Two analysts counting radiology's share of 2025 authorizations got 75 percent and 71.5 percent, 211 of 295, a gap of 3.5 points from counting method and cutoff. A number of "AI devices" means little until it says whether it counts submissions, products or indications, and at what date.

A maker's count of its own authorizations is a claim to check

Company figures for clearances, indications and "firsts" are written for investors and customers. Each can be checked against the FDA's databases, where every authorization has a number, a date and a decision summary, and against the list's own entries. Where the two disagree, the regulator's record is the evidence.

Makers in India

Indian makers reach patients first through other regulators. Qure.ai's head-CT triage program qER was cleared by the FDA on 17 June 2020, under the classification created for ContaCT, and the company's regulatory page lists CE certificates under the EU's device regulation for several products. Its page also reports an approval from India's regulator without giving a number. India's own requirements for software devices were set out in one place only in July 2026, in the regulator's final guidance on medical device software, described in Chapter 26. An Indian manufacturer announced in August 2026 what it called India's first manufacturing license for diagnostic software, a claim about in vitro diagnostic software that is not specific to AI. This guide found no reliable source naming the first AI software device licensed by India's regulator.

Generative AI at the door

Generative AI, models that produce text, images or speech rather than a score, reached the FDA's advisory process before it reached the list. The agency's Digital Health Advisory Committee met for the first time in November 2024 to discuss the whole product lifecycle for generative AI devices, and again in November 2025 on generative AI for mental health. At the second meeting the device center's director said that none of the digital mental-health devices authorized so far involved generative AI. The list page, as updated in September 2026, says the FDA will explore ways to identify and tag devices that incorporate foundation models, "from large language models (LLMs) to multimodal architectures"; no entry was tagged.

A device whose maker says it uses large language models was cleared six months before its maker announced it. UpDoc, an insulin-management program for adults with type 2 diabetes, was cleared through a 510(k) on 23 December 2025 as a drug-dose calculator, class II. Its public summary describes conversational data collection by voice or chat and insulin instructions computed from parameters set by the patient's clinician, and reports non-clinical testing only. It does not use the words "large language model" or "generative AI." The company announced on 25 June 2026 that this was the first FDA clearance of software using patient-facing large language models; the FDA has not confirmed the claim, and when a reporter asked whether generative AI made treatment decisions, the chief executive would not say. The clearance Aidoc announced in January 2026, built on what the company calls a foundation model, is a different case: a large image model, not a language model, doing a triage task of a kind already regulated.

50 — Where the language model sits
INTERFACE as the 510(k) summary describes UpDoc patient conversation service: voice or chat maker says it uses large language models structured data: readings, doses taken CLEARED LOGIC dosing logic computed from clinician-set parameters clinician sets parameters insulin instruction where the model sits decides the evidence evidence can concentrate on whether the service captures what the patient meant DECISION what the record would need to show patient language model writes the recommendation insulin instruction would need test sets, references and subgroup results for an output that varies with phrasing (Parts III and IV) cleared 23 Dec 2025 (510(k), drug-dose calculator, class II); announced 25 Jun 2026; the FDA summary does not use the words "large language model"
SpecificationCodeDataFaultReadoutFocal detail
Plate 50 — A language model that collects a patient's answers for fixed, cleared rules is an interface; a model that writes the recommendation is the decision and needs the evidence of Parts III and IV. UpDoc's public summary describes the first.

The evidence question for generative AI is the central question of this guide in a new place. A language model that collects a patient's answers and passes them to fixed, cleared rules is an interface, and the evidence can concentrate on whether it captures what the patient meant. A model that writes the recommendation itself is the decision, and it would need the test sets, references and subgroup results of Parts III and IV for an output that can vary with every phrasing. The public record should say which of the two a device is. For UpDoc, the summary's description, with doses computed from clinician-set parameters, points to the first.

Read the record before the press release

For any software device, find its FDA submission number, decision date and summary, and note the product code and the device it was compared with. Check which indications the authorization covers, whether it includes a change control plan, and what testing it reports. Treat a maker's counts and "firsts" as claims until the record supports them, and look for where any language model sits relative to the clinical decision.

+ What this chapter established
  • About three in four entries on the FDA's list of AI devices are radiology devices, helped by digital archives, natural references and early classifications.
  • Scanner makers lead the counts, AI software specialists follow, and the makers of therapy devices carry most conventional device software.
  • Counts of AI devices depend on whether they count submissions, products or indications; the regulator's record settles disagreements.
  • UpDoc, cleared in December 2025, is claimed by its maker as the first device with a patient-facing language model; the FDA record does not say so.

+ Part VIII · Quality and regulation, as of October 2026

What a software device must show.

On 27 July 2026 a regulation the EU calls the Digital Omnibus on AI entered into force, three days after its publication. Among other changes to the AI Act of 2024, it moved the date from which AI in products such as medical devices must meet the act's high-risk requirements, from 2 August 2027 to 2 August 2028. It nearly did more: the European Parliament had proposed moving medical devices to a lighter category of the act's product list, and the final compromise left them where they were. Rules for software devices change like this, by amendment, guidance and revised date, while the mechanism beneath them holds. The four chapters of this part therefore put the lasting mechanism first and the clauses last: what makes software a device and sets its class, what evidence it must file, which updates need a regulator, and what the rules on security and AI add, including for the hospital that deploys it. Every status in them is as of October 2026.

26 — When software is a device

Intended use draws the line.

+ The questionWhich software is regulated as a medical device in the US, the EU and India, and in what class?

Two guidances on one day

On 6 January 2026 the FDA reissued two guidances that together mark where its oversight of software stops. The first, on general wellness products, said that certain non-invasive products estimating blood pressure, oxygen saturation, blood glucose or heart rate variability can be wellness products, outside device regulation, when their outputs are "intended solely for wellness uses." The conditions are strict. Such a product may suggest that a user see a health professional when a reading falls outside a wellness range, but it may not name a disease, call a value abnormal, use clinical thresholds or give alerts for managing a disease, and it may not show values that look clinical unless they have been validated.

The second, on clinical decision support software, revised the criteria for software that advises clinicians without being a device. It added enforcement discretion for software that gives a single recommendation where only one is clinically appropriate, and it moved its treatment of time-critical decisions to the criterion about whether a clinician can independently review the basis of a recommendation, the change Chapter 16 described. The FDA corrected the decision-support guidance on 29 January 2026. Neither document changed the law. Both moved the practical line by explaining how the agency reads it, which is how the boundary of software regulation usually moves.

The US line

In US law a device is an instrument or related article intended for the diagnosis of disease or other conditions, or for the cure, mitigation, treatment or prevention of disease, and software can meet that definition. The 21st Century Cures Act of December 2016 carved five kinds of software function out of that definition. Software for the administrative support of a health care facility is excluded, as is software for maintaining a healthy lifestyle unrelated to any disease, software serving as an electronic patient record, and software that transfers, stores, converts or displays laboratory or device data without interpreting it. The fifth exclusion is clinical decision support, which must meet four criteria at once.

The software must not acquire, process or analyze a medical image, a signal from an in vitro diagnostic test, or a pattern or signal from a signal acquisition system. It must display, analyze or print medical information. It must support or provide recommendations to a health care professional. And it must let that professional independently review the basis for its recommendations, so that the professional does not rely primarily on them. IDx-DR fails the first criterion, because it analyzes images, and the fourth, because no professional reviews the images before the result is given. It is a device on two counts.

51 — Four criteria for decision support outside device rules
EXCLUDED FROM THE DEVICE DEFINITION (21ST CENTURY CURES ACT, 2016) administrative support healthy lifestyle, unrelated to disease electronic patient records transfer, store, convert or display lab or device data, without interpreting FIFTH EXCLUSION: DECISION SUPPORT, IF ALL FOUR CRITERIA HOLD 1 does not acquire, process or analyze a medical image, an IVD signal or a signal pattern 2 displays, analyzes or prints medical information 3 supports or provides recommendations to a health care professional 4 the professional can independently review the basis, not relying primarily on it yes yes yes yes non-device clinical decision support no no no device software function if it meets the device definition; then classed I, II or III by risk IDx-DR no would also fail here: no one reviews the images first analyzes images: a device FDA guidance revised 6 Jan 2026, corrected 29 Jan 2026; time-critical decisions handled under criterion 4
SpecificationCodeReadoutFocal detail
Plate 51 — US law excludes some software from the device definition, but decision support qualifies only if it meets all four criteria. IDx-DR analyzes images and gives its answer without a professional reviewing them, so it is a device on two counts.

A function that is a device is then classed by risk, I, II or III, through a classification regulation that names the device type, as IDx-DR's 2018 De Novo created one for retinal diagnostic software in class II. The class sets the route to market: most class II software reaches it through a 510(k) or a De Novo, class III through premarket approval.

The EU line

The EU's medical device regulation classifies software by its Rule 11. Software that provides information used to take decisions for diagnosis or therapy is class IIa; it is class IIb if those decisions could cause a serious deterioration of a person's health or a surgical intervention, and class III if they could cause death or an irreversible deterioration. Software that monitors physiological processes is IIa, or IIb where it monitors vital parameters whose variations could result in immediate danger. All other software is class I. The rule's practical effect is that almost any software informing a clinical decision is at least class IIa, and so needs a notified body, an independent organization designated to assess conformity, before it can carry the CE mark.

The Medical Device Coordination Group's guidance on qualifying and classifying software, first endorsed in October 2019, was revised in June 2025. According to a regulatory consultancy's review, the revision made no substantive change, but it used the term "medical device artificial intelligence" for the first time, with a reference to the AI Act, and added examples, including software that informs a prognosis or prediction.

The class follows the worst credible consequence of a wrong output

Every scheme in this chapter asks what could happen to a patient if the software's information or action were wrong. A maker that classifies by how the software usually performs, or by the clinician who is expected to catch its errors, will usually reach a lower class than the regulator.

India's grid

India's regulator published its final guidance on medical device software on 21 July 2026, after a draft in October 2025. It drops the terms software as a medical device and software in a medical device, and distinguishes standalone software from software that is part of, or drives, a hardware device. Software that drives or influences hardware takes the hardware's class. Standalone software is classed on a grid that follows the 2014 international framework of Chapter 1, with India's classes A to D.

Solved example: one function classified three ways

Take autonomous screening for diabetic retinopathy from retinal photographs, used by primary-care staff, as in the opening scene. In the US it is a device, because it analyzes images and no clinician reviews the basis of its answer; its type is class II. Under the EU's Rule 11 it provides information used for a diagnostic decision, so it is at least IIa; whether it is IIb turns on whether a missed referral could cause a serious deterioration of health, an argument the maker must make and the notified body must accept. On India's grid, a situation that is serious and software that diagnoses give class C; a critical situation would give D, and software that only informed a clinician in the same serious situation would be class A. India's guidance adds that software used by non-clinical users in a serious situation without specialist support may be treated as critical, which would move this function to D if lay operators ran it at home. One function, three schemes: class II, IIa or IIb, and C or D, each decided by the claim and the user.

52 — One function, three classifications
INDIA'S GRID FOR STANDALONE SOFTWARE TREAT OR DIAGNOSE DRIVE INFORM CRITICAL D C B SERIOUS C B A NON-SERIOUS B A A autonomous retinopathy screening lay users without specialist support may be treated as critical decided by the claim and the user India, final guidance 21 July 2026; grid follows the 2014 international framework US class II device type (De Novo, 2018) EU RULE 11 IIa or IIb information for diagnostic decisions → IIa; IIb if a wrong decision could cause serious deterioration of health
SpecificationReadoutFocal detail
Plate 52 — The same screening function is class II in the US, IIa or IIb under the EU's Rule 11, and C, or D with lay users, on India's grid. Each scheme asks what a wrong output could do to the patient and how much the output decides.

The rest of India's scheme follows the same logic. Where several classes could apply, the highest wins, and the regulator keeps a list of software it has classified. Manufacturing licenses for classes A and B are issued by state authorities and for classes C and D by the central authority. The guidance states that it reflects current practice under India's Medical Devices Rules of 2017 rather than creating a new control. Internationally, the medical device regulators' forum published a further document in January 2025 on characterizing software risk, which suggests assuming, where possible, that the software will fail, so that the risk assessment rests on the consequences rather than on an estimate of probability.

Write the intended use before the classification

State in one paragraph what the software is for, who uses it, in which clinical situation, and what its output decides. Classify it under each market's scheme from that paragraph, record the reasoning, and check it against the regulators' examples. Keep the intended use under change control, because a new claim or a new user can change the class without a line of code changing.

+ What this chapter established
  • The FDA moved its boundary in January 2026 by guidance: wellness estimates and single-recommendation decision support, under unchanged law.
  • US law excludes five kinds of software; decision support qualifies only if a clinician can review its basis and it does not analyze images or signals.
  • EU Rule 11 makes almost any decision-informing software class IIa or higher; India classes standalone software A to D on a grid.
  • The same screening function is class II in the US, IIa or IIb in the EU and C or D in India, decided by the claim and the user.
27 — The evidence a software device files

Documents, validation and clinical evidence.

+ The questionWhat does each regulator ask to see before a software or AI device is sold, and what do the validation rules ask of the sites and makers that use software under GMP?

Two levels of documentation

A documentation level, in the FDA's guidance on premarket submissions for device software, final since 14 June 2023, is the amount of software evidence a submission must contain, set by what a failure could do. The level is Enhanced where a failure or latent flaw in the software "could present a hazardous situation with a probable risk of death or serious injury," judged before any risk control is counted; otherwise it is Basic. The guidance replaced one from 2005 and applies to every device with software, AI or not.

Most of the file is the same at both levels: the evaluation of the level itself, a description of the software, the risk management file, the software requirements, an architecture diagram, the version history and the list of unresolved anomalies. The Enhanced level adds the detailed design specification, the configuration management and maintenance plans, and the protocols and reports of unit and integration testing, where Basic asks for a summary of testing and the system-level protocol and report. At the Basic level a maker can describe its development practices by declaring conformity to IEC 62304, which the FDA has recognized in full since 2019. Chapters 4 and 7 traced where each of these records comes from; the guidance is the list of which ones leave the company.

53 — Basic and Enhanced documentation
could a failure or latent flaw present a hazardous situation with a probable risk of death or serious injury, judged before risk controls? the level is set by the worst case, not the design's defenses no yes BASIC ENHANCED SAME AT BOTH LEVELS documentation level evaluation · software description risk management file · software requirements architecture diagram · version history · unresolved anomalies SOFTWARE DESIGN SPECIFICATION — in the submission DEVELOPMENT AND MAINTENANCE PRACTICES summary, or IEC 62304 declaration + configuration management and maintenance plans TESTING summary + system-level protocol and report + unit and integration protocols and reports FDA guidance on device software content, final 14 June 2023 (replaced the 2005 guidance)
SpecificationCodeReadoutFocal detail
Plate 53 — The FDA sets software documentation at Basic or Enhanced by what a failure could do before any risk control counts. Most of the file is the same at both levels; Enhanced adds the design specification, development plans and unit and integration test records.

Behind the submission sits the quality system. Since 2 February 2026 the FDA's quality rule for devices, renamed the Quality Management System Regulation, has incorporated the international standard ISO 13485:2016 by reference, and the agency's inspectors have stopped using their old inspection technique. The surgical energy guide in this library treats IEC 62304 and the documentation levels as items on a generator maker's compliance list; this chapter has followed them from the software's side.

What the FDA asks of AI

For AI, the FDA's main document is still a draft. Its guidance on lifecycle management and marketing submissions for AI-enabled device software functions was published in January 2025 and remained a draft in October 2026. It proposes sections a submission should contain beyond the general software file: a description of the device and its user interface, labeling, risk assessment, data management, the model's description and development, validation, plans for monitoring performance in the field, cybersecurity, and a public summary.

Its expectations read like Parts III and IV of this guide. Test data should be independent of training data, for example drawn from different clinical sites, and sequestered from the developers. Validation data should represent the intended population, and relying on a single collection site is generally not appropriate; the draft gives at least three geographically diverse US sites as an example of what may be. Performance should be reported across subgroups such as sex, age, race, ethnicity, disease severity, site and acquisition equipment, at each operating point. It suggests, without requiring, a model card: a short structured summary of the model, its data and its performance, placed in the labeling. The final guidance on change control plans of Chapter 21 completes the set.

The EU and India

In the EU, a software device's evidence goes into the technical documentation that a notified body reviews. For clinical evidence, the coordination group's guidance of March 2020 sets out three components. The first is a valid clinical association: the software's output is associated with the clinical condition it targets. The second is technical performance: the software generates its intended output accurately and reliably from its input data. The third is clinical performance: the output is clinically relevant in accordance with the intended purpose. All three are kept current by post-market evaluation. The guidance predates the AI wave and does not mention AI. IEC 62304 is not a harmonized standard under the EU's device regulation, and neither is the security standard of Chapter 23: in the list as consolidated on 17 June 2026, neither appears. A maker can still use them as the state of the art, but they give no presumption of conformity.

India's final guidance of July 2026 uses the same three components. For AI it adds dataset requirements of its own: the composition of training, validation and test data, including demographic distribution, geographic origin and clinical diversity; whether data are real-world, from public databases or synthetic; and a justification where models were trained or validated outside India. Performance shown on non-Indian data should be supplemented with validation in representative Indian populations. The guidance lists IEC 62304, IEC 81001-5-1, ISO 14971 and ISO 13485 among the standards that may apply.

A draft guidance signals direction while the final one governs

The FDA's AI-enabled device guidance, its draft on change control plans for all devices and the EU's draft manufacturing annex on AI all describe where regulators are heading, and makers sensibly design toward them. A submission is still judged against the final documents in force on the day it is filed, so a file needs to show which version of each it followed.

Standards in motion

Standards change more slowly than guidance, and their status is easy to misreport. IEC 62304's first edition dates from 2006, with an amendment in 2015. A second edition, widened to health software in general, has been in preparation for years, and its status can be read directly from the international project record.

Solved example: reading a standard's stage code

ISO and IEC track every project with a stage code: 30 is the committee stage, where drafts circulate among experts; 40 the enquiry, a draft circulated to national bodies for vote; 50 the approval of a final draft; 60.60 publication. The project record for the second edition of IEC 62304, as mirrored by a national standards body, showed stage 30.20, a committee draft ballot initiated, from 7 August 2026, with the next milestone due on 27 November 2026. A consultancy that follows the committee reported almost 1,500 comments on the first committee draft and expected the final draft no earlier than 2028. A vendor's web page, by contrast, said the final draft stage had begun on 22 May 2026 and that publication was scheduled for 12 August 2026. A standard cannot be at stage 30 and stage 50 at once. The record wins: in October 2026 the 2015 consolidated first edition is the one in force, and anything said about the second edition, such as the proposed replacement of the three safety classes by two levels of process rigor, reported by the German association VDE, describes a draft.

54 — Where IEC 62304's second edition stands
20 PREPARATORY 30 COMMITTEE drafts among experts 40 ENQUIRY national bodies vote 50 APPROVAL (FINAL DRAFT) final draft vote 60.60 PUBLISHED project record: 30.20 committee draft ballot, from 7 Aug 2026; next milestone 27 Nov 2026 a standard cannot be at stage 30 and 50 at once a vendor's web page: final draft from 22 May 2026, publication 12 Aug 2026 (contradicted) consultancy: final draft not before 2028; IEC forecast publication Oct 2028 IN FORCE IN OCTOBER 2026 IEC 62304 edition 1.1 (2006 + amendment 2015), recognized by the FDA in full almost 1,500 comments on the first committee draft (2025) draft proposals (e.g. two process-rigor levels replacing classes A, B, C) may change
SpecificationFaultReadoutFocal detail
Plate 54 — Stage codes make a standard's status checkable. In October 2026 the second edition of IEC 62304 was at committee draft stage 30.20, whatever one vendor's page said, and edition 1.1 remained the one in force.

Validation rules for those who use software

The rules for sites and makers that use software, rather than sell it, run in parallel. The FDA's rule on electronic records and signatures, 21 CFR Part 11, has applied since 1997. Its guidance on computer software assurance, final in September 2025 and revised on 3 February 2026 under a new title that refers to quality management system software, sets out the risk-based testing of Chapter 9 for software used in production and in the quality system; it does not cover the device software itself. In the EU, a revised Annex 11 on computerized systems and a new Annex 22 on AI, the draft described in Chapter 19, went out for consultation from July to October 2025. On 8 October 2026 the EU's list of good manufacturing practice guidelines still showed the Annex 11 of 2011 and no Annex 22.

Build the submission from the records, release by release

Determine the documentation level before design starts, and keep each required record, from requirements to unresolved anomalies, current with every release so that a submission is an export rather than a project. For AI, keep the data management records, the sequestered test sets and the subgroup results the FDA, EU and Indian documents ask for, and record which version of each guidance and standard the file follows.

+ What this chapter established
  • The FDA asks for Basic or Enhanced software documentation, set by whether a failure could probably cause death or serious injury before controls.
  • The FDA's AI guidance is still a draft; it asks for independent, sequestered and representative test data and subgroup results, and suggests a model card.
  • The EU and India ask for clinical association, technical performance and clinical performance; India adds Indian-population data.
  • IEC 62304's second edition is a committee draft, and Part 11, CSA and the 2011 Annex 11 govern software used under GMP, with a revised Annex 11 and a new Annex 22 still in draft.
28 — When an update needs a regulator

Changes, jurisdiction by jurisdiction.

+ The questionWhich software changes can a maker release on its own, and which need a regulator first?

One release, three answers

The same two-week release can be routine in one jurisdiction and a new submission in another. Suppose a team finishes version 2.4 of an AI screening program. It retrains the model on more data from the same kinds of clinics, moves the operating threshold within the range already validated, applies a security patch to the operating system and upgrades an image-processing library to a new major version. The code was built, verified and released through the pipeline of Chapter 9, and every record is in place. Whether it may now reach patients depends on where they are.

The three regulators ask the same underlying question, the one this guide has asked of every change: could it move a claim the evidence supports, or create a risk the evidence did not cover? They encode the answer differently. The FDA asks a series of questions that lead to a likely answer; the EU sorts changes into significant and not; India divides them into major and minor. Each also offers a way to agree some changes in advance.

The FDA's questions

The FDA's guidance on deciding when to submit a 510(k) for a software change, final since 25 October 2017, applies to cleared devices and to those authorized through De Novo. It works as a flowchart. A change made solely to strengthen cybersecurity, with no other effect on the device, is likely not to need a new submission, and neither is one made solely to return the device to the specification of its most recently cleared version. A change that introduces a new risk, or modifies an existing one, that could result in significant harm and is not effectively mitigated likely does need one. So does a new or modified risk control for a hazard that could cause significant harm. The last question is whether the change could significantly affect clinical functionality or performance specifications; if so, a submission is likely needed. Changes to what the device is indicated for go through the FDA's general guidance on device changes.

55 — The FDA's software change questions
software change to a cleared or De Novo device solely to strengthen cybersecurity, no other impact? yes likely no new 510(k): document solely to return to the most recently cleared specification? no yes likely no new 510(k): document new or modified risk, or risk control, for a hazard that could cause significant harm? no yes likely new 510(k) could it significantly affect clinical functionality or performance specifications? no yes likely new 510(k) no document the assessment and release change inside an authorized change control plan → no new submission (Chapter 21) retraining is meant to change performance a retrained model usually lands here FDA guidance, final 25 Oct 2017 changes to indications: the FDA's general change guidance
SpecificationCodeReadoutFocal detail
Plate 55 — The FDA's flowchart lets security-only fixes and returns to specification ship with documentation, but sends changes that affect significant risks or performance toward a new 510(k). A retrained model usually meets the performance question unless a change control plan covers it.

The flowchart leaves the decision to the maker, which must document its reasoning and is accountable for it in an inspection. For a model, the performance question usually decides: retraining is meant to change performance, so a retrained model is likely to need a new submission unless an authorized change control plan already covers it. That is why the plans of Chapter 21 matter. UpDoc's, for example, lists five categories of change the maker may make without a new submission, none of which alters the cleared dosing logic.

The EU: significant change and substantial modification

Under the EU's medical device regulation, a maker tells the notified body that certified a device of planned changes, and changes that could affect conformity need that body's approval. The most detailed published list of what counts for software is in the coordination group's guidance on significant changes. That guidance was written for devices still certified under the older directives, but its software chart sets out the reasoning plainly. A change is significant if it brings a new or modified architecture or database structure, or a change of an algorithm, or if a closed-loop algorithm replaces a user's input. So is a new operating system or any new component, or a major change to one. A new diagnostic or therapeutic feature, a new channel of interoperability, a new user interface or a new presentation of medical data are significant too. Safety-neutral bug fixes, security updates, changes to appearance and gains in efficiency are not significant, provided they do not affect diagnosis or treatment. One significant answer makes the whole change significant.

The AI Act adds its own test. A substantial modification is a change to an AI system after it has been placed on the market that was not foreseen in its initial conformity assessment and that affects its compliance or modifies its intended purpose. A high-risk system that undergoes one needs a new conformity assessment. Changes that the provider predetermined at the initial assessment, and documented, are not substantial modifications: the EU's form of a change control plan. For medical device AI, those provisions apply from 2 August 2028, and the EU's joint guidance on the device rules and the AI Act, from June 2025, notes that retraining may trigger reassessment under both and that a plan agreed in advance can reduce how often.

Retraining is a design change in every jurisdiction

A model retrained on new data has new parameters and, by design, new performance. Every regulator in this chapter treats that as a change to evaluate against the device's claims, and none treats it as maintenance. A plan agreed before authorization is the main way to make retraining routine.

India: major and minor

India's final guidance on medical device software, published in July 2026, divides software changes into two lists. Major changes need the licensing authority's approval before they are made. They include design changes that affect specifications, indications or performance; changes to software design or system requirements, new clinical claims or new types of input data; and changes of intended use. Label changes other than font, color or layout are major, and so are version changes that affect intended use, safety, effectiveness or risk controls. Minor changes are notified: bug fixes and security patches that do not affect intended use, safety or performance, version changes with no such effect, and re-tuning of performance within validated ranges. Software is identified by its version, revision level and build or release date. Under India's rules a minor change must be notified within 30 days.

For in vitro diagnostic devices, an addendum of 13 March 2026 to the regulator's questions and answers states that each new version of approved software needs a post-approval change application, the route the immunoassay guide in this library describes for analyzers. For AI, the July guidance adds an optional algorithm change protocol with four parts: a data management plan, a plan for evaluating and monitoring performance, a retraining plan where retraining is intended, and a plan for software updates and rollback.

Solved example: version 2.4 in three jurisdictions

Take the four changes of the opening section. In the US, the operating-system security patch falls under the flowchart's first question and is likely not to need a submission; the threshold move within the validated range and the library upgrade need a documented assessment of whether they could significantly affect performance; the retraining is likely to need a 510(k) unless an authorized plan covers it. In the EU, the security patch is not significant, but the retraining is a change of an algorithm and the library's new major version is a major change of a component. By the reasoning of the coordination group's chart, at least two of the four are significant, and the threshold move may be a third if it changes the algorithm's output; under the regulation, the notified body must approve any of them that could affect conformity. In India, the patch and the threshold move are minor and notified within 30 days, the library upgrade is minor if it affects nothing the guidance lists, and the retraining is major if it changes performance specifications. One release yields one likely submission in the US, at least two significant changes in the EU and one major change in India, before the AI Act applies to any of them.

56 — Version 2.4 in three jurisdictions
US (FDA) EU (MDCG CHART) INDIA (CDSCO) RETRAIN ON MORE DATA FROM THE SAME KINDS OF CLINICS likely 510(k), unless a change control plan covers it significant: change of an algorithm major if performance specifications change every market treats retraining as a change MOVE THRESHOLD WITHIN THE VALIDATED RANGE documented assessment may be significant if the output changes minor: notify within 30 days OPERATING-SYSTEM SECURITY PATCH likely no submission not significant minor: notify IMAGE LIBRARY: NEW MAJOR VERSION documented assessment significant: major change of a component minor if no listed effect RELEASE TOTAL 1 likely submission at least 2 significant changes 1 major change regulator first document or notify depends on the output AI Act substantial-modification rules for device AI apply from 2 Aug 2028
SpecificationCodeFocal detail
Plate 56 — One release of four changes yields one likely submission in the US, at least two significant changes in the EU and one major change in India. Retraining needs a regulator everywhere unless a plan agreed in advance covers it; a security patch needs one nowhere.

Patches and borrowed code

Security patches and updates to third-party components are where the schemes come closest. All three treat a patch that only closes a vulnerability as something a maker can release without prior approval, after documenting its assessment and, in India, notifying the regulator, which keeps the regular and out-of-cycle fixes of Chapter 24 possible. They part company on larger component changes: a new major version of an operating system or library is significant under the EU's chart, while the FDA and India ask what it could do to safety and performance. A maker that keeps its SBOM, its regression evidence and its assessment of each component change in one record can answer all three from the same file.

Assess every release against every market before it ships

For each change in a release, record which claims and risks it could affect, the verification run, and the regulatory decision for each market: documented assessment or submission in the US, significant or not in the EU, major or minor in India. Group changes that need a regulator into planned releases, keep security patches separate so they can ship fast, and write a change control plan for the changes a model will need.

+ What this chapter established
  • The FDA's 2017 software-change flowchart lets security-only and return-to-specification changes ship, but changes that affect risk or performance usually need a 510(k).
  • The EU treats algorithm, architecture and major component changes as significant; the AI Act adds substantial modification, with predetermined changes excluded.
  • India lists major software changes needing approval and minor ones to be notified, and offers an optional algorithm change protocol.
  • One release can need one submission in the US, notified-body review in the EU and approval in India; plans agreed in advance make change routine.
29 — Security, AI and the deploying site

Section 524B, the AI Act and the deployer's duties.

+ The questionWhat do the cybersecurity and AI-specific rules add, what do they ask of the hospital that deploys the device, and how does one device meet two EU regulations at once?

Section 524B

Ninety days after an appropriations law was enacted in the US on 29 December 2022, a new section of the Food, Drug, and Cosmetic Act took effect, on 29 March 2023. Section 524B applies to any cyber device: one that includes software validated, installed or authorized by its maker, can connect to the internet, and has characteristics that could be vulnerable to cybersecurity threats. A maker submitting such a device must do four things. It must provide a plan to monitor, identify and address vulnerabilities after release, including coordinated disclosure. It must design, develop and maintain processes giving reasonable assurance that the device and related systems are secure, and make updates and patches available on a regular cycle and, for critical vulnerabilities, out of cycle. It must provide a software bill of materials covering commercial, open-source and off-the-shelf components. And it must meet any further requirements the FDA sets by regulation.

The FDA gave makers a transition: until 1 October 2023 it worked with sponsors whose submissions lacked the new information, and after that date it could refuse to accept them. Its premarket cybersecurity guidance, the source of Chapters 22 to 24, was revised on 3 February 2026 to align with the new quality management system rule, without new technical expectations, according to two law and certification firms' reviews. In March 2026 a standards body announced that the FDA had recognized its consensus report on cybersecurity considerations unique to machine-learning devices.

One device, two EU regulations

In the EU, cybersecurity is part of the medical device regulation itself. Its general safety and performance requirements ask that software be developed according to the state of the art, taking account of the development lifecycle, risk management including information security, verification and validation. They also require makers to set out the minimum requirements for hardware, IT networks and IT security measures, including protection against unauthorized access, needed to run the software as intended. The coordination group's cybersecurity guidance of December 2019, revised in July 2020, adds that a device should not rely on security controls in its operating environment. The EU's Cyber Resilience Act, which applies to most products with digital elements from 11 December 2027, excludes devices covered by the medical device and diagnostic regulations, so for a hospital it matters for other software it buys, not for its medical devices.

AI adds a second regulation on top. Under the AI Act, an AI system is high-risk if it is, or is a safety component of, a product covered by the EU laws listed in its first annex and that product must undergo assessment by a third party. Medical devices are on that list, and almost all decision-informing software needs a notified body under Rule 11, so most medical device AI is high-risk. The high-risk requirements cover risk management, data governance, technical documentation, automatic logging, information for those who deploy the system, human oversight, and accuracy, robustness and cybersecurity. The Digital Omnibus of July 2026 kept medical devices in that annex, moved the date these requirements apply to medical device AI to 2 August 2028, and, according to the consolidated text, clarified that AI used solely for purposes other than safety does not qualify as a safety component.

The two regulations meet in one assessment. A device that is also high-risk AI follows the conformity assessment of the medical device regulation, and the notified body checks the AI Act's requirements within it. The joint guidance of the device coordination group and the EU's AI board, published in June 2025, describes one technical documentation, an integrated quality system and integrated risk management, rather than two parallel files.

57 — One device, two EU regulations
MEDICAL DEVICE REGULATION general safety and performance requirements: lifecycle, risk management incl. information security, V&V minimum IT and security requirements for the site Rule 11 class (IIa and above → notified body) AI ACT, HIGH-RISK REQUIREMENTS risk management data governance technical documentation automatic logging information for deployers human oversight accuracy, robustness, cybersecurity one technical documentation, integrated quality system and risk management notified body: device assessment, AI Act checked within it CE MARKING two regulations, one assessment DEPLOYER (the hospital) FROM 2 AUG 2028 use per instructions human oversight by competent, trained staff relevant, representative input data monitor; tell the provider; suspend if at risk keep logs at least 6 months inform workers in use Cyber Resilience Act: excludes devices under the MDR and IVDR; applies to other software (from 11 Dec 2027)
SpecificationReadoutFocal detail
Plate 57 — A medical device that is high-risk AI meets two EU regulations in one technical file and one notified-body assessment. The hospital that uses it becomes a deployer with duties of its own from 2 August 2028; the Cyber Resilience Act does not reach the device.
A device that leans on the hospital's network for its security hands the risk to the hospital

Segmentation, firewalls and access control at the site are valuable layers, and the EU's minimum IT requirements belong in the instructions for use. A device designed to be safe only behind them, though, transfers its security to an organization that did not design it and may change its network tomorrow.

The deployer's duties

The AI Act gives duties to the organization that uses a high-risk system under its own authority, which it calls the deployer: for a hospital AI device, the hospital. It must use the system in accordance with its instructions and assign human oversight to people with the necessary competence, training and authority. Input data under its control must be relevant and sufficiently representative for the intended purpose. It must monitor the system's operation, inform the provider without undue delay of risks, and immediately of serious incidents, and suspend use where it sees a risk. It must keep the logs the system generates automatically, where they are under its control, for at least six months, and inform workers before the system is used in their workplace. The Omnibus also rewrote the act's article on AI literacy: providers and deployers must take measures to support it, without having to guarantee any level of literacy in an individual.

These are, almost line for line, the site practices of Chapters 18 to 20 made into law: integration within the specification, a local check that the inputs match the conditions of the evidence, monitoring with named people, and a channel back to the maker. For hospitals using AI that is a medical device, they apply from 2 August 2028.

India

India's guidance of July 2026 brings security into the same file as safety. According to trade reports of its text, it asks makers to keep a bill of materials that includes an SBOM, covering third-party, open-source and commercial components, and to link it to monitoring and mitigating known vulnerabilities. Cybersecurity is treated as a risk area, vulnerabilities are to be monitored after deployment, and devices are to keep operating safely during incidents, including partial loss of connectivity, denial of service and loss of integrity. It lists IEC 81001-5-1 among the standards that may apply. Its AI dataset rules were set out in Chapter 27.

Patient data are governed by the Digital Personal Data Protection Act of 2023 and its rules, notified in November 2025 and phased in. The rules on the data protection board applied at once; the registration of consent managers applies from 13 November 2026; and most obligations on those who process data, including security safeguards and notifying breaches, apply from 13 May 2027. The penalty for failing to keep reasonable security safeguards can reach 250 crore rupees.

Solved example: counting down from October 2026

From 8 October 2026, the next dates that reach a software device or the site that runs it fall over two years. India's consent-manager rules apply in about one month, on 13 November 2026, and its main data protection duties, including breach notification, in about seven months, on 13 May 2027. The Cyber Resilience Act's full application follows on 11 December 2027, about 14 months away, for the hospital's non-device software. The AI Act's high-risk requirements for medical device AI, including the hospital's duties as deployer, apply on 2 August 2028, about 22 months away. A device authorized today and updated every quarter will see about seven releases before the last of these dates, and each release will be assessed against the rules in force on its own release day.

58 — Dates ahead
2023 2024 2025 2026 2027 2028 29 Mar 2023: section 524B in force (US) 27 Jul 2026: Digital Omnibus in force (EU) 11 Sep 2026: CRA vulnerability reporting (non-device products) 8 OCT 2026: TODAY 13 Nov 2026: India consent managers (about 1 month) 13 May 2027: India data protection duties incl. breach notice (about 7 months) 11 Dec 2027: Cyber Resilience Act fully applies, not to devices (about 14 months) 2 Aug 2028: AI Act high-risk rules for device AI and deployers (about 22 months) device AI and its deployers quarterly releases: about 7 before Aug 2028
CodeReadoutFocal detail
Plate 58 — From October 2026 the next rules reaching a software device or its site fall over 22 months, ending with the AI Act's high-risk duties for device AI on 2 August 2028. A device released quarterly will ship about seven versions before then, each judged against the rules of its day.
Map every device to the rules that reach it, with dates

For each device, list the rules that apply in each market and the date each takes effect: section 524B for cyber devices, the device regulation and the AI Act together in the EU, India's software guidance and data protection rules. Assign each duty to the maker or the site, write the site's duties into the responsibility agreement of Chapter 18, and review the list whenever a regulation changes its dates.

The four-hour operators, again

The operators of the opening scene had four hours of training and a locked program. What their trial showed in 2017 was narrow and solid: that program, on that camera, run by people like them, gave the right screening decision for adults like those enrolled often enough to beat bars agreed in advance. It could not show what the program would do in a clinic whose systems sent images differently, after a camera or model changed, or against someone trying to make it fail.

The same device filed in October 2026 would carry evidence for each of those gaps. Because its images and answers cross the internet, it is a cyber device, so its file would hold a threat model, an SBOM, security test results and a disclosure plan. Its software documentation would be at the level its potential for harm sets. Following the FDA's draft for AI, it would show test data drawn from several sites and kept from its developers, with results by subgroup and by camera, and very likely a change control plan for the retraining it will need. In India it would be expected to add validation in Indian populations; in the EU, from August 2028, logging and human-oversight measures assessed by its notified body. The hospital that installed it would test its interfaces, check its inputs against the specification, run it silently on local patients before relying on it, watch its positive and declined rates, and tell the maker when something moved.

None of that changes the invariant. Every way the software can fail was built in on the day it shipped or arrived with an update, and for the model the training data is part of what was built. The evidence that it will be right for the next patient is a measurement of what was built, taken on patients chosen to resemble the next one; what keeps it true after an update is the discipline that measures again. Four hours were enough to take the pictures. Everything else in this guide is what it takes to keep trusting the answer.

+ What this chapter established
  • Since March 2023, US cyber devices must file a vulnerability plan, show secure processes with patching, and provide an SBOM.
  • In the EU, a medical device that is high-risk AI meets the device regulation and the AI Act in one assessment, with AI duties from 2 August 2028.
  • Hospitals deploying such AI become deployers, with duties for oversight, input data, monitoring and logs that mirror good site practice.
  • India's 2026 guidance brings SBOMs and cyber resilience into the device file, and its data protection duties phase in through May 2027.
§ Lessons

Lessons.

Fourteen working rules follow from the chapters, for anyone who builds, buys, installs, regulates or writes about software and AI in medical devices. Each names the chapters that explain why it holds.

  1. Write the intended use before anything else. The claim turns code into a device, sets its class in each market and fixes what the evidence must show; a new claim or a new user can change all three without a line of code changing (Chapters 1 and 26).
  2. Look for the fault and the input that reaches it, not a wear-out rate. Software does not age: every failure was built in on the day it shipped or arrived with an update, and it shows itself only when some input reaches it (Chapter 2).
  3. Build confidence from the process, because testing alone cannot supply it. Paths outnumber any test campaign and very low failure rates cannot be demonstrated by running the software, so the evidence is requirements derived from risk and traced to tests, an architecture that keeps the dangerous part small, and rigor scaled to harm (Chapters 3 to 5 and 7).
  4. Account for every component someone else wrote. Record each one's exact version, published anomalies, support status and replacement plan, generate the SBOM from the build, and treat pretrained models and public datasets as components too (Chapter 6).
  5. Let every sprint and every pipeline run leave its records. IEC 62304 prescribes activities and records, not phases, so agile increments and automated builds can produce its evidence, provided a definition of done includes verification and a person still decides each release (Chapters 8 and 9).
  6. Keep the data that judges a model away from the people and data that built it. Separate test sets by patient and, where possible, by site, sequester them from developers, and watch for the leakage paths that let a model learn the hospital instead of the disease (Chapters 10 and 11).
  7. Ask who decided the right answer and where yes begins. The reference standard limits what measured accuracy can mean, and the operating point, the endpoints and the treatment of declined cases must be fixed before the test and reported with their intervals (Chapters 12, 13 and 17).
  8. Test where the device will be used, and report by subgroup. Accuracy is a measurement on particular patients, equipment and settings; it moves when any of them moves, and an average can hide a group the model fails (Chapters 14 and 15).
  9. Measure the reader and the model together when a clinician uses the output. An assistant that is usually right can still make a reader worse, so the evidence for an assistive device is the pair's performance, with automation bias in view (Chapter 16).
  10. Treat installation as part of the evidence. Test every interface with local data, confirm that local inputs fall inside the specification, run the device silently on local patients before anyone relies on it, and write down who notifies whom of changes on either side (Chapters 18 and 19).
  11. Watch inputs and output rates from the first day, and close the loop on a schedule. Shifts show in days in what the device receives and returns; accuracy shows only in reference reads of sampled negatives and matched outcomes, with named people acting on agreed limits (Chapter 20).
  12. Agree in advance what a model may change, and fix the tests each version must pass. A change control plan is the main route to routine retraining: authorized plans in the US, an optional algorithm change protocol in India and predetermined changes under the EU's AI Act from August 2028; outside such a plan, retraining is a change a regulator must see (Chapters 21 and 28).
  13. Design the device to be attacked and to be patched. Authenticate every interface, sign the code and keep a fallback image, build the threat model from the architecture, put the SBOM to work when vulnerabilities are published, and publish a way to report them (Chapters 22 to 24).
  14. Read the record, and date every status. Makers' counts and "firsts" are claims until a regulator's database supports them, and every rule, guidance and standard in this field is quoted as of a day; the mechanism lasts longer than the clause (Chapters 25 to 29).
§ Glossary

Glossary.

Terms are defined as they are used in this guide.

510(k)
A US premarket submission by which the FDA clears a device shown to be substantially equivalent to one already on the market. Once a De Novo has created a device type, later devices of that type can follow this route; IDx-DR's new versions were cleared this way in June 2021 and June 2022.
AAMI TIR45
A technical information report from AAMI, the Association for the Advancement of Medical Instrumentation, that maps agile development, in short cycles each ending with tested software, onto IEC 62304 in four layers: project, release, increment (a sprint of one to four weeks) and story (a few days of user-visible function). The FDA recognizes its 2012 and 2023 editions in full.
AI Act
The EU's 2024 law on artificial intelligence, under which most medical device AI is high-risk, because medical devices are on its product list and almost all decision-informing software needs a notified body. Its high-risk requirements, from risk management and data governance to logging and human oversight, apply to medical device AI from 2 August 2028, as moved by the Digital Omnibus of July 2026.
AI model
In this guide, an algorithm whose decision rules were fitted to example data rather than written line by line. Its accuracy is a measurement on a test set and holds only for patients like those in it, against the reference chosen, for the exact model that was locked.
Algorithm
The procedure the software follows. When its decision rules are fitted to example data rather than written line by line, the guide calls it an AI model.
Algorithm change protocol
In the FDA's 2019 discussion paper, the methods a maker would use to make, test and control changes to a machine-learning model within an agreed region of potential changes. India's July 2026 guidance offers an optional AI protocol of the same name in four parts: data management, performance evaluation and monitoring, retraining, and updates with rollback.
All-pairs testing
Testing every pair of parameter settings rather than every combination, which grows slowly with the number of settings. In 98 percent of 109 FDA recall reports that NIST researchers could trace, it would have revealed the failure; the HAMILTON-C6 ventilator fault needed four conditions at once, so a pairwise campaign would not have been sure to find it.
Attack surface
Every way data or commands can enter a device: radio and Bluetooth links, Wi-Fi and Ethernet, USB and serial ports, a cloud connection, a service interface used only by technicians. The MiniMed pumps' unauthenticated radio link was the part of theirs that mattered.
AUC
Area under the curve: the area under a receiver operating characteristic curve, from 0.5 for a coin toss to 1 for perfect separation, which summarizes how well a score separates diseased from healthy cases regardless of where the threshold is set. It hides the operating point and can stay stable while calibration drifts; the Google retinal network of 2016 scored 0.991.
Automation bias
The tendency to accept a machine's output in place of one's own judgment, which studies suggest is strongest when the person is uncertain. In a 2023 study, wrong suggestions labeled as AI cut radiologists' accuracy on mammograms from about 80 percent to between 20 and 46 percent, with the least experienced readers affected most.
Autonomous, assistive and triage devices
The three roles of software that analyzes clinical data. An autonomous device gives the answer itself, as IDx-DR does; an assistive device gives a reader extra information and the reader decides, so the reader's performance with and without it is what counts; a triage device works in parallel with the usual workflow, alerting a specialist while radiologists still read every image.
Baseline
The results of a site's local performance check, kept as the reference for later monitoring: the distribution of inputs, the rate of declined cases, and the local sensitivity, specificity and predictive values. In agile work the word also names the controlled set of requirements approved at each release.
Calibration
Whether a model's scores mean what they say, such as predicted risks matching the true rate; discrimination is whether it ranks sick patients above healthy ones. Seven kidney-injury models at US Veterans Affairs hospitals kept their discrimination for nine years while their calibration drifted, so a fixed threshold would have alerted ever more often.
CDSCO
Central Drugs Standard Control Organisation: India's device regulator. Its final guidance on medical device software, of 21 July 2026, classes standalone software A to D, divides software changes into major and minor, and asks makers to justify AI models trained or validated outside India; licenses for classes A and B come from state authorities, for C and D from the central authority.
CE mark
The marking under which a device is sold in the EU. Under Rule 11 almost any software that informs a clinical decision is at least class IIa, so it needs a notified body's assessment before it can carry the mark.
Change control
The discipline that judges every change to a device against the claims it could move and the risks it could create before it reaches patients; for an AI model it adds a plan agreed with the regulator in advance that bounds what may change and fixes the tests each new version must pass. It is the guide's answer to what keeps the evidence true after an update.
Clinical decision support
Software that advises clinicians, which US law excludes from the device definition only if it meets four criteria at once: it does not analyze a medical image or a diagnostic signal or pattern; it displays, analyzes or prints medical information; it supports recommendations to a health care professional; and that professional can independently review their basis. IDx-DR fails the first and the fourth.
Computerized system validation (CSV)
Validation by its user of a computerized system that affects product quality, patient safety or data integrity, which GxP rules require of laboratories, blood services, clinical-trial organizations and drug manufacturers. Most follow GAMP 5, scaling the effort to the system's impact, complexity and novelty and building on the supplier's evidence.
Confidence interval
The range that expresses the uncertainty of an estimate made from a sample, at a stated confidence, usually 95 percent; it narrows as the number of cases behind the estimate grows. IDx-DR's observed sensitivity of 87.4 percent has an exact interval of 81.9 to 91.7 percent, and a sensitivity of 90 percent spans 78.6 to 95.7 from 50 patients but 87.1 to 92.3 from 500.
Configuration management
Identifying and controlling every item that goes into a build, from source code and SOUP versions to build scripts and the compiler with its settings, so that the same inputs produce the same software and the version tested is the version shipped. IEC 62304 requires it at every class.
Coordinated vulnerability disclosure
The practice by which someone who finds a weakness reports it privately to the maker, the maker investigates, and the two agree when the details are published, ideally with a fix; a government coordinator such as the US cybersecurity agency CISA often stands between them. US law has required it in connected-device makers' vulnerability plans since 2023.
Coverage
A measure of how much of the code a test campaign executed. The FDA's 2002 guidance describes a ladder from statement coverage, every line run once, which it calls insufficient on its own, through decision, condition and multiple-condition coverage, to full path coverage, which it calls generally not achievable; coverage shows what ran, not whether it was right.
CRC
Cyclic redundancy check: a checksum computed from a file by a published method, which detects corruption by a faulty memory chip or a noisy cable. Anyone who changes the file can compute a matching one, so the FDA tells makers not to rely on it as a security control; only a signature made with the maker's private key proves a file is genuine.
CSA
Computer software assurance: the FDA's risk-based approach to validating software used in production and in the quality system, finalized in September 2025 and reissued in February 2026. Functions whose failure could foreseeably compromise safety get documented, scripted testing and the rest lighter, unscripted methods; it does not cover the device software itself.
CVSS
Common Vulnerability Scoring System: the usual measure of a vulnerability's technical severity, built from factors such as the attack vector, its complexity and the impact on confidentiality, integrity and availability. The 2019 MiniMed weakness scored 7.1, High, on version 3; the score says nothing about how many patients use a device or what harm follows.
Cyber device
Under section 524B of the US Food, Drug, and Cosmetic Act, a device that includes software validated, installed or authorized by its maker, can connect to the internet, and has characteristics that could be vulnerable to cybersecurity threats. The FDA reads connection broadly, down to a USB port used briefly for service; IDx-DR, whose images and answers cross the internet, would be one.
Cyber Resilience Act
The EU law, CRA for short, on the cybersecurity of products with digital elements, which applies in full from 11 December 2027. It excludes devices covered by the medical device and diagnostic regulations, so for a hospital it governs other software it buys, not its medical devices.
Dataset shift
Any difference between the distribution of cases a model was fitted and tested on, meaning patients, images and labels in particular proportions, and the one it meets in use. It can change the inputs (a new camera or population), the prevalence, the right answer itself, called concept shift (a new grading guideline), or the acquisition chain (a scanner update, compression, a new format), and time produces it too.
De Novo
The US route by which the FDA grants authorization to a new kind of device and creates a classification regulation for its type, so that later devices of the type can be cleared through a 510(k). A De Novo is granted, not cleared: IDx-DR's, on 11 April 2018, created the class II type of retinal diagnostic software device.
Decision summary
The FDA's published account of the review behind a De Novo, with a shorter summary for most 510(k) clearances, covering the indications and limitations, the software, the special controls, the clinical study and the risks with their mitigations. IDx-DR's is the most complete public record of its evidence, though its date and one interval need checking against the counts.
Declined case
A case for which a device gives no usable answer, such as IDx-DR's result of insufficient image quality, which its labeling directs be referred. Headline figures usually leave such cases out: counting IDx-DR's 10 declined diseased participants as misses lowers its sensitivity from 87.4 to 83.2 percent, and counting them as referrals raises it to 88.0.
Definition of done
The conditions a story must meet before it counts as finished in agile work: requirement written and reviewed, design and code reviewed, unit and integration tests passed, trace links in place, risk analysis updated and any SOUP change recorded. Applied strictly, it lets the evidence for a release accumulate in small pieces as the work is done.
Deployer
Under the EU's AI Act, the organization that uses a high-risk AI system under its own authority; for a hospital AI device, the hospital. From 2 August 2028 it must follow the instructions, assign competent human oversight, keep its input data relevant and representative, monitor the system, report risks and serious incidents to the provider, and keep the logs for at least six months.
Device software function
Software that meets the legal definition of a medical device, whether it runs inside a pump or on a server by itself. What makes it one is its intended use, not anything in its code.
DICOM
Digital Imaging and Communications in Medicine: the standard format hospitals use for medical images and image-based results. Each image carries a header of tagged fields, such as the modality, the manufacturer, the model name and the equipment's software versions, which a device can read to refuse inputs outside its specification; IDx-DR accepted DICOM images from its 2021 version.
Documentation level
The FDA's measure, since June 2023, of how much software evidence a submission must contain: Enhanced where a failure could present a hazardous situation with a probable risk of death or serious injury, judged before any risk control is counted, and Basic otherwise. Enhanced adds the detailed design and the unit and integration test protocols and reports.
End of support
The date after which a software supplier no longer publishes security fixes. Extended support for Windows XP ended on 8 April 2014; devices that outlive their operating system's support stay exposed, as WannaCry showed in 2017 on unpatched or unsupported Windows systems; the FDA asks makers to state each component's support level and end-of-support date.
Enrichment
Recruiting more of one kind of participant or case than the population holds, such as the extra participants with poorly controlled diabetes in the IDx-DR trial, whose headline figures were corrected for it. In reader studies, enrichment with diseased cases can change how readers behave and biases predictive values.
Evidence
In this guide, records that someone outside the team could check: requirements, test results, study data, monitoring reports.
Exploitability
How easily an adversary could use a vulnerability. Because an attacker chooses the inputs, the FDA's cybersecurity guidance judges security risk by exploitability and by the harm that would follow, not by probability, and treats known vulnerabilities as reasonably foreseeable.
External and prospective tests
An external test uses data from sites, devices or periods that contributed nothing to a model's development and shows how it travels; a prospective test runs the device forward in its intended setting, on patients enrolled for the purpose, with its real workflow and operators. The IDx-DR trial was prospective, in primary care.
Fault
A flaw in a program's logic or data, present from the moment it was written, that causes a failure when an input reaches it, in every copy of that version. The HAMILTON-C6 ventilator fault waited inside three software versions for four ordinary events to coincide.
FHIR
Fast Healthcare Interoperability Resources: HL7's newer standard, built from web resources that can be addressed individually. Release 5 has been current since March 2023, and release 6 was in its second normative ballot in July 2026.
Foundation model
A class of large model, ranging in the FDA's words from large language models to multimodal architectures; in September 2026 the agency said it would explore tagging devices that incorporate one, and none was yet tagged. Aidoc's January 2026 clearance of 14 acute CT findings from one model rests on what the company calls a foundation model, a large image model.
GAMP 5
Good Automated Manufacturing Practice: the guide from the industry body ISPE that regulated companies follow to validate the computerized systems they buy and build, in its second edition of July 2022. It is risk-based, sorts software into categories from infrastructure to custom software, supports iterative and agile methods, and was joined in July 2025 by a separate GAMP guide to AI.
Generative AI
AI models that produce text, images or speech rather than a score. Where such a model sits relative to the clinical decision decides the evidence a device needs: one that only collects a patient's answers for fixed, cleared rules is an interface, while one that writes the recommendation is the decision.
Good machine-learning practice
Principles for developing AI devices, published by the FDA, Health Canada and the UK's MHRA in 2021 and finalized by the International Medical Device Regulators Forum in January 2025. Among them: test sets independent of training data, reference standards fit for purpose, data representative of the intended population, assessment of the human-AI team, and monitoring of deployed models.
GxP
Good practice: the collective name for the good-practice rules, such as good manufacturing practice, under which laboratories, blood services, clinical-trial organizations and drug manufacturers work. They require any computerized system that affects product quality, patient safety or data integrity to be validated by its user.
Hazard, hazardous situation and harm
The chain that ISO 14971 traces: a hazard is a potential source of harm, a hazardous situation is one in which a sequence of events exposes someone to it, and harm is what may follow. Each chain is judged by the severity of the harm and the probability of the sequence; for software, regulators suggest assuming the failure will happen.
HL7
Health Level Seven: a standards body, and its version 2 standard, which carries orders and reports between hospital systems as messages; the body says version 2 is used by 95 percent of US healthcare organizations.
IEC 62304
The international standard for the life cycle of medical device software, published on 9 May 2006, amended in June 2015 and recognized in full by the FDA since January 2019. It prescribes activities and their records, not their order, and sets three software safety classes; a second edition widened to health software was a committee draft in October 2026, not expected before 2028.
IEC 80001-1
The standard, in its 2021 edition, for organizations applying risk management before, during and after connecting a medical device or health software to their IT infrastructure, covering safety, effectiveness and security. It is marked for revision and is to be replaced by a standard in the 81001 series.
IEC 81001-5-1
The standard, published in December 2021, for building security into health software. It follows IEC 62304's scope and clause order and takes its security activities from IEC 62443-4-1, adding security risk management, threat modeling, security testing and vulnerability handling to an existing process; the FDA names it as one way to meet its secure product development framework.
IHE
Integrating the Healthcare Enterprise: an initiative that writes profiles telling vendors how to combine standards for a task. Its profiles for AI results and AI workflow, which define how results are encoded in DICOM and how analysis requests are sent and managed, were still at trial implementation in October 2026.
Intended use
What the maker says software is for, in its labeling and claims. It, not anything in the code, makes software a medical device, sets its class in each market and fixes what the evidence must show; IDx-DR's covers adults with diabetes not diagnosed with retinopathy, imaged on the Topcon NW400.
ISO 13485
The international standard for the quality management systems of device makers, in its 2016 edition. The FDA's device quality rule has incorporated it by reference since 2 February 2026, and it requires makers to validate software used in the quality system and in production, in proportion to risk.
ISO 14971
The international standard for the risk management of medical devices, whose 2019 edition was confirmed in 2025. It traces hazards through hazardous situations to harms and gives each unacceptable risk a control.
Large language model
A generative AI model that produces text, abbreviated LLM by the FDA. UpDoc's maker announced in June 2026 that its insulin program, cleared in December 2025, was the first FDA clearance of software using patient-facing large language models, a claim the FDA has not confirmed; its public summary points to an interface rather than the decision.
Leakage
Any route by which the data that judges a model influences how it is built, inflating the reported score: the same patient on both sides of a split, preprocessing statistics computed with the test images included, duplicates in both sets, or repeated looks at the test set until it becomes a second tuning set.
Legacy software
Software legally on the market but without enough evidence that it was built to IEC 62304. The standard's 2015 amendment lets it stay in use after a risk analysis using field experience, a gap analysis against the standard, a plan to close the gaps that matter and a documented rationale.
Local performance check
Running a device on a sample of a site's own cases, silently or retrospectively, and comparing its outputs with a local reference, after acceptance has confirmed the installation, interfaces and fallback. Showing a sensitivity above 75 percent when about 87 percent is expected takes about 69 diseased cases, or 690 consecutive cases where 10 percent have the disease.
Locked model
A model whose parameters are fixed after training, so that the same input always gives the same output and it can be verified like any other software; a continuously learning model, by contrast, keeps adjusting them from new data in use. IDx-DR's analysis was locked before the 2017 trial, and almost every authorized AI device is locked.
Machine-learning model
A fixed structure chosen by engineers whose parameters are set by training on data rather than by hand. The structure is code that can be reviewed; the parameters cannot be read as rules, so evidence about the model comes from its behavior on carefully chosen data.
MDR
Medical Device Regulation: the EU's law for medical devices, which classifies software by Rule 11, so that almost any software informing a clinical decision needs a notified body, and makes makers state the minimum hardware, network and IT security their software needs. Its counterpart for diagnostic tests is the in vitro diagnostic regulation, IVDR.
Medical Device Coordination Group (MDCG)
The EU body whose guidance interprets the device regulations: on qualifying and classifying software (2019, revised June 2025), clinical evidence for software (March 2020), significant changes, cybersecurity (2019, revised 2020) and, jointly with the EU's AI board in June 2025, meeting the device rules and the AI Act in one assessment.
Model card
A short structured summary of a model, its data and its performance, placed in the labeling; the FDA's draft guidance on AI-enabled devices suggests one without requiring it.
Modification protocol
The part of a predetermined change control plan that says how each planned change will be developed, validated and implemented: how new data are collected and kept separate, how the model is retrained and evaluated, the acceptance criteria it must meet and how users are told. Its acceptance criteria, fixed before any change is made, carry the plan's weight.
More than mild diabetic retinopathy
The condition IDx-DR screens for: in the 2017 trial, a worse eye at level 35 or higher on the severity scale of the Early Treatment Diabetic Retinopathy Study (ETDRS), or macular edema, as graded at the Wisconsin reading center. It was present in 24 percent of the analyzable participants.
Notified body
An independent organization designated in the EU to assess a device's conformity before it can carry the CE mark. Under Rule 11 almost all software that informs a clinical decision needs one, and for a device that is also high-risk AI, the notified body checks the AI Act's requirements within the same assessment.
OTS
Off-the-shelf software: in the FDA's term, a generally available software component used by a device maker that cannot claim complete control of its life cycle. A commercial operating system is OTS, and so is a compiler or test tool that never ships in the device but must be validated for its use.
Overfitting
Fitting a model so closely to its training examples that it follows their noise and predicts new data worse. In the guide's example, a ten-parameter polynomial passes through all ten training points yet misses new points by 1.67 on average, where a straight line misses them by 0.78.
PACS
Picture archiving and communication system: the hospital system that stores images after the scanner sends them, from which an imaging model receives studies and to which it often returns its results.
Parameter
One of the adjustable numbers inside a model, a weight that multiplies one value before it is passed on; ResNet-50, a widely used image network of some fifty layers, has 25,557,032. Parameters are not written but found by training.
Part 11
21 CFR Part 11, the FDA's rule on electronic records and signatures, in force since 1997, which requires records that are attributable, complete and protected, with audit trails.
Pipeline
An automated sequence that takes every submitted change, builds the software from source, runs the tests and analysis tools and packages the result, with no step done by hand. For a device maker it makes every internal build a complete record; a person still judges the unresolved anomalies and approves the release.
Predetermined change control plan (PCCP)
A plan, authorized with a device, that lets specified changes be made without a new submission. The FDA's final guidance for AI-enabled software, of December 2024, asks for a description of the planned modifications, a modification protocol and an impact assessment of their benefits and risks, all within the intended use.
Predictive value
What a device's answer means in a given population: the positive predictive value (PPV) is the share of positive answers that are right, the negative predictive value (NPV) the share of negative answers that are right. Both change with prevalence: IDx-DR's positive answers are right about 72 percent of the time at 24 percent prevalence and about 39 percent at 7 percent.
Premarket approval (PMA)
The US route to market for class III devices. A predetermined change control plan can be authorized through it, as through a 510(k) or a De Novo.
Prevalence
How common a condition is in the population tested: about 24 percent for more than mild retinopathy among the IDx-DR trial's analyzed participants, 7 percent for sepsis in the Michigan hospitalizations. It leaves sensitivity and specificity unchanged but changes predictive values, and a change in it is one kind of dataset shift.
Quality Management System Regulation (QMSR)
The FDA's quality rule for devices under its name since 2 February 2026, when it began to incorporate ISO 13485:2016 by reference and the agency's inspectors stopped using their old inspection technique.
Random and systematic failure
A random failure is one of a hardware part, which wears and fails at its own moment, so that a rate per hour describes a population; a systematic failure occurs every time the same conditions recur, in every copy of the same version, which is how software fails. Two copies of a program fed the same input fail together, so redundancy does not protect against a software fault.
Reader study
A study of clinicians reading cases with and without an assistive device. In the fully crossed design of the FDA's guidance, every reader reads every case in both conditions, sessions are separated by at least four weeks to avoid memory bias, and the primary measure is usually the area under the ROC curve with the aid against without it.
Reference standard
The best available judgment of each case's true state, against which a device is scored, often called ground truth. For IDx-DR it was the majority grade of three masked readers at the Wisconsin reading center, from four stereo photograph pairs of each retina and a macular scan; an imperfect reference lowers a perfect model's apparent accuracy and can hide errors both share.
Regression testing
Re-running the tests that cover what a change touched, and for anything beyond a trivial change the whole system-level suite, so that behavior verified before the change is verified again after it. Because 79 percent of the software recalls the FDA counted in 1992 to 1998 followed changes after release, it is the most important habit in maintenance.
Release
In agile work, a period of one to several months ending with software that could be delivered. A release to patients is a regulated event: IEC 62304 requires it to be verified as complete, its known anomalies evaluated, the released version archived and reproducible, and the release approved.
Reproducible build
A build in which rebuilding a past release from its recorded inputs gives a bit-for-bit identical result, so that the version tested and the version shipped can be shown to be the same.
Responsibility agreement
A written statement of which party, the hospital, the IT supplier or the device maker, does what across the life of a device's connection to a hospital's systems: who tests the interfaces, who approves changes on each side, who is told when a scanner or the archive is updated, and what clinicians do when the AI is unavailable.
ROC curve
Receiver operating characteristic curve: a plot of sensitivity against the false-positive rate at every possible threshold of a model's score, showing the trade between missed cases and false alarms. The Google retinal network's curve carried a high-specificity point, 90.3 percent sensitivity at 98.1 percent specificity, and a high-sensitivity point, 97.5 percent at 93.4 percent.
Root of trust
The protected start-up code and key, held in memory that cannot be rewritten, at the start of a secure boot chain. It must be in the hardware from the first unit shipped, because a device built without it cannot gain one by update.
Rule 11
The EU device regulation's rule for classifying software: IIa if it informs diagnostic or therapeutic decisions, IIb if those decisions could cause a serious deterioration of health or a surgical intervention, III if they could cause death or an irreversible deterioration. Monitoring software is IIa, or IIb for vital parameters whose variations could mean immediate danger; all other software is class I.
Rule of three
When no failures are seen in a number of independent, realistic trials, the upper bound of the 95 percent confidence interval for the failure probability is about three divided by that number. Showing a rate of 1 failure in a billion hours this way would take about three billion failure-free hours, some 342,000 years.
SBOM
Software bill of materials: a machine-readable list of every component in a piece of software, with supplier, name, version and dependencies, now extended to models and datasets. Required for US cyber devices since March 2023, it tells a maker or a hospital which devices contain a newly vulnerable component within hours rather than weeks.
Section 524B
The section of the US Food, Drug, and Cosmetic Act, in effect since 29 March 2023, that requires the maker of a cyber device to file a plan for monitoring and addressing vulnerabilities after release, including coordinated disclosure; to keep processes that make the device secure, with patches on a regular cycle and out of cycle for critical vulnerabilities; and to provide an SBOM.
Secure boot
A start-up chain in which each stage checks the next before handing over control. At every power-on, protected code recomputes the firmware's hash, a fingerprint that changes completely if a single bit changes, checks it against the maker's signature made with a private key, and runs the firmware only if the two agree, so altered code never runs.
Secure product development framework
The FDA's name for a development process with security built in, from threat modeling and security requirements to security testing and the handling of vulnerabilities after release; its 2026 guidance names IEC 81001-5-1 and IEC 62443-4-1 as examples that can satisfy it.
Segregation
In IEC 62304, any mechanism that prevents one software item from negatively affecting another: a separate processor, operating-system memory protection, or checks on data crossing a boundary. Demonstrated segregation lets the strictest class apply only to the dangerous code, in the guide's pump example 8,000 of 200,000 lines.
Sensitivity
The share of people with the condition whom a device flags: 173 of 198, 87.4 percent, in the IDx-DR trial. Its precision depends on the number of diseased patients tested, not on the total.
Significant change
Under the EU's device rules, a change that brings in the notified body. The coordination group's software chart counts a new or modified architecture, database structure or algorithm, a new operating system or component or a major change to one, and a new feature, interoperability channel or user interface as significant, but not safety-neutral bug fixes, security updates or cosmetic changes.
Silent trial
Running a model on live clinical data with its outputs hidden from clinicians, so that its local performance is measured before anyone acts on it. At the Hospital for Sick Children in Toronto one took a kidney-ultrasound model from an area under the curve of 0.90 to 0.50; reprocessing the live images, which arrived as unprocessed PNG files rather than processed JPEGs, restored most of the loss.
Software as a medical device
Software that is a medical device on its own, running on ordinary computers, phones or cloud servers rather than inside hardware; IDx-DR's analysis program is an example. India's 2026 guidance drops the term in favor of standalone software.
Software requirement
A statement of what the software must do, or must never do, written so that a test, an inspection or an analysis can show whether it holds. A display that is easy to read is not one; a dose shown with its unit in characters at least 5 mm high until the user confirms it is.
Software safety class
IEC 62304's class A, B or C, set by the worst harm software could contribute to once risk controls outside it are counted: A where no unacceptable risk remains, B for non-serious injury, C for death or serious injury. Until classified, software is treated as class C; the class sets how much process the standard demands, and only controls outside the software can lower it.
Software system, item and unit
IEC 62304's division of a program: the software system is divided into software items, any identifiable parts, and these in turn until they reach software units, items not divided further. IDx-DR has three separately versioned items: the client, the service that passes images and results, and the analysis.
SOUP
Software of unknown provenance: in IEC 62304, a software item already developed and generally available that was not developed for the device, or one developed earlier without adequate records. For each, a maker states the requirements the device places on it and its exact version, evaluates its published anomalies and keeps it under configuration control.
Special controls
The requirements the FDA sets for a device type when it creates one, which every later device of that type must meet. IDx-DR's included a protocol on which changes could significantly affect safety or effectiveness and a cybersecurity vulnerability and management process.
Specificity
The share of people without the condition whom a device clears: 556 of 621, 89.5 percent, in the IDx-DR trial, and about 79 percent in studies of young people outside its label.
Subgroup analysis
Reporting a device's performance separately for groups within the intended population, such as by sex, age, ethnicity, skin tone, site or equipment, each with its confidence interval, with the groups named in the protocol before the data are seen. Dermatology models that had scored 0.88 to 0.94 on their own test sets fell as low as chance on the darkest skin.
Substantial modification
Under the EU's AI Act, a change to an AI system after it is placed on the market that was not foreseen in its initial conformity assessment and affects its compliance or modifies its intended purpose; a high-risk system that undergoes one needs a new conformity assessment. Changes predetermined and documented at the initial assessment are not substantial modifications.
Threat model
A list of a device's entry points, the assets behind them, such as dose settings, patient data and the software itself, and the adversaries who might want them, with what an attacker could do on each path and what would stop it. The FDA expects it to be part of the design, updated as the architecture changes.
Threshold
The value that turns a model's score into a decision, refer above it and not below. Moving it trades missed cases against false alarms; the point chosen, the operating point, is fixed from the clinical costs before the test set is opened.
Traceability
The links from each hazard to the requirements that control it, from each requirement to the design that implements it and the tests that verify it, and from each test to its result, so that a reviewer can follow any risk to its control and its evidence.
Training
Finding a model's parameters from examples whose right answers are known: the model scores a batch, a measure of its error called the loss is computed, and every parameter is nudged in the direction that would have made the loss smaller, over millions of batches.
Training, tuning and test sets
The three separate bodies of data behind a model: the training set its parameters are fitted to, the tuning set used for choices training does not make (which engineers usually call the validation set), and the test set, held back and used once to estimate performance on the next patient. Only the test set, kept independent and split by patient, judges the model fairly.
Unit, integration and system testing
The three levels of verification by test: unit testing checks each smallest piece in isolation against its detailed design, integration testing checks that units and items work together across their boundaries, and system testing checks the complete software, usually on the real hardware, against the software requirements.
Unresolved anomaly
A known defect that software ships with. Submissions list each one with an assessment of its effect on safety and effectiveness, and a release may carry it only if a person has judged that it leaves no unacceptable risk.
Valid clinical association
The first of three components of clinical evidence for software in the EU and India: the software's output is associated with the clinical condition it targets. The others are technical performance, producing the intended output accurately and reliably from the input, and clinical performance, an output clinically relevant to the intended purpose.
Validation
Confirmation, by examination and objective evidence, that software specifications conform to user needs and intended uses; it asks whether the specification was right and reaches to usability and clinical evidence. In machine learning the same word names the tuning set, and a submission that uses both senses must say which it means.
Verification
Objective evidence that the outputs of a stage of development meet the requirements set for that stage, that is, that the software was built as specified. It runs at unit, integration and system level, with reviews, code inspection and static analysis finding faults without running the code.
VEX
Vulnerability Exploitability eXchange: a statement of whether a published vulnerability affects a specific product, with one of four values, not affected, affected, fixed or under investigation. An SBOM says which components a device contains; a VEX says whether a weakness in one of them can be exploited there.
Watchdog
A timer, often a separate circuit, that the software must reset at regular intervals; if the program hangs, the timer runs out and forces the device into a safe state. A program that runs on time while computing a wrong dose keeps resetting it, so a watchdog cannot catch a wrong output.
§ Sources

Sources.

Each source is listed under the chapter whose text first relies on it, with the opening scene and the front matter first. Company documents, papers by a company's founders or staff, and secondary reports are marked as such. Regulatory status, standards editions and product facts are as of October 2026.

  1. Opening — Abràmoff M.D., Lavin P.T., Birch M., Shah N., Folk J.C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digit Med 1, 39 (2018). doi:10.1038/s41746-018-0040-6 (first author founded the company; trial funded by IDx)
  2. Opening — US Food and Drug Administration. De Novo DEN180001, IDx-DR (IDx, LLC): database record and decision summary; received 12 Jan 2018, granted 11 Apr 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN180001; accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf
  3. Opening — Digital Diagnostics. IDx rebrands to lead AI health care revolution (5 Oct 2020). digitaldiagnostics.com/idx-rebrands-to-lead-ai-health-care-revolution/ (manufacturer)
  4. Opening — US Food and Drug Administration. FDA permits marketing of artificial intelligence-based device to detect certain diabetes-related eye problems. Press release, 11 Apr 2018. drugdiscoverytrends.com/fda-permits-marketing-of-ai-based-device-to-detect-certain-diabetes-related-eye-problems/ (read via the drugdiscoverytrends.com reprint; the fda.gov page returned 404)
  5. Opening — US Food and Drug Administration. General Principles of Software Validation; Final Guidance for Industry and FDA Staff. 11 Jan 2002 (Section 6 since superseded by the computer software assurance guidance). fda.gov/media/73141/download
  6. Ch. 1 — US Code of Federal Regulations (eCFR). 21 CFR 886.1100, Retinal diagnostic software device (codified at 87 FR 3205, 21 Jan 2022). ecfr.gov/current/title-21/chapter-I/subchapter-H/part-886/subpart-B/section-886.1100
  7. Ch. 1 — Central Drugs Standard Control Organisation, India. Guidance document on Medical Device Software under MDR-2017. Doc No. CDSCO/MD/GD/MDSW/01/2026. 21 Jul 2026. cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/Guidance-document-on-Medical-Device-Software-under-MDR-2017.pdf (text read to section 12.5.2)
  8. Ch. 1 — US Food and Drug Administration. Changes to Existing Medical Software Policies Resulting from Section 3060 of the 21st Century Cures Act. Final guidance, docket FDA-2017-D-6294. 27 Sep 2019. fda.gov/media/109622/download
  9. Ch. 1 — US Food and Drug Administration. General Wellness: Policy for Low Risk Devices. Final guidance, docket FDA-2014-N-1039. 6 Jan 2026 (supersedes the version of 27 Sep 2019). fda.gov/media/90652/download
  10. Ch. 1 — US Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final guidance, docket FDA-2022-D-2628. 4 Dec 2024, reissued 18 Aug 2025. fda.gov/media/166704/download
  11. Ch. 1 — Akin Gump Strauss Hauer & Feld. IMDRF releases international framework for regulating device software (Oct 2014), on IMDRF/SaMD WG/N12 FINAL:2014, Software as a Medical Device: Possible Framework for Risk Categorization and Corresponding Considerations (18 Sep 2014). akingump.com/en/insights/alerts/imdrf-releases-international-framework-for-regulating-device (secondary; IMDRF page not opened)
  12. Ch. 2 — US Food and Drug Administration. Ventilator Software Correction: Hamilton Medical Issues Correction for HAMILTON-C6 Medical Ventilators to Address Risk of Failed Ventilation Restart. Recall notice, posted 11 Jul 2024. fda.gov/medical-devices/medical-device-recalls-and-early-alerts/ventilator-software-correction-hamilton-medical-issues-correction-hamilton-c6-medical-ventilators
  13. Ch. 2 — US Food and Drug Administration. Medical Device Recalls: Class 1 recall Z-2020-2024, ventilator HAMILTON-C6 (Hamilton Medical), software versions 1.1.4–1.1.6; initiated 15 May 2024, posted 18 Jun 2024. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?id=207956
  14. Ch. 2 — Wallace D.R., Kuhn D.R. Failure modes in medical device software: an analysis of 15 years of recall data. Int J Reliab Qual Saf Eng 8, 351–371 (2001). doi:10.1142/S021853930100058X (read via the NIST copy, tsapps.nist.gov/publication/get_pdf.cfm?pub_id=917180)
  15. Ch. 2 — Wallace D.R., Kuhn D.R. Lessons from 342 medical device failures (NIST; conference version, c. 1999) (read via inf.ed.ac.uk/teaching/courses/seoc/2006_2007/resources/CS_342failures.pdf)
  16. Ch. 2 — US Food and Drug Administration. Tandem Diabetes Care, Inc. Recalls Version 2.7 of the Apple iOS t:connect Mobile App… (used with the t:slim X2 insulin pump). Recall notice, posted 28 Aug 2024. fda.gov/medical-devices/medical-device-recalls-and-early-alerts/tandem-diabetes-care-inc-recalls-version-27-apple-ios-tconnect-mobile-app-used-conjunction-tslim-x2
  17. Ch. 2 — US Food and Drug Administration. Medical Device Recalls: Class 1 recall Z-1609-2024, t:connect mobile app (Tandem Diabetes Care), root cause recorded as software design change; initiated Mar 2024, posted 6 May 2024. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?id=206914
  18. Ch. 3 — Dijkstra E.W. The humble programmer (ACM Turing Lecture 1972; EWD340). Commun ACM 15, 859–866 (1972). doi:10.1145/355604.361591 (read as the EWD340 transcription, cs.utexas.edu/~EWD/transcriptions/EWD03xx/EWD340.html; journal details not checked)
  19. Ch. 3 — Eypasch E., Lefering R., Kum C.K., Troidl H. Probability of adverse events that have not yet occurred: a statistical reminder. BMJ 311, 619 (1995). doi:10.1136/bmj.311.7005.619 (abstract read)
  20. Ch. 3 — Butler R.W., Finelli G.B. The infeasibility of quantifying the reliability of life-critical real-time software. IEEE Trans Softw Eng 19, 3–12 (1993). doi:10.1109/32.210303 (read as the NASA preprint, ntrs.nasa.gov/api/citations/20040139817/downloads/20040139817.pdf)
  21. Ch. 3 — NASA Science. Universe overview (age of the universe, about 13.8 billion years). science.nasa.gov/universe/overview/
  22. Ch. 4 — IEC. IEC 62304:2006, Medical device software – Software life cycle processes, ed. 1.0 (9 May 2006). IEC Webstore page, webstore.iec.ch/en/publication/6792
  23. Ch. 4 — IEC. IEC 62304:2006+AMD1:2015 CSV, Medical device software – Software life cycle processes, consolidated ed. 1.1 (26 Jun 2015; stability date 2028). IEC Webstore page, webstore.iec.ch/en/publication/22794; preview pages, including the introduction to Amendment 1, read via webstore.ansi.org
  24. Ch. 4 — US Food and Drug Administration. Recognized Consensus Standards database: IEC 62304 ed. 1.1 2015-06 consolidated (and ANSI/AAMI IEC 62304:2006/A1:2016), recognition no. 13-79, complete standard; entry 14 Jan 2019. accessdata.fda.gov/scripts/cdrh/cfdocs/cfStandards/detail.cfm?standard__identification_no=38829
  25. Ch. 4 — IEC. Project records for IEC 62304 ED2, Health software – Software life cycle processes: iec:proj:122433 (stage 30.20, CD study/ballot initiated, 7 Aug 2026; next stage due 27 Nov 2026) and the earlier project iec:proj:23605 (deleted 7 Mar 2023); with iec:proj:11630 and iec:proj:21252 for editions 1.0 and 1.1. iss.rs/en/project/show/iec:proj:122433 (read via iss.rs, the Institute for Standardization of Serbia’s mirror of IEC project data; the iec.ch page was blocked)
  26. Ch. 4 — VDE. IEC 62304 Edition 2: Änderungen (VDE Health Fachinformation, 22 Jul 2026). vde.com/iec-62304-edition-2-aenderungen (secondary)
  27. Ch. 4 — Johner Institute. IEC 62304 2nd edition: all areas of application and changes (8 Jul 2026). blog.johner-institute.com/iec-62304-medical-software/iec-62304-2nd-edition-all-areas-of-application-and-changes/ (secondary)
  28. Ch. 4 — ISO. ISO 14971:2019, Medical devices – Application of risk management to medical devices, 3rd ed. (Dec 2019; confirmed 7 Mar 2025), catalogue page. iso.org/standard/72704.html
  29. Ch. 4 — International Medical Device Regulators Forum, SaMD Working Group. Characterization Considerations for Medical Device Software and Software-Specific Risk. IMDRF/SaMD WG/N81 FINAL:2025. 27 Jan 2025. imdrf.org/sites/default/files/2025-01/IMDRF_SaMD%20WG_Software-Specific%20Risk_N81%20Final_0.pdf
  30. Ch. 4 — Wikipedia. Software safety classification (web page quoting IEC 62304:2006+A1:2015, accessed 8 Oct 2026). en.wikipedia.org/wiki/Software_safety_classification (secondary)
  31. Ch. 4 — LDRA Ltd. Ease the Heartache of Medical Device Software Certification: Achieving Cost-effective Compliance with IEC 62304 – Amendment 1:2015. Technical briefing v2 (Apr 2018). ldra.com/wp-content/uploads/ldra/IEC_62304_Technical_Briefing_v2_04_18.pdf (secondary)
  32. Ch. 4 — Johner Institute. Software safety classes according to IEC 62304 (rewritten 15 Oct 2025). blog.johner-institute.com/iec-62304-medical-software/safety-class-iec-62304/ (secondary)
  33. Ch. 4 — US Food and Drug Administration. Content of Premarket Submissions for Device Software Functions. Final guidance, docket FDA-2021-D-0775. 14 Jun 2023 (supersedes the guidance of 11 May 2005). fda.gov/media/153781/download
  34. Ch. 4 — Matrix Requirements. IEC 62304:2015 impact on class A software (blog). matrixone.health/blog/iec-62304-2015-impact-on-class-a-software (secondary; checked in the subject review)
  35. Ch. 5 — US Food and Drug Administration. Infusion Pumps Total Product Life Cycle: Guidance for Industry and FDA Staff. Final guidance. 2 Dec 2014. fda.gov/media/78369/download
  36. Ch. 5 — US Food and Drug Administration. 510(k) K203629, IDx-DR (Digital Diagnostics, Inc.): database record and 510(k) summary; received 11 Dec 2020, decision 10 Jun 2021. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K203629; accessdata.fda.gov/cdrh_docs/pdf20/K203629.pdf
  37. Ch. 5 — US Food and Drug Administration. 510(k) K213037, IDx-DR v2.3 (Digital Diagnostics, Inc.): database record and 510(k) summary; received 21 Sep 2021, decision 17 Jun 2022. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K213037; accessdata.fda.gov/cdrh_docs/pdf21/K213037.pdf
  38. Ch. 6 — Microsoft. Windows XP, product lifecycle (Microsoft Learn, accessed 8 Oct 2026). learn.microsoft.com/en-us/lifecycle/products/windows-xp (manufacturer)
  39. Ch. 6 — National Audit Office (UK). Investigation: WannaCry cyber attack and the NHS. HC 414. 27 Oct 2017. nao.org.uk/reports/investigation-wannacry-cyber-attack-and-the-nhs/ (summary page read)
  40. Ch. 6 — Dworetzky T. ‘Ransomware’ attack hit U.S. medical devices, too. DOTmed News (18 May 2017). dotmed.com/legal/print/story.html?nid=37359 (secondary, reporting Bayer and Siemens Healthineers)
  41. Ch. 6 — Palo Alto Networks Unit 42. 2020 Unit 42 IoT Threat Report (10 Mar 2020). unit42.paloaltonetworks.com/iot-threat-report-2020/ (manufacturer)
  42. Ch. 6 — US Food and Drug Administration. Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions. Guidance for Industry and FDA Staff, docket FDA-2021-D-1158. 3 Feb 2026 (supersedes the version of 27 Jun 2025). fda.gov/media/119933/download
  43. Ch. 6 — Johner Institute. Off-the-shelf software (OTS) versus SOUP (web page, undated; quoting IEC 62304:2006+A1:2015). blog.johner-institute.com/iec-62304-medical-software/off-the-shelf-software-ots-versus-soup/ (secondary)
  44. Ch. 6 — US Food and Drug Administration. Off-The-Shelf Software Use in Medical Devices. Final guidance. 11 Aug 2023 (supersedes the edition of 27 Sep 2019; first issued 9 Sep 1999). fda.gov/media/71794/download
  45. Ch. 6 — Black Duck Software. 2026 Open Source Security and Risk Analysis (OSSRA) Report (Mar 2026). blackduck.com/resources/analyst-reports/open-source-security-risk-analysis.html (manufacturer)
  46. Ch. 6 — OWASP CycloneDX. CycloneDX v1.5 released (26 Jun 2023). cyclonedx.org/news/cyclonedx-v1.5-released/
  47. Ch. 6 — National Telecommunications and Information Administration (US Department of Commerce). The Minimum Elements For a Software Bill of Materials (SBOM). 12 Jul 2021. ntia.gov/files/ntia/publications/sbom_minimum_elements_report.pdf
  48. Ch. 6 — Medcrypt. Blog post on CISA’s 2026 Minimum Elements for a Software Bill of Materials (29 Jul 2026). medcrypt.com/blog/cisa-sbom-minimum-elements-2026-update (secondary; written by a device-security vendor; CISA’s own 2026 PDF was blocked and not opened)
  49. Ch. 6 — devops.com. CISA’s 2026 SBOM guidance adds hash requirements and AI coverage (31 Jul 2026). devops.com/cisas-2026-sbom-guidance-adds-hash-requirements-and-ai-coverage/ (secondary)
  50. Ch. 6 — HSToday. CISA updates software bill of materials guidance to strengthen supply chain security (2026). hstoday.us/subject-matter-areas/cybersecurity/cisa-updates-software-bill-of-materials-guidance-to-strengthen-supply-chain-security/ (secondary)
  51. Ch. 6 — ISO/IEC. ISO/IEC 5962:2021, Information technology — SPDX® Specification V2.2.1, ed. 1 (Aug 2021; to be revised), catalogue page. iso.org/standard/81870.html
  52. Ch. 6 — Ecma International. ECMA-424, CycloneDX Bill of materials specification, 1st ed. (Jun 2024) and 2nd ed. (Dec 2025). ecma-international.org/publications-and-standards/standards/ecma-424/
  53. Ch. 6 — National Telecommunications and Information Administration. Vulnerability-Exploitability eXchange (VEX) – An Overview (27 Sep 2021). ntia.gov/files/ntia/publications/vex_one-page_summary.pdf
  54. Ch. 6 — United States Code. 21 U.S.C. 360n-2, Ensuring cybersecurity of devices (FD&C Act section 524B), added by Pub. L. 117-328, div. FF, title III, sec. 3305, 29 Dec 2022; effective 29 Mar 2023. law.cornell.edu/uscode/text/21/360n-2 (read via Cornell LII)
  55. Ch. 6 — US Food and Drug Administration. Cybersecurity in Medical Devices Frequently Asked Questions (FAQs) (content current as of 26 Jun 2025). fda.gov/medical-devices/digital-health-center-excellence/cybersecurity-medical-devices-frequently-asked-questions-faqs
  56. Ch. 7 — US Food and Drug Administration. Medical Device Recalls: Class 2 recall Z-1233-2023, Monaco RTP System (Elekta), builds 5.11.00–5.11.03, root cause recorded as software design; initiated 28 Feb 2023, posted 8 Mar 2023. accessdata.fda.gov/scripts/cdrh/cfdocs/cfRes/res.cfm?id=198827
  57. Ch. 8 — AAMI. AAMI TIR45:2012 and AAMI TIR45:2023, Guidance on the use of AGILE practices in the development of medical device software (not opened; editions as listed in the FDA Recognized Consensus Standards database)
  58. Ch. 8 — US Food and Drug Administration. Recognized Consensus Standards database: AAMI TIR45:2012, recognition no. 13-36 (entry 15 Jan 2013), and AAMI TIR45:2023, recognition no. 13-143 (entry 26 May 2025; declarations of conformity to 13-36 accepted until 2 Jul 2028). accessdata.fda.gov/scripts/cdrh/cfdocs/cfStandards/detail.cfm?standard__identification_no=46298; accessdata.fda.gov/scripts/cdrh/cfdocs/cfstandards/detail.cfm?standard__identification_no=46295
  59. Ch. 8 — Johner Institute. TIR 45: Agile software development (2016, updated for the 2023 edition). blog.johner-institute.com/iec-62304-medical-software/tir-45-agile-software-development/ (secondary)
  60. Ch. 8 — AAMI. Key Updates: AAMI TIR45:2023 – Guidance on Agile Practices (training page, undated). aami.org/training/training-suites/software-cybersecurity/key-updates-aami-tir45-2023-guidance-on-agile-practices (publisher’s training page)
  61. Ch. 8 — ISPE. ISPE GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition) (Jul 2022), product page. ispe.org/publications/guidance-documents/gamp-5-guide-2nd-edition
  62. Ch. 8 — Wyn S., Clark C. What you need to know about GAMP 5 Guide, 2nd Edition. Pharmaceutical Engineering (Jan/Feb 2023). ispe.org/pharmaceutical-engineering/january-february-2023/what-you-need-know-about-gampr-5-guide-2nd-edition
  63. Ch. 8 — OpenRegulatory. IEC 62304:2006 mapping of requirements to documents (9 Jun 2022). openregulatory.com/document_templates/iec-623042006-mapping-of-requirements-to-documents (secondary; clause titles of IEC 62304, standard not opened)
  64. Ch. 8 — OpenRegulatory. Doing a software release in compliance with IEC 62304 (24 May 2023, updated 1 Oct 2024). openregulatory.com/articles/software-release-iec-62304 (secondary)
  65. Ch. 9 — Google Cloud DORA. Accelerate State of DevOps Report 2024 (22 Oct 2024). dora.dev/research/2024/dora-report/2024-dora-accelerate-state-of-devops-report.pdf (manufacturer)
  66. Ch. 9 — US Food and Drug Administration. Quality Management System Regulation (QMSR) web page; and Medical Devices; Quality System Regulation Amendments, final rule, FR Doc. 2024-01709 (2 Feb 2024; effective 2 Feb 2026). fda.gov/medical-devices/postmarket-requirements-devices/quality-management-system-regulation-qmsr
  67. Ch. 9 — ISO. ISO/TR 80002-2:2017, Medical device software — Part 2: Validation of software for medical device quality systems, ed. 1 (13 Jun 2017), catalogue page. iso.org/standard/60044.html
  68. Ch. 9 — US Food and Drug Administration. Computer Software Assurance for Production and Quality System Software. Final guidance, docket FDA-2022-D-0795. 24 Sep 2025. fda.gov/regulatory-information/search-fda-guidance-documents/computer-software-assurance-production-and-quality-system-software
  69. Ch. 9 — US Food and Drug Administration. Computer Software Assurance for Production and Quality Management System Software. Final guidance, docket FDA-2022-D-0795. 3 Feb 2026 (supersedes the guidance of 24 Sep 2025 and Section 6 of the 2002 software validation guidance). fda.gov/media/188844/download
  70. Ch. 9 — Johner Institute. IT security for legacy devices (web page, undated; quoting the IEC 62304 Amendment 1 definition of legacy software). blog.johner-institute.com/iec-62304-medical-software/it-security-for-legacy-devices/ (secondary)
  71. Ch. 9 — Spyro-soft. How to implement a legacy software gap analysis as required by IEC 62304:2015 (blog, undated). spyro-soft.com/blog/how-to-implement-a-legacy-software-gap-analysis-as-required-by-iec-62304-2015 (secondary)
  72. Ch. 10 — Gulshan V., Peng L., Coram M., et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 316, 2402–2410 (2016). doi:10.1001/jama.2016.17216 (authors include company staff)
  73. Ch. 10 — PyTorch. torchvision.models.resnet50, documentation (stable release, accessed 8 Oct 2026). docs.pytorch.org/vision/stable/models/generated/torchvision.models.resnet50.html
  74. Ch. 10 — US Food and Drug Administration. De Novo DEN190040, Caption Guidance (Bay Labs, Inc.): database record and decision summary; received 27 Aug 2019, granted 7 Feb 2020. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN190040; accessdata.fda.gov/cdrh_docs/reviews/DEN190040.pdf
  75. Ch. 10 — US Food and Drug Administration. 510(k) K201992, Caption Guidance (Caption Health, Inc.): 510(k) summary; decision 18 Sep 2020. accessdata.fda.gov/cdrh_docs/pdf20/K201992.pdf
  76. Ch. 10 — PIC/S and EMA GMP/GDP Inspectors Working Group. Draft Annex 22: Artificial Intelligence. Consultation draft (Jul 2025). picscheme.org/docview/9715
  77. Ch. 10 — US Food and Drug Administration. Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) – Discussion Paper and Request for Feedback. 2 Apr 2019. fda.gov/media/122535/download
  78. Ch. 11 — Zech J.R., Badgeley M.A., Liu M., Costa A.B., Titano J.J., Oermann E.K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLOS Med 15, e1002683 (2018). doi:10.1371/journal.pmed.1002683 (authors include company staff)
  79. Ch. 11 — International Medical Device Regulators Forum, AI/ML-enabled Working Group. Good machine learning practice for medical device development: Guiding principles. IMDRF/AIML WG/N88 FINAL:2025. 27 Jan 2025. imdrf.org/sites/default/files/2025-02/IMDRF_AIML%20WG_GMLP_N88%20Final.pdf
  80. Ch. 12 — Krause J., Gulshan V., Rahimy E., Karth P., Widner K., Corrado G.S., et al. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology 125, 1264–1272 (2018). doi:10.1016/j.ophtha.2018.01.034 (read as arXiv:1710.01711v3; journal version not opened) (authors include company staff)
  81. Ch. 12 — Medicines and Healthcare products Regulatory Agency, US Food and Drug Administration, Health Canada. Good machine learning practice for medical device development: guiding principles (27 Oct 2021). gov.uk/government/publications/good-machine-learning-practice-for-medical-device-development-guiding-principles
  82. Ch. 12 — US Food and Drug Administration. De Novo DEN170073, ContaCT (Viz.ai, Inc.): database record and decision summary; received 29 Sep 2017, granted 13 Feb 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN170073; accessdata.fda.gov/cdrh_docs/reviews/DEN170073.pdf
  83. Ch. 12 — Daneshjou R., et al. Diverse Dermatology Images (DDI) dataset, project page (accessed 8 Oct 2026). ddi-dataset.github.io/
  84. Ch. 13 — NIST/SEMATECH. e-Handbook of Statistical Methods, section 7.2.4.1, Confidence intervals (web page, undated). itl.nist.gov/div898/handbook/prc/section2/prc241.htm
  85. Ch. 13 — MedCalc Software. Sensitivity and specificity (MedCalc manual, web page, undated). medcalc.org/en/book/sensitivity-specificity.php (secondary)
  86. Ch. 13 — Wong A., Otles E., Donnelly J.P., Krumm A., McCullough J., DeTroyer-Cooley O., et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med 181, 1065–1070 (2021). doi:10.1001/jamainternmed.2021.2626
  87. Ch. 14 — Wolf R.M., Channa R., Liu T.Y.A., et al. Autonomous artificial intelligence increases screening and follow-up for diabetic retinopathy in youth: the ACCESS randomized control trial. Nat Commun 15, 421 (2024). doi:10.1038/s41467-023-44676-z (authors include the company’s founder; youth accuracy of the earlier SEE study as reported in this paper)
  88. Ch. 14 — Heaven W.D. Google’s medical AI was super accurate in a lab. Real life was a different story. MIT Technology Review (27 Apr 2020). technologyreview.com/2020/04/27/1000658/google-medical-ai-accurate-lab-real-life-clinic-covid-diabetes-retina-disease/ (secondary)
  89. Ch. 14 — Beede E. Healthcare AI systems that put people at the center. Google Keyword blog (25 Apr 2020). blog.google/technology/health/healthcare-ai-systems-put-people-center/ (manufacturer)
  90. Ch. 14 — Beede E., Baylor E., Hersch F., Iurchenko A., Wilcox L., Ruamviboonsuk P., et al. A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (2020). research.google/pubs/a-human-centered-evaluation-of-a-deep-learning-system-deployed-in-clinics-for-the-detection-of-diabetic-retinopathy/ (abstract read) (authors include company staff)
  91. Ch. 14 — Ruamviboonsuk P., Tiwari R., Sayres R., Nganthavee V., Hemarat K., Kongprayoon A., et al. Real-time diabetic retinopathy screening by deep learning in a multisite national screening programme: a prospective interventional cohort study. Lancet Digit Health 4, e235–e244 (2022). doi:10.1016/S2589-7500(22)00017-6 (abstract read, via DOAJ) (authors include company staff; funded by Google and Rajavithi Hospital)
  92. Ch. 15 — Daneshjou R., et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. arXiv:2203.08807 v1 (Mar 2022). arxiv.org/pdf/2203.08807 (preprint)
  93. Ch. 15 — Daneshjou R., Vodrahalli K., Novoa R.A., et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv 8, eabq6147 (2022). doi:10.1126/sciadv.abq6147 (abstract read)
  94. Ch. 15 — Seyyed-Kalantari L., Zhang H., McDermott M.B.A., Chen I.Y., Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med 27, 2176–2182 (2021). doi:10.1038/s41591-021-01595-0 (text read; tables not accessible)
  95. Ch. 15 — US Food and Drug Administration. Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests. Final guidance. 13 Mar 2007. fda.gov/media/71147/download
  96. Ch. 15 — Emergo by UL. India CDSCO finalizes guidance on medical device software (30 Jul 2026). emergobyul.com/news/india-cdsco-finalizes-guidance-medical-device-software (secondary)
  97. Ch. 15 — Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act). OJ L (12 Jul 2024). artificialintelligenceact.eu/article/3/ (Articles 3, 4, 6, 8–15, 26, 43 and 113, as amended, read via artificialintelligenceact.eu, a transcription of the OJ text)
  98. Ch. 15 — Emergo by UL. Decoding MDCG 2025-6: interplay between the MDR/IVDR and the AI Act (24 Jun 2025). emergobyul.com/news/decoding-mdcg-2025-6-interplay-between-mdrivdr-and-aia (secondary)
  99. Ch. 16 — Dratsch T., Chen X., Rezazade Mehrizi M., Kloeckner R., Mähringer-Kunz A., Püsken M., et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology 307, e222176 (2023). doi:10.1148/radiol.222176 (abstract read)
  100. Ch. 16 — US Food and Drug Administration. Clinical Decision Support Software. Final guidance, docket FDA-2017-D-6569. 6 Jan 2026, corrected 29 Jan 2026 (supersedes the September 2022 version); with the guidance web page. fda.gov/media/109618/download; fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software
  101. Ch. 16 — US Food and Drug Administration. Clinical Performance Assessment: Considerations for Computer-Assisted Detection Devices Applied to Radiology Images and Radiology Device Data in Premarket Notification (510(k)) Submissions. Final guidance, docket FDA-2009-D-0503. 28 Sep 2022. fda.gov/media/77642/download
  102. Ch. 16 — US Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles (13 Jun 2024). fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles
  103. Ch. 17 — Digital Diagnostics. LumineticsCore product page (last modified 5 Aug 2026; accessed 8 Oct 2026). digitaldiagnostics.com/idx-dr/ (manufacturer)
  104. Ch. 18 — NEMA, DICOM Standards Committee. DICOM PS3.6 2026d, Data Dictionary, chapter 6: Registry of DICOM Data Elements. dicom.nema.org/medical/dicom/current/output/chtml/part06/chapter_6.html
  105. Ch. 18 — HL7 International. HL7 Version 2 Product Suite, product brief (web page, undated). hl7.org/implement/standards/product_brief.cfm?product_id=185
  106. Ch. 18 — HL7 International. FHIR Publication (Version) History (web page, accessed 8 Oct 2026). hl7.org/fhir/history.html
  107. Ch. 18 — HL7 International. FHIR R5 (v5.0.0), Resource (26 Mar 2023). hl7.org/fhir/R5/resource.html
  108. Ch. 18 — HL7 International. FHIR v6.0.0-ballot5, CodeSystem FHIR-version (page generated 17 Jul 2026). hl7.org/fhir/6.0.0-ballot5/codesystem-FHIR-version.html
  109. Ch. 18 — IHE Radiology Technical Committee. IHE Radiology Technical Framework Supplement: AI Results (AIR), Rev. 1.3, Trial Implementation. 8 Aug 2025. ihe.net/uploadedFiles/Documents/Radiology/IHE_RAD_Suppl_AIR.pdf
  110. Ch. 18 — IHE Radiology Technical Committee. IHE Radiology Technical Framework Supplement: AI Workflow for Imaging (AIW-I), Rev. 1.1, Trial Implementation. 6 Aug 2020. ihe.net/uploadedFiles/Documents/Radiology/IHE_RAD_Suppl_AIW-I.pdf
  111. Ch. 18 — Martinez-Gutierrez J.C., Kim Y., Salazar-Marioni S., et al. Automated large vessel occlusion detection software and thrombectomy treatment times: a cluster randomized clinical trial. JAMA Neurol 80, 1182–1190 (2023). doi:10.1001/jamaneurol.2023.3206 (abstract read, via scholars.aku.edu)
  112. Ch. 18 — UTHealth Houston, McWilliams School of Biomedical Informatics. News release on the study of artificial intelligence software and endovascular thrombectomy treatment times (2023). sbmi.uth.edu/news/story/uthealth-houston-study-artificial-intelligence-software-improves-endovascular-thrombectomy-treatment-times-for-stroke-patients (secondary)
  113. Ch. 18 — Medical Device Coordination Group. MDCG 2019-16 rev.1, Guidance on Cybersecurity for medical devices. Dec 2019, rev.1 Jul 2020 (quoting Regulation (EU) 2017/745, Annex I, sections 17.2 and 17.4). health.ec.europa.eu/system/files/2022-01/md_cybersecurity_en.pdf
  114. Ch. 18 — IEC. IEC 80001-1:2021, Application of risk management for IT-networks incorporating medical devices – Part 1: Safety, effectiveness and security in the implementation and use of connected medical devices or connected health software, ed. 2.0 (21 Sep 2021). IEC Webstore page, webstore.iec.ch/en/publication/34263; ISO catalogue page (stage 90.92, to be revised, 4 Feb 2026), iso.org/standard/72026.html
  115. Ch. 18 — ISO. ISO/TR 80001-2-6:2014, Application of risk management for IT-networks incorporating medical devices – Part 2-6: Application guidance – Guidance for responsibility agreements (catalogue entry read via implementer.digitalhealth.gov.au, Australian Digital Health Agency)
  116. Ch. 19 — Brady A.P., Allen B., Chong J., Kotter E., Kottler N., Mongan J., et al. Developing, purchasing, implementing and monitoring AI tools in radiology: practical considerations. A multi-society statement from the ACR, CAR, ESR, RANZCR & RSNA. Can Assoc Radiol J 75, 226–244 (2024). doi:10.1177/08465371231222229
  117. Ch. 19 — Kwong J.C.C., Erdman L., Khondker A., Skreta M., Goldenberg A., McCradden M.D., et al. The silent trial: the bridge between bench-to-bedside clinical AI applications. Front Digit Health 4, 929508 (2022). doi:10.3389/fdgth.2022.929508
  118. Ch. 19 — IntuitionLabs. GAMP 5 categories explained (16 Oct 2025, updated 8 Aug 2026). intuitionlabs.ai/articles/gamp-5-categories-explained (secondary)
  119. Ch. 19 — US Code of Federal Regulations (eCFR). 21 CFR Part 11, Electronic Records; Electronic Signatures (62 FR 13464, 20 Mar 1997). ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
  120. Ch. 19 — Blumenthal R., Erdmann N., Heitmann M., Lemettinen A.-L., Stockton B.M. Machine learning risk and control framework. Pharmaceutical Engineering (Jan/Feb 2024). ispe.org/pharmaceutical-engineering/january-february-2024/machine-learning-risk-and-control-framework
  121. Ch. 19 — ISPE. ISPE GAMP Guide: Artificial Intelligence (Jul 2025), product page. ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence
  122. Ch. 19 — Stockton B., Staib E., Heitmann M. New GAMP Guide addresses challenges posed by AI-enabled computerized systems. Pharmaceutical Engineering (Sep/Oct 2025). ispe.org/pharmaceutical-engineering/september-october-2025/new-gampr-guide-addresses-challenges-posed-ai
  123. Ch. 19 — European Commission, DG SANTE. Stakeholders’ Consultation on EudraLex Volume 4 – Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22 (7 Jul to 7 Oct 2025). health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en
  124. Ch. 19 — European Commission. EudraLex – Volume 4, Good Manufacturing Practice guidelines (web page, accessed 8 Oct 2026). health.ec.europa.eu/medicinal-products/eudralex/eudralex-volume-4_en
  125. Ch. 20 — Michigan Medicine Health Lab. Study of 24 U.S. hospitals shows onset of COVID-19 led to spike in sepsis alerts (6 Dec 2021), reporting Wong A., et al. Quantification of sepsis model alerts in 24 US hospitals before and during the COVID-19 pandemic. JAMA Netw Open 4, e2135286 (2021). doi:10.1001/jamanetworkopen.2021.35286. michiganmedicine.org/health-lab/study-24-us-hospitals-shows-onset-covid-19-led-spike-sepsis-alerts (secondary; the paper itself was blocked and not opened)
  126. Ch. 20 — Davis S.E., Lasko T.A., Chen G., Siew E.D., Matheny M.E. Calibration drift in regression and machine learning models for acute kidney injury. J Am Med Inform Assoc 24, 1052–1061 (2017). doi:10.1093/jamia/ocx030 (abstract read)
  127. Ch. 20 — US Food and Drug Administration, CDRH Digital Health Center of Excellence. Request For Public Comment: Measuring and Evaluating Artificial Intelligence-enabled Medical Device Performance in the Real-World. Docket FDA-2025-N-4203. 30 Sep 2025. fda.gov/medical-devices/digital-health-center-excellence/request-public-comment-measuring-and-evaluating-artificial-intelligence-enabled-medical-device
  128. Ch. 21 — US Food and Drug Administration. Artificial Intelligence in Software as a Medical Device (web page with milestones, content current as of 25 Mar 2025). fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-software-medical-device
  129. Ch. 21 — US Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Predetermined Change Control Plans for Machine Learning-Enabled Medical Devices: Guiding Principles (Oct 2023; web page republished 18 Aug 2025). fda.gov/medical-devices/software-medical-device-samd/predetermined-change-control-plans-machine-learning-enabled-medical-devices-guiding-principles
  130. Ch. 21 — US Food and Drug Administration. Predetermined Change Control Plans for Medical Devices. Draft guidance, docket FDA-2024-D-2338. Aug 2024. fda.gov/regulatory-information/search-fda-guidance-documents/predetermined-change-control-plans-medical-devices
  131. Ch. 21 — Dayma K., Patel P., Hildreth K., Jamaspishvili T. Predetermined change control plan adoption and documentation transparency in U.S. Food and Drug Administration–cleared radiology artificial intelligence/machine learning devices. Radiol Artif Intell 8(5) (online 29 Jul 2026). doi:10.1148/ryai.260385 (abstract read)
  132. Ch. 21 — US Food and Drug Administration. 510(k) K253281, UpDoc (Updoc, Inc.): database record and 510(k) summary; received 29 Sep 2025, decision 23 Dec 2025. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K253281; accessdata.fda.gov/cdrh_docs/pdf25/K253281.pdf
  133. Ch. 22 — Abbott. Important Cybersecurity Advisory: pacemaker firmware update, letter to physicians (US), 28 Aug 2017. cardiovascular.abbott/content/dam/cv/cardiovascular/pdf/reports/Pacemaker-Firmware-Update-Doctor-Letter-Aug2017-US.pdf (manufacturer)
  134. Ch. 22 — US Food and Drug Administration. Medical Device Recalls: recall event 78093 (St. Jude Medical pacemakers, Merlin PCS programmer software and Merlin@home software; records Z-0029 to Z-0038-2018, Class 2), initiated 28 Aug 2017, posted 12 Jun 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?start_search=1&event_id=78093
  135. Ch. 22 — American College of Cardiology. FDA approves firmware addressing cybersecurity vulnerabilities in Abbott implantable pacemakers (31 Aug 2017). acc.org/latest-in-cardiology/articles/2017/08/31/12/13/fda-approves-firmware-addressing-cybersecurity-vulnerabilities-in-abbott-implantable-pacemakers (secondary)
  136. Ch. 22 — CSO Online. 465,000 Abbott pacemakers vulnerable to hacking, need a firmware fix (4 Sep 2017). csoonline.com/article/3222068/465000-abbott-pacemakers-vulnerable-to-hacking-need-a-firmware-fix.html (secondary)
  137. Ch. 22 — US Food and Drug Administration. FDA warns patients and health care providers about potential cybersecurity concerns with certain Medtronic insulin pumps. Press release, 27 Jun 2019. biospace.com/fda-warns-patients-and-health-care-providers-about-potential-cybersecurity-concerns-with-certain-medtronic-insulin-pumps (read via the BioSpace reprint; the FDA safety communication returned 404)
  138. Ch. 22 — US Food and Drug Administration. Medical Device Recalls: Class 2 recall Z-1581-2020, MiniMed insulin pump MMT-508 (Medtronic), initiated 27 Jun 2019, posted 26 Mar 2020, status open; and recall event 83433 (16 records, MiniMed 508 and Paradigm models). accessdata.fda.gov/scripts/cdrh/cfdocs/cfRES/res.cfm?id=175194; accessdata.fda.gov/scripts/cdrh/cfdocs/cfRes/res.cfm?start_search=1&event_id=83433
  139. Ch. 22 — Cybersecurity and Infrastructure Security Agency (NCCIC). ICS Medical Advisory ICSMA-19-178-01: Medtronic MiniMed 508 and Paradigm Series Insulin Pumps (27 Jun 2019, date inferred from the advisory number). cisa.gov/news-events/ics-medical-advisories/icsma-19-178-01
  140. Ch. 22 — Cybersecurity and Infrastructure Security Agency (NCCIC). ICS Medical Advisory ICSMA-18-219-02: Medtronic MiniMed 508 and Paradigm Series insulin pumps, remote controllers (7 Aug 2018; updated). cisa.gov/news-events/ics-medical-advisories/icsma-18-219-02
  141. Ch. 22 — US Food and Drug Administration. Cybersecurity (Digital Health Center of Excellence web page, content current as of 6 Jul 2026), listing the Playbook for Threat Modeling Medical Devices (30 Nov 2021). fda.gov/medical-devices/digital-health-center-excellence/cybersecurity
  142. Ch. 22 — US Food and Drug Administration. Cybersecurity Vulnerabilities with Certain Patient Monitors from Contec and Epsimed: FDA Safety Communication. 30 Jan 2025, updated 2 Jul 2025. fda.gov/medical-devices/safety-communications/cybersecurity-vulnerabilities-certain-patient-monitors-contec-and-epsimed-fda-safety-communication
  143. Ch. 22 — MITRE and Medical Device Innovation Consortium. Playbook for Threat Modeling Medical Devices (Nov 2021; FDA-funded). Release: mitre.org/news-insights/news-release/mitre-and-medical-device-innovation-consortium-create-playbook-threat; mdic.org/resource/playbook-for-threat-modeling-medical-devices/
  144. Ch. 23 — Regenscheid A. Platform Firmware Resiliency Guidelines. NIST SP 800-193. May 2018. csrc.nist.gov/pubs/sp/800/193/final (abstract page read)
  145. Ch. 23 — IEC. IEC 81001-5-1:2021, Health software and health IT systems safety, effectiveness and security – Part 5-1: Security – Activities in the product life cycle, ed. 1.0 (Dec 2021). standards.iteh.ai/catalog/standards/iso/227a3148-4206-42e3-ad21-05f80dda507a/iec-81001-5-1-2021 (standard’s front matter read via a reseller preview)
  146. Ch. 23 — ISO. IEC 81001-5-1:2021, catalogue page (published 21 Dec 2021; stage 90.92, to be revised, since 19 Sep 2025). iso.org/standard/76097.html
  147. Ch. 23 — Blue Goat Cyber. Did AAMI SW96 replace TIR57? (blog post, 21 Jul 2026). bluegoatcyber.com/blog/did-aami-sw96-replace-tir57-fda-2026 (secondary)
  148. Ch. 23 — AAMI. ANSI/AAMI SW96:2023, Standard for medical device security – Security risk management for device manufacturers (not opened; year and alignment with ISO 14971 as reported by AHA News and MedTech Intelligence)
  149. Ch. 23 — AHA News. FDA recognizes ANSI/AAMI medical device standard to enhance cybersecurity (8 Nov 2023). aha.org/news/headline/2023-11-08-fda-recognizes-ansiaami-medical-device-standard-enhance-cybersecurity (secondary)
  150. Ch. 23 — MedTech Intelligence. FDA recognizes AAMI SW96 cybersecurity guidance document (15 Nov 2023). medtechintelligence.com/news_article/fda-recognizes-aami-sw96-cybersecurity-guidance-document/ (secondary)
  151. Ch. 24 — Armis. URGENT/11 (research page; updated 1 Oct 2019 and 15 Dec 2020). armis.com/research/urgent-11/ (manufacturer)
  152. Ch. 24 — US Food and Drug Administration. FDA informs patients, providers and manufacturers about potential cybersecurity vulnerabilities for connected medical devices and health care networks that use certain communication software. Press release, 1 Oct 2019. fda.gov/news-events/press-announcements/fda-informs-patients-providers-and-manufacturers-about-potential-cybersecurity-vulnerabilities
  153. Ch. 24 — Whooley S. B. Braun, Baxter, Carestream, Green Hills affected by Ripple20 cyber vulnerabilities. MassDevice (29 Jun 2020). massdevice.com/b-braun-baxter-carestream-green-hills-affected-by-ripple20-cyber-vulnerabilities/ (secondary)
  154. Ch. 24 — FIRST. FIRST has officially published the latest version of CVSS (v4.0). Press release, 1 Nov 2023. first.org/newsroom/releases/20231101
  155. Ch. 24 — FIRST. Common Vulnerability Scoring System v4.0: Specification Document (document version 1.2). first.org/cvss/v4-0/specification-document
  156. Ch. 24 — Chase M., Christey Coley S. Rubric for Applying CVSS to Medical Devices. MITRE (web page dated 21 Oct 2020). mitre.org/md-cvss-rubric
  157. Ch. 24 — US Food and Drug Administration. New Medical Device Development Tool (MDDT) Qualification for Cybersecurity. Bulletin, 20 Oct 2020. content.govdelivery.com/accounts/USFDA/bulletins/2a6cf79
  158. Ch. 24 — US Food and Drug Administration. Postmarket Management of Cybersecurity in Medical Devices. Final guidance. 28 Dec 2016. fda.gov/files/medical%20devices/published/Postmarket-Management-of-Cybersecurity-in-Medical-Devices---Guidance-for-Industry-and-Food-and-Drug-Administration-Staff.pdf
  159. Ch. 24 — ISO/IEC. ISO/IEC 29147:2018, Information technology — Security techniques — Vulnerability disclosure, ed. 2 (23 Oct 2018; to be revised since 26 Sep 2025), catalogue page. iso.org/standard/72311.html
  160. Ch. 24 — ISO/IEC. ISO/IEC 30111:2019, Information technology — Security techniques — Vulnerability handling processes, ed. 2 (1 Oct 2019; to be revised since 26 Sep 2025), catalogue page. iso.org/standard/69725.html
  161. Ch. 24 — European Commission. Guidance – MDCG endorsed documents and other guidance (web page, accessed 8 Oct 2026). health.ec.europa.eu/medical-devices-sector/new-regulations/guidance-mdcg-endorsed-documents-and-other-guidance_en
  162. Ch. 25 — US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (list; page updated 4 Sep 2026). fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
  163. Ch. 25 — IntuitionLabs. FDA-approved AI medical devices list: complete 2026 guide (19 Jul 2026), citing counts by TheImagingWire and Innolitics. intuitionlabs.ai/articles/fda-approved-ai-medical-devices-list (secondary)
  164. Ch. 25 — US Food and Drug Administration. Digital Health Advisory Committee meeting of 6 Nov 2025, Generative Artificial Intelligence-Enabled Digital Mental Health Medical Devices: agenda, discussion questions and brief summary. fda.gov/media/189389/download; fda.gov/media/189392/download; fda.gov/media/190450/download
  165. Ch. 25 — GE HealthCare. GE HealthCare drives growth with investment in AI-enabled medical devices and tops FDA’s list of AI authorizations for 4th year with 100. Press release, 23 Jul 2025. gehealthcare.com/en/about/newsroom/press-releases/ge-healthcare-drives-growth-with-investment-in-ai-enabled-medical-devices-and-tops-fdas-list-of-ai-authorizations-for-4th-year-with-100 (manufacturer)
  166. Ch. 25 — US Food and Drug Administration. 510(k) K200921, qER (Qure.ai, Mumbai): database record; received 6 Apr 2020, decision 17 Jun 2020. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K200921
  167. Ch. 25 — DAIC. Qure.ai’s chest X-ray reporting tool receives additional FDA clearances (26 Feb 2026). dicardiology.com/content/qureais-chest-x-ray-reporting-tool-receives-additional-fda-clearances (secondary, reporting the manufacturer)
  168. Ch. 25 — Qure.ai. Regulatory and privacy (web page, accessed 8 Oct 2026). qure.ai/regulatory-and-privacy (manufacturer)
  169. Ch. 25 — HLTH. Aidoc wins FDA nod for comprehensive foundation model AI (26 Jan 2026). hlth.com/insights/news/aidoc-wins-fda-nod-for-comprehensive-foundation-model-ai-2026-01-26 (secondary, reporting the manufacturer)
  170. Ch. 25 — Tijori Alerts. Lords Mark receives India’s first Biomescan SaMD manufacturing licence (10 Aug 2026). tijorialerts.com/company-updates/lords-mark-receives-indias-first-biomescan-samd-manufacturing-licence-139868/ (secondary, reporting the manufacturer)
  171. Ch. 25 — US Food and Drug Administration. November 20–21, 2024: Digital Health Advisory Committee Meeting Announcement (Total Product Lifecycle Considerations for Generative Artificial Intelligence-Enabled Medical Devices). fda.gov/advisory-committees/advisory-committee-calendar/november-20-21-2024-digital-health-advisory-committee-meeting-announcement-11202024
  172. Ch. 25 — McGuireWoods. A pathway for clinical AI developers opens: FDA clears first software as a medical device with patient-facing LLM. Legal alert, 6 Jul 2026. mcguirewoods.com/client-resources/alerts/2026/7/a-pathway-for-clinical-ai-developers-opens-fda-clears-first-software-as-a-medical-device-with-patient-facing-llm/ (secondary, reporting the manufacturer’s claim)
  173. Ch. 25 — Aguilar M. A ‘historic’ FDA clearance raises the question: is the LLM an interface or the decision-maker? STAT (2 Jul 2026). statnews.com/2026/07/02/fda-clearance-raises-questions-updoc-use-generative-ai-diabetes-treatment/ (secondary; read in part, paywalled)
  174. Ch. 26 — Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI). OJ L (24 Jul 2026); in force 27 Jul 2026. eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32026R1744 (read in two parts, with gaps)
  175. Ch. 26 — Taylor Wessing. Update: AI-enabled medical devices and IVDs confirmed as high risk (Life Sciences Legal Lens, vol. 2; 2026). taylorwessing.com/en/international-life-sciences-newsletter/life-sciences-legal-lens-vol-2/update-ai-enabled-medical-devices-and-ivds-confirmed-as-high-risk (secondary)
  176. Ch. 26 — Covington & Burling. 5 key takeaways from FDA’s revised clinical decision support (CDS) software guidance (8 Jan 2026). cov.com/en/news-and-insights/insights/2026/01/5-key-takeaways-from-fdas-revised-clinical-decision-support-cds-software-guidance (secondary)
  177. Ch. 26 — OpenRegulatory. MDCG 2019-11 explained (web page dated 12 Aug 2026). openregulatory.com/mdcg/mdcg-2019-11 (secondary)
  178. Ch. 26 — GMP Insiders. MDCG 2019-11 Rev.1: qualification and classification of software under the MDR and IVDR (web page). gmpinsiders.com/mdcg-2019-11-rev1-mdr-ivdr/ (secondary)
  179. Ch. 26 — Sobande S. European revision of primary software guidance MDCG 2019-11, revision 1. Emergo by UL (20 Jun 2025). emergobyul.com/news/european-revision-primary-software-guidance-mdcg-2019-11-revision-1-small-changes-meaningful (secondary)
  180. Ch. 26 — Rao A. CDSCO draft guidance on medical device software. India Briefing (12 Nov 2025). india-briefing.com/news/cdsco-draft-guidance-medical-software-40691.html (secondary)
  181. Ch. 26 — Nishith Desai Associates. From code to compliance: the CDSCO guidance on medical device software (3 Aug 2026) (secondary)
  182. Ch. 26 — US Food and Drug Administration. How to Determine if Your Product is a Medical Device (content current 29 Sep 2022). fda.gov/medical-devices/classify-your-medical-device/how-determine-if-your-product-medical-device
  183. Ch. 26 — US Food and Drug Administration. Classify Your Medical Device (content current 26 Aug 2026). fda.gov/medical-devices/overview-device-regulation/classify-your-medical-device
  184. Ch. 27 — US Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft guidance, docket FDA-2024-D-4488. Cover date 7 Jan 2025 (web page published 6 Jan 2025). fda.gov/media/184856/download; fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketing (PDF read to partway through Performance Validation)
  185. Ch. 27 — Medical Device Coordination Group. MDCG 2020-1, Guidance on Clinical Evaluation (MDR) / Performance Evaluation (IVDR) of Medical Device Software. Mar 2020. ec.europa.eu/docsroom/documents/40323
  186. Ch. 27 — Commission Implementing Decision (EU) 2021/1182 on harmonised standards for medical devices, consolidated text to 17 Jun 2026. eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02021D1182-20260617
  187. Ch. 27 — Critical Software. IEC 62304 Edition 2 changes (web article, undated). asd.criticalsoftware.com/en/newsroom/iec-62304-edition-2-changes-august-2026 (secondary)
  188. Ch. 28 — US Food and Drug Administration. Deciding When to Submit a 510(k) for a Software Change to an Existing Device. Final guidance, docket FDA-2016-D-2021. 25 Oct 2017. fda.gov/media/99785/download
  189. Ch. 28 — Medical Device Coordination Group. MDCG 2020-3, Guidance on significant changes regarding the transitional provision under Article 120 of the MDR. Mar 2020 (rev.1 of Sep 2023 not opened). ec.europa.eu/docsroom/documents/40301
  190. Ch. 28 — Central Drugs Standard Control Organisation, India. IVD Medical Devices FAQ, Doc No. CDSCO/IVD/FAQ/04/2022, with addendum of 28 Mar 2025. cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/FAaddendum.pdf (read by the G-06 researcher)
  191. Ch. 28 — Artixio. CDSCO medical device post approval change application (19 Aug 2026). artixio.com/post/cdsco-medical-device-post-approval-change-application (secondary)
  192. Ch. 28 — Central Drugs Standard Control Organisation, India. Addendum No. 02 to Doc No. CDSCO/IVD/FAQ/04/2022 (13 Mar 2026). cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/Addendum-Doc-No-CDSCOIVDFAQ-042022dated-13032026.pdf (read by the G-06 researcher)
  193. Ch. 28 — Regulation (EU) 2017/745 on medical devices, Annex X section 5 (changes to the approved type). OJ L 117 (5 May 2017). Text read as reproduced at advisera.com/13485academy/mdr/conformity-assessment-based-on-type-examination
  194. Ch. 29 — DLA Piper. FDA issues revised cybersecurity premarket submission guidance (27 Feb 2026). dlapiper.com/en/insights/publications/2026/02/fda-issues-revised-cybersecurity-premarket-submission-guidance (secondary)
  195. Ch. 29 — NSF. FDA updates cybersecurity guidance: shift toward QMSR alignment rather than new requirements (3 Feb 2026). nsf.org/life-science-regulatory-news/fda-updates-cybersecurity-guidance-shift-toward-qmsr-alignment-rather-than-new-requirements (secondary)
  196. Ch. 29 — AAMI. Food and Drug Administration adds AAMI cybersecurity guidance to Recognized Consensus Standards Database (AAMI CR515:2025, Cybersecurity considerations unique to machine learning-enabled medical devices). Press release via Newsfile, 24 Mar 2026. newsfilecorp.com/release/289622/Food-and-Drug-Administration-Adds-AAMI-Cybersecurity-Guidance-to-Recognized-Consensus-Standards-Database (standards developer’s press release)
  197. Ch. 29 — Regulation (EU) 2024/2847 of the European Parliament and of the Council of 23 October 2024 (Cyber Resilience Act). OJ L (20 Nov 2024); in force 10 Dec 2024. eur-lex.europa.eu/eli/reg/2024/2847/oj (recital 25 read on EUR-Lex; Articles 2 and 71 read via springlex.eu)
  198. Ch. 29 — Pure Global. India CDSCO medical device software guidance 2026 (9 Aug 2026). pureglobal.com/news/india-cdsco-medical-device-software-guidance-2026 (secondary)
  199. Ch. 29 — BioSpectrum India. CDSCO releases guidance document on medical device software (27 Jul 2026). biospectrumindia.com/news/93/28210/cdsco-releases-guidance-document-on-medical-device-software.html (secondary)
  200. Ch. 29 — LexCounsel. CDSCO publishes guidance on medical device software under MDR 2017. Mondaq (5 Aug 2026). mondaq.com/cdsco-publishes-guidance-on-medical-device-software-under-mdr-2017/1826286 (secondary)
  201. Ch. 29 — Press Information Bureau, Government of India. Digital Personal Data Protection Rules, 2025: explainer (17 Nov 2025). static.pib.gov.in/WriteReadData/specificdocs/documents/2025/nov/doc20251117695301.pdf
  202. Ch. 29 — S.S. Rana & Co. MeitY notifies final Digital Personal Data Protection Rules, 2025 (14 Nov 2025). ssrana.in/articles/meity-notifies-final-digital-personal-data-protection-rules-2025/ (secondary)
  203. Ch. 29 — Hogan Lovells. India’s Digital Personal Data Protection Act 2023 brought into force (Nov 2025). ca.hoganlovells.com/en/publications/indias-digital-personal-data-protection-act-2023-brought-into-force- (secondary)
§ PDF edition

Take the PDF with you.

Reading online needs no sign-up. For the PDF edition, laid out for print and offline reading, tell us where to send it and we will email you.

We use your details to send the guide, and to tell you about new guides only if you tick the box. See the privacy policy.

Created with Claude (Anthropic).
Educational reference material, not engineering, regulatory or clinical advice. Product and company names referenced are the property of their respective owners.