The next patient, the next update.
Software can never be tested on every path, and an AI model learns only from the data its makers chose. What evidence shows either will give the right answer for the next patient, and what keeps that true after an update?
Four hours of training.
The people holding the camera had signed a statement that they had never taken a picture of the inside of an eye. They worked in family-medicine and primary-care offices, ten of them spread across the United States, and each had been given one standardized training session of four hours, with no refresher afterward. Their task was to photograph the retinas of adults with diabetes who had come in for ordinary care: two pictures of each eye on a Topcon NW400, one centered on the optic disc and one on the fovea, the small pit at the center of sharp vision. About three participants in four needed no drops to widen the pupil. The other quarter were dilated and photographed again.
The pictures went to a program that had been locked before the first participant arrived. Within seconds it returned one of three answers: more than mild diabetic retinopathy detected, not detected, or image quality insufficient. Nobody at the clinic interpreted the photographs. Between January and July 2017, 900 people enrolled.
Every participant who completed the study was photographed a second time, by photographers certified by the Wisconsin Fundus Photograph Reading Center, with a far more demanding protocol: four overlapping stereo pairs of each retina and a scan of the macula by optical coherence tomography. Three experienced readers graded the photographs without seeing the program's answer, and their majority decided which participants truly had the disease; the macular scans were graded separately. Before the trial began, its designers, with input from the FDA, had set the targets it was measured against: sensitivity of 85 percent and specificity of 82.5 percent.
Of the 198 analyzable participants whose reference images showed more than mild disease, the program flagged 173, or 87.4 percent. Of the 621 without it, it cleared 556, or 89.5 percent. Among participants whose reference images could be graded, it gave a usable answer for 96.1 percent. The trial was funded by the developer, IDx, and its first author, the University of Iowa ophthalmologist Michael Abràmoff, had founded the company; the paper discloses the funding and his financial ties to IDx. On 11 April 2018 the FDA granted the system a De Novo authorization and described it as the first device authorized to provide a screening decision "without the need for a clinician to also interpret the image or results."
The trial measured one locked program, run by those operators on that camera, for those patients. It said nothing about the same program in a clinic whose computers send images differently, after the camera is replaced or the model retrained, or when someone sets out to make it fail. A device that will be used for years on people who were never in its trial needs evidence of a different kind, and each later version needs more of it.
Before we start.
Software in a medical device does what it was written, or trained, to do, on whatever input arrives. A trial like the one above shows how it behaved for 819 analyzable people. A maker, a regulator and a hospital need something more: a reason to believe the next person will get the right answer too, including people the trial never saw, and that the reason will survive the next release. That is the question this guide answers: software can never be tested on every path, and an AI model learns only from the data its makers chose; what evidence shows either will give the right answer for the next patient, and what keeps that true after an update?
Four terms carry the argument. A device software function is software that meets the legal definition of a medical device, whether it runs inside a pump or on a server by itself. An algorithm is the procedure the software follows. An AI model, in this guide, is an algorithm whose decision rules were fitted to example data rather than written line by line. And evidence means records that someone outside the team could check: requirements, test results, study data, monitoring reports.
The answer comes in four parts. Testing samples what software does; it cannot exhaust it, because the paths through even a modest program and the inputs it may meet outnumber any test campaign. Confidence therefore comes from how the software was built: requirements derived from the harm a fault could do and traced to tests, an architecture that keeps the dangerous part small enough to test thoroughly, a full account of the code someone else wrote, and proof that each change broke nothing already verified. Agile teams and automated pipelines can produce that evidence, provided each increment leaves its records. An AI model adds a second limit. Its accuracy is a measurement on a test set, and it holds only for patients like those in that set, scored against a reference someone chose, for the exact model that was locked and fed the inputs it was specified for. The evidence that it will be right for the next patient is a test kept apart from training and drawn from the setting of use, results reported for each group of patients with their uncertainty, a check at each hospital that the specified conditions hold, and monitoring once it is in use. A device can also be made to fail by someone trying, and that threat changes while the code stands still, so a device must be built to be patched and watched for years, starting from a list of what is inside it. What keeps all of this true after an update is change control: every change judged against the claims it could move, and, for AI, a plan agreed with the regulator in advance that says what may change and how the new model must be tested.
One fact shapes every chapter. Software does not wear out. A pump's motor and a camera's sensor age and fail at rates that can be measured; software never does. Every way it can fail was built in on the day it shipped or arrived with an update, waiting for the input that triggers it, and for an AI model the training data is part of what was built.
How to read this guide
The guide runs through 29 chapters in eight parts. Parts I and II explain why software fails differently from hardware and what a software lifecycle does about it, including when much of the code comes from elsewhere and when the team ships every two weeks. Parts III and IV turn to models learned from data: how they are built and scored, and why they do worse on patients unlike those they were tested on. Part V follows a model into a hospital and through its later versions; by its end the central question has its answer. Part VI covers attack and defense, Part VII surveys the products and makers of October 2026, and Part VIII sets out what the FDA, the EU and India require, as of that month. IDx-DR returns throughout as the thread's device, and the recalled insulin pumps return in the security chapters. Every chapter works one example with real numbers, and plates carry their own numbers so that each can be cited alone. These boxes recur:
Blue edge. Carries the structural point of a section, or a calculation worked once with real values and stated assumptions.
Orange edge. Names a common misreading, a trap, or the limit of a claim.
Green edge. Maps the idea onto the reader's own work: specifying, building or testing device software, evaluating an AI product, deploying it in a hospital, or planning the evidence for a submission.
- Each chapter closes with what it established, in four lines.
No formula needs more than arithmetic. Proportions such as sensitivity are percentages; uncertainty is given as a range, and the one rule used to size a test, that a run of failure-free trials bounds a failure rate, is worked once with numbers. Counting paths uses powers, written as 10⁹ for a billion. Where a figure is a company's own claim about its product, the text says so, and where a paper's authors work for the company whose product they test, the chapter notes it. The guide uses US spelling.
Colors in the plates follow a fixed key. Requirements, claims and acceptance criteria are gold; software the maker wrote, and a model running as software, is teal, with third-party code hatched in the same teal; data of every kind, from training sets to the images a device receives, is violet; an attacker, an attack path or a vulnerability is red; and a defect, a wrong output or a drift in performance is magenta. The output reported to a clinician or patient is always blue, and one orange mark in each plate points at the detail that matters most. A key strip under every plate lists only the colors that plate uses.
This is a reference for understanding, not a development procedure, a regulatory opinion or clinical advice. Products, authorizations, guidances, standards and laws are stated as of October 2026 and will date; several of those named here were drafts or under revision in that month, and the text says which. Where a figure comes from a company about its own product, or from a paper whose authors work for that company, the text says so. Named companies and products are examples of a principle, not recommendations.
+ Part I · Built in on the day it shipped
Why software fails differently.
Between 1992 and 1998 the FDA counted 3,140 recalls of medical devices, and 242 of them, 7.7 percent, were caused by software. Of those 242, 192 were caused by defects introduced when the software was changed after it had first been distributed. Both numbers point to the same peculiarity: software is not a part that ages, so its failures are mistakes made by people, in the first version or in a later one. The three chapters of this part set out what software in a device is trusted to do, why it fails when it does, and why no amount of testing can prove that it will not.
What the code is trusted to decide.
+ The questionWhat does the software in a medical device actually do, and when is the software itself the device?
A camera, a client and a server
The system in the opening scene has three pieces, and only its software is the device the FDA authorized. A commercial retinal camera, the Topcon NW400, takes the photographs. A computer attached to the camera runs a program the maker calls the client, which lets the operator pick two images of each eye and sends them "over a secure internet connection" to a server in a data center. There a second program, the analysis, examines the images and returns one of three answers to the client screen: more than mild diabetic retinopathy detected, not detected, or image quality insufficient. The FDA's decision summary describes this architecture and names the software version reviewed in 2018, IDx-DR 2.0.0.
The camera is a separate product, outside the authorization. What the FDA classified in April 2018 was software, in a new category it called a retinal diagnostic software device, class II, under a regulation written for this authorization. Nothing in that device touches the patient, emits energy or moves. Its effect on health passes entirely through a decision: whether a person with diabetes is referred to an eye specialist now or told to come back in a year.
When the software is the device
Most medical devices with software carry it inside: the firmware of an infusion pump, the program that runs a ventilator's valves, the code in an imaging scanner. That software is part of the device and is judged with it. A growing share of devices are software alone, running on ordinary computers, phones or cloud servers. Regulators call these software as a medical device, or in newer documents simply medical device software that stands alone, and IDx-DR's analysis program is an example.
What turns code into a device is not anything in the code. It is the intended use: what the maker says the software is for, in its labeling and claims. A program that draws a graph of blood glucose readings for a person's own interest is not a device in the US; the same graph with a rule that tells the patient how much insulin to take is. The US definition, the EU's rules and India's 2026 guidance draw the boundary differently in detail, and Chapter 26 sets those differences out. They agree on the principle: the claim makes the device, and the claim also fixes what the evidence must show.
Two companies can ship identical image-analysis code, one to help researchers count lesions in a study and one to tell clinicians which patients to refer. Only the second makes a medical claim, and only the second needs a medical device's evidence. Changing a sentence of labeling can therefore change the regulatory status of software that has not changed at all.
Inform, drive, diagnose or treat
Once software is a device, the evidence it needs scales with two questions. The first is how serious the patient's situation is: whether a wrong answer could lead to death or irreversible harm, to a serious but treatable deterioration, or to something minor. The second is how much the software's output decides: whether it treats or diagnoses directly, drives the next step of care, or only informs a clinician who still decides. International regulators set out these two axes in a framework for software as a medical device in 2014, and India's 2026 guidance uses the same grid to assign its classes.
IDx-DR sits high on the second axis. Its output is the screening decision itself; no clinician interprets the image. Diabetic retinopathy threatens sight but rarely kills, and a person told to rescreen in a year will usually be seen again, so the situation is serious rather than critical. A stroke-triage program that alerts a specialist to a suspected blocked artery, by contrast, informs a clinician who still reads the scan, but in a situation where minutes count. The two products sit in different cells of the same grid, and Chapter 16 returns to the difference between a program that decides and one that advises.
Take four functions. A pump's bolus calculator computes an insulin dose from a glucose reading and the carbohydrates entered; it sets the treatment itself, and an error of a few units can cause dangerous hypoglycemia, so it sits at the top of both axes. A triage program that flags suspected large-vessel occlusion on a CT angiogram informs a stroke team in a critical situation: lower on the second axis, highest on the first. A logbook app that stores and displays glucose readings without interpreting them generally falls outside the US device definition, because displaying device data is excluded by statute. A step counter for fitness is a wellness product, not a device. The same phone could run all four; the evidence each needs ranges from none to a clinical trial.
The claim is the unit of evidence
Because the claim defines the device, it also defines what can go wrong. The FDA's decision summary for IDx-DR lists three risks to health: a false positive leading to an unnecessary referral, a false negative delaying evaluation and treatment, and an operator failing to capture images good enough to analyze. Every control it required answers one of those three: clinical testing, software verification and validation built on a hazard analysis, training and human-factors testing for operators, labeling, and a protocol stating which changes could significantly affect its safety or effectiveness. Chapter 21 shows how that last requirement anticipated, by six years, the plans for changing AI models that the FDA finalized in 2024.
The list is short because the claim is narrow. IDx-DR is indicated for adults with diabetes who have not been diagnosed with retinopathy, used by health care providers, with images from the Topcon NW400. It does not look for glaucoma, and at the time of authorization it was not labeled for patients younger than 22 or for pregnant patients. Each limit removes a population or a condition from what the evidence has to cover. The rest of this guide is about the evidence inside those limits, and about what happens at their edges, where real use tends to wander.
A useful intended-use statement names the decision the software supports, the patients it applies to, the users, the setting and the inputs it accepts, including the devices that produce them. Draft it before architecture begins, and treat each change to it as a design change: it fixes the class of the device, the hazards to analyze, the population a clinical study must sample, and the edges where use outside the claim is most likely.
- IDx-DR is a software device: a client and a server-side analysis that turn photographs from a named camera into one of three screening answers.
- Software becomes a medical device through its intended use, not through anything in its code, so a change of claim can change its status.
- The evidence a software device needs rises with how serious the patient's situation is and with how much the software's output decides.
- The claim defines the risks and the controls; narrow indications shrink what the evidence must cover, and real use tends to test the edges.
Built in on the day it shipped.
+ The questionIf code never wears out, where do its failures come from?
Four conditions at once
In May 2024 Hamilton Medical began correcting the software of its HAMILTON-C6 ventilators, and the FDA classed the action as the most serious type of recall. The fault needed four things to happen together. A clinician pressed the key that briefly raises the oxygen supplied and disconnected the breathing tube to suction the patient's airway. During that disconnection a sensor error occurred, for example because the tubing of the flow sensor was kinked. The ventilator entered its sensor-fail mode. And the patient was reconnected while that mode was still active. In that sequence, and only in it, ventilation might not restart. The FDA's notice lists one injury and one death.
Each of the four events is ordinary on an intensive-care unit. The combination is rare, and it lay inside three versions of the software, 1.1.4 to 1.1.6, in every ventilator that ran them, until version 1.2.3 removed it. No part had worn and no component had drifted. The ventilators that could fail this way were exactly as they had been on the day the software was installed.
A fault waits for its input
The FDA's guidance on software validation, issued in 2002 and still in force apart from one section replaced in 2025, puts the difference plainly: "Unlike hardware, software is not a physical entity and does not wear out." A hardware part fails for physical reasons that accumulate: a bearing wears, a capacitor dries, a solder joint cracks under thousands of heating cycles. Its failures are random in the engineering sense: each unit fails at its own moment, and a rate per hour of operation describes the population well.
Software fails only when an input reaches a fault, a flaw in its logic or data, that was there from the start. Engineers call these failures systematic: they occur every time the same conditions recur, in every copy of the same version. The time to failure depends on when the triggering input arrives, which depends on how the device is used, not on how long it has been running. A rarely used function can carry a fault for years before anyone meets it, and a busy hospital can meet it on the first day.
A version that has run in the field for years without a reported failure has shown that the inputs it met so far did not reach a fault. A new hospital, a new workflow or a change in another device can supply the input that the earlier users never did. Field history is useful evidence about the common paths and weak evidence about the rare ones.
Why reliability arithmetic does not transfer
The distinction changes how safety is argued. For hardware, two independent parts that each fail once in many thousand hours rarely fail in the same hour, so redundancy multiplies safety. For software, two copies of the same program fed the same input fail together, so copying the code adds nothing against its own faults.
Suppose a pump motor fails at random at a rate of 1 in 100,000 per hour, and a second, independent motor is fitted as a backup. The chance that both fail in the same hour is 1 in 100,000 multiplied by 1 in 100,000: 1 in 10,000,000,000. Now suppose the pump's control program has a fault that a particular input triggers once in 100,000 hours of use, and a second processor runs an identical copy as a backup. The input that reaches the first copy reaches the second, so the chance that both fail in the same hour is still 1 in 100,000. Independence, which made the hardware redundancy work, is absent; protection against a software fault has to come from a design that is different or from a check that does not depend on the same code, which is the subject of Chapter 5.
What do such faults look like? A NIST analysis of the 383 software-related device recalls the FDA recorded from 1983 to 1997 could assign a type to 342 of them. Logic faults, a wrong condition or an unhandled case, made up 43 percent; calculation faults made up 24 percent; the rest were spread across data, requirements, timing, interfaces and other types. None is a matter of wear. Each is a decision written into the program that was wrong for some input. The same analysis found that the recalled devices had caused no reported deaths or serious injuries, so the figures describe how software fails, not how often it harms.
A change is a new chance to fail
If faults are built in, the moment of building matters, and building does not stop at the first release. In the FDA's count from the 1990s, 192 of the 242 software-related recalls, 79 percent, came from defects introduced by changes made after the software was first distributed. A fix or a new feature reaches parts of the program that were already verified, and anything it disturbs there becomes a new fault.
In March 2024 Tandem Diabetes Care recalled version 2.7 of its t:connect app for iPhone, which works with the company's t:slim X2 insulin pump. The FDA's notice describes the mechanism: the app could crash and be relaunched by the phone's operating system again and again, and the repeated relaunches caused excessive Bluetooth communication with the pump, which drained its battery until it shut down early and insulin delivery stopped. The recall covered 85,863 installed apps, and by mid-April 2024 the company had received 224 reports of injury. The FDA recorded the root cause as a software design change. The pump had not changed at all; a new version of the program that talked to it had.
The two halves of the central question start here. A program that was right for one set of inputs can be wrong for another that arrives later, and a program that was right can become wrong when it is changed. Chapter 3 asks why testing alone cannot close either gap, and Part II describes what the software lifecycle does instead.
When a software failure is reported, assume the fault is present in every unit running that version and in every version that shares the code. Reconstruct the exact sequence of inputs, write it as a test that fails on the faulty version, and keep that test in the regression suite permanently, so that no later change can bring the fault back unnoticed.
- Software faults are present from the moment the code is written; a failure occurs when an input reaches one, in every copy of that version.
- Software failures are systematic, so the time to failure depends on use, and years without failure say little about rare paths.
- Identical copies of a program fail together, so redundancy that works for hardware does not protect against a software fault.
- Changes are a major source of new faults: 79 percent of the software recalls the FDA counted in 1992 to 1998 followed changes after release.
More paths than seconds.
+ The questionWhy can't a team simply test every case before release?
Thirty decisions, a billion paths
A routine with 30 independent yes-or-no decisions has 1,073,741,824 distinct paths through it, because each decision doubles the number of paths that came before. A device program has thousands of decisions, many of them inside loops that can run any number of times, and each path can be taken with different data values and at different moments relative to other events. The count of distinct behaviors is not large in the way a big inventory is large. It is beyond counting.
Suppose a program had only 100 independent decisions and a test rig could run a billion paths a second. Running every path once would take about 40 trillion years, nearly 3,000 times the age of the universe. Faster computers do not change the conclusion, because each extra decision doubles the work. The FDA's software validation guidance states the consequence without arithmetic: "Except for the simplest of programs, software cannot be exhaustively tested," and "path coverage is generally not achievable."
Presence, not absence
The computer scientist Edsger Dijkstra drew the conclusion that shaped software engineering. In his 1972 Turing Award lecture he said that program testing "can be a very effective way to show the presence of bugs, but is hopelessly inadequate for showing their absence." A test that passes shows that one path, with one set of values, gave the expected answer. It says nothing directly about the paths not taken.
That does not make testing useless. It makes testing a sample, and like any sample its value depends on how it was drawn. A test campaign chosen at random from the space of possible inputs would almost never reach the rare combinations that matter, such as the four simultaneous conditions of the HAMILTON-C6 recall in Chapter 2. A campaign chosen from what the software is required to do, and from the ways it could cause harm, reaches far more of them. The FDA's guidance says testing alone "cannot fully verify that software is complete and correct," which is why Part II describes the rest of the evidence a lifecycle produces.
A coverage tool reports which statements or branches of the code a test campaign executed. A campaign can execute every statement and still miss a wrong condition, because executing a line with one value does not test it with others. The FDA's guidance calls statement coverage "insufficient to provide confidence" on its own, and even full branch coverage leaves most paths unexercised.
How many failure-free runs prove a rate?
Statistics gives a precise answer to a narrower question: if a program runs many times without failing, how high could its failure rate still be? The rule of three answers it. When no failures are seen in a number of independent trials, the one-sided 95 percent upper confidence bound for the failure probability is about three divided by that number.
A team runs a function 300 times under realistic conditions and sees no failure. The rule of three says the true failure probability could still be as high as 3 in 300, or 1 percent, at 95 percent confidence. To bring the bound down to 1 in 1,000, the team needs about 3,000 failure-free runs; to 1 in 1,000,000, about 3,000,000. For a rate of 1 failure in 1,000,000,000 hours, a target used for life-critical aviation software, it needs about 3,000,000,000 failure-free hours of realistic operation, which is about 342,000 years. NASA researchers Ricky Butler and George Finelli made the same point in 1993: their table puts the test time needed to show a failure probability of 1 in 1,000,000,000 over a ten-hour mission, on one system, at about 1.1 million years, and even 10,000 copies tested in parallel would need 114 years.
The bound applies only when the test runs resemble real use, so the rare inputs of the field appear in the test at the rate they appear in practice. That is the second weakness of statistical testing for safety: the inputs that cause harm are, by their nature, the ones nobody thought to make common in the test.
Choosing the inputs that matter
If testing cannot be complete, the question becomes which inputs to try. Three ideas do most of the work. Inputs can be grouped into classes the software should treat alike, with tests at the boundaries between classes, where off-by-one mistakes live. Combinations of settings can be covered systematically rather than exhaustively. And the hazard analysis of Chapter 4 can name the specific sequences, like the ventilator's four conditions, that would cause the most harm, so that each becomes a deliberate test.
The second idea has unusually direct evidence from medical devices. When NIST researchers Dolores Wallace and Richard Kuhn studied FDA recall reports from 1983 to 1997, 109 described their failures in enough detail to see what triggered them. In all but 3 of those 109 (97 percent; the authors give 98), testing every pair of parameter settings would have revealed the failure; the other 3 needed more than two conditions to occur together. Testing all pairs of settings grows slowly with the number of settings, so it is affordable where testing all combinations is not. The HAMILTON-C6 case shows the limit: its failure needed four conditions, so a pairwise campaign would not have been sure to find it, and only an analysis of what happens when a sensor fails during suctioning would have pointed to that sequence.
Testing, then, is necessary and never sufficient. The confidence a regulator accepts comes from testing chosen by requirements and risk, applied at several levels, repeated after every change, and backed by a process that makes faults less likely to be written in the first place. That process is the subject of the next part.
For each hazard the software could contribute to, list the inputs and sequences that could lead to it, including combinations of user actions, sensor states and timing. Make those deliberate tests; cover the remaining settings with boundary values and all-pairs combinations; and record which hazards rely on testing alone and which also rely on design controls, so that a reviewer can see where the testing argument is thin.
- Paths through software multiply with every decision, so even modest programs have far more behaviors than any campaign can test.
- Testing can show that faults exist but cannot show their absence; coverage measures what ran, not whether it was right.
- By the rule of three, proving a very low failure rate by failure-free testing would take an impossible amount of realistic operation.
- Useful testing is a deliberate sample: input classes and boundaries, all pairs of settings, and sequences named by the hazard analysis.
+ Part II · Confidence from process
The software lifecycle.
IEC 62304, the international standard for the life cycle of medical device software, was published on 9 May 2006 and amended in June 2015. It prescribes activities and the records each leaves behind, not an order in which to do them, and the FDA has recognized the amended edition in full since January 2019. A second edition, widened to health software in general, reached its second committee draft in August 2026 and is not expected to be final before 2028. Because testing cannot be complete, the confidence a regulator accepts comes largely from how the software was made. The six chapters of this part follow that process: requirements derived from risk, an architecture that isolates what is dangerous, an account of the code taken from others, testing at several levels, and the ways agile teams and automated pipelines meet the same obligations.
What it must do, and must never do.
+ The questionWhat has to be written down before a line of code is written, and why does risk decide how much?
A requirement is a promise that can be checked
A software requirement is a statement of what the software must do, or must never do, written so that a test, an inspection or an analysis can show whether it holds. "The display shall be easy to read" is not a requirement in this sense; "each dose value shall be shown with its unit, in characters at least 5 mm high, and shall remain on screen until the user confirms it" is. The difference is that the second can fail a test.
Requirements come from the intended use of Chapter 1. A claim to detect more than mild diabetic retinopathy from two photographs per eye implies requirements about which images are accepted, what happens when an image is too dark or blurred, what the three possible answers are and how each is shown, how results are tied to the right patient, and how long the analysis may take. Some come from users and clinicians, some from standards, some from the hazards the software could cause, and some from the rest of the system it lives in. Written down and reviewed, they become the yardstick for everything that follows: design is checked against them, tests are written from them, and a change is judged by which of them it touches.
Risk decides how much
Not every requirement carries the same weight. Risk management for medical devices follows ISO 14971, whose 2019 edition was confirmed in 2025. It traces a chain: a hazard, a potential source of harm; a sequence of events that leads to a hazardous situation, in which someone is exposed to it; and the harm that may follow. Each chain is judged by the severity of the harm and the probability that the harm occurs, and each unacceptable risk gets a control. The immunoassay guide in this library follows the chain for an analyzer, from a fault to a wrong result to a harm; software adds a twist to it.
The twist is probability. Chapter 2 showed that software failures are systematic, so there is no failure rate per hour to put into the chain. International regulators, in a 2025 document on software-specific risk, suggest that it can help to start from a probability of software failure of one and, where possible, to estimate the probability of harm from other factors, judging the risk by severity and by the controls that act if the failure happens. A risk control may itself be software, a check in another part of the program; or it may be hardware, procedure or labeling outside the software altogether.
Three classes of software
IEC 62304 turns this into three software safety classes. In the amended edition, software is class A if it cannot contribute to a hazardous situation, or if any it contributes to does not lead to unacceptable risk once risk controls outside the software are taken into account. It is class B if, after those external controls, it could still contribute to a hazardous situation whose possible harm is a non-serious injury, and class C if that harm could be death or serious injury. Until a team has classified its software, the standard treats it as class C.
The class sets how much process the standard demands. All three classes need a development plan, requirements, testing of the whole software against those requirements, risk management, configuration management, problem resolution and controlled release. In outline, class B adds a documented architecture and verification of software units and their integration; class C adds a detailed design of each unit and more demanding criteria for verifying them. The class also applies to parts of a program, so Chapter 5 shows how an architecture that separates dangerous functions can put most of the code in a lower class.
An infusion pump's software computes a delivery rate from a prescription. A wrong rate could kill, so before any controls the worst credible harm is death. Suppose the pump also has an independent hardware circuit, outside the software system, that stops the motor if the delivered volume exceeds a limit the clinician has set, and the team's analysis shows that with this circuit in place the worst remaining harm from a software error is a non-serious injury. Under IEC 62304 as amended, the software may then be class B, because the external control counts. The FDA's 2023 guidance on premarket software documentation instead asks whether a failure "could present a hazardous situation with a probable risk of death or serious injury" before any risk controls are considered. The same pump software would need the FDA's Enhanced documentation level. The guidance itself explains why a declaration of conformity to the whole of IEC 62304 is not needed: the two schemes classify differently.
A check that runs inside the same program, on the same processor and from the same code base, can fail together with the function it guards, as the redundant copies of Chapter 2 did. IEC 62304 credits risk controls external to the software system when it sets the class. A team that lowers a class by pointing to a software check inside the system has misread the standard and has also left the risk in place.
Tracing the threads
The records that make this auditable are links. Each hazard links to the requirements that control it; each requirement links to the design elements that implement it and to the tests that verify it; each test links to its result. With the links in place, a reviewer can start anywhere: from a hazard, find whether its control was implemented and tested; from a failed test, find which hazard is now uncontrolled; from a change, find every requirement and test it touches. Without them, a stack of documents is only a stack.
The FDA's decision summary for IDx-DR shows the top of such a chain in public. It lists three risks to health, false positives, false negatives and operators failing to capture usable images, and maps each to its mitigations: clinical performance testing, software verification and validation built on a hazard analysis, a change protocol, labeling, operator training and human-factors testing. Below that summary, invisible to the public, sit the requirements and tests the reviewers read. Chapter 7 describes the testing; Chapter 27 lists what the FDA, the EU and India ask to see of it.
Classify the software, and each software item, before architecture is fixed, and write down the hazards, the external controls credited and the worst remaining harm behind each class. Record the FDA documentation level separately, assessed before risk controls. Revisit both whenever a requirement, a control or the intended use changes, because a class that rested on a hardware limit fails silently if that limit is removed.
- A software requirement states what the software must or must never do in a form that a test or inspection can check.
- ISO 14971 traces hazards to harms; because software failures are systematic, regulators suggest assuming a software failure will occur.
- IEC 62304 classes software A, B or C by the worst harm after controls outside the software, and the class sets the process required.
- Traceability links hazards, requirements, design and tests, so a reviewer can follow any risk to its control and its evidence.
Keeping the dangerous part small.
+ The questionHow does the way software is divided into parts change what must be proven about each?
A timer that does not trust the program
A watchdog is a timer, often a separate circuit, that the software must reset at regular intervals. If the program hangs, stuck in a loop or waiting for something that never arrives, it stops resetting the timer, the timer runs out, and the watchdog forces the device into a safe state: it stops a motor, raises an alarm or restarts the processor. The FDA's guidance on infusion pumps, finalized in 2014, asks makers to describe their watchdog timer, lists watchdog tests among the safety mechanisms a pump should have, and lists a watchdog failure among the hardware causes of system failure.
The watchdog embodies the main idea of software architecture for safety. A function that must not fail is protected by something that does not share its weaknesses. The watchdog does not run the dose calculation and does not need to understand it; it only needs proof of life at regular intervals, and its own logic is small enough to verify thoroughly. The same guidance adds a caution that applies to every safety mechanism: the safety case must also cover the hazards that the mechanism itself could start, such as, to take an example of our own, a watchdog that resets a pump in the middle of an infusion.
A program that runs normally while computing a wrong dose keeps resetting its watchdog on time, so the watchdog never acts. Protection against a wrong output needs a different mechanism: an independent check of the result, a limit enforced outside the software, or a second calculation by different means. Naming which failure each safety mechanism catches, and which it cannot see, is part of the design.
Items, units and boundaries
IEC 62304 describes a program as a software system divided into software items, any identifiable parts, which are divided in turn until they reach software units, the items not divided any further. The division is not a filing scheme. It decides where faults can travel. If one item can overwrite another's memory, starve it of processor time or hand it corrupted data without detection, a fault in the first is a fault in the second, and both must be treated as one.
The standard calls the remedy segregation: any mechanism that prevents one software item from negatively affecting another. Segregation can be physical, with the critical function on its own processor; it can use the memory protection of an operating system, so that one process cannot write into another; or it can rest on checks at the boundary, so that data crossing it is validated. Each mechanism has to be shown to work, because a claimed boundary that a fault can cross is worse than none: it lowers the rigor applied to the code on the far side without lowering the risk.
Segregation is what lets the class of Chapter 4 apply to parts of a program rather than the whole. If the items that could contribute to serious harm are separated from the rest, only they need the class C activities, and the user interface, the logging and the reporting code can be developed at class B or A. Architecture is therefore a decision about evidence as much as about structure.
Suppose a pump's software has 200,000 lines, of which the dose calculation and the delivery control, the only parts whose failure could cause death or serious injury, take 8,000. Without demonstrated segregation, the whole program is class C, and the detailed design and stricter unit-verification criteria of that class apply to all 200,000 lines. If the 8,000 lines run on a separate processor that the rest cannot write to, and every command crossing the boundary is checked against the prescription, the class C work applies to 8,000 lines, 4 percent of the program. The remaining 96 percent still needs its own class, its own tests and evidence that the boundary holds.
Where the parts live
The architecture also has to say where each part runs and what it depends on. Much of the code in a modern device was written by someone else: the operating system, the network stack, a database, a library that reads images, a cloud service. Each is an item the maker did not write and cannot fully inspect, and the architecture has to place it, bound it and say what the device does when it misbehaves. Chapter 6 is about those items.
The thread's device shows the split in its public record. IDx-DR has three software items with their own version numbers: the client on the camera's computer, a service on the server that passes images and results between client and analysis, and the analysis that grades the images. The 510(k) cleared in June 2021 moved the client from version 2.0.1, as its summary lists the authorized software, to 3.2.0, the analysis from 2.0.1 to 2.1.1 and the service from 1.0.0 to 1.1.2; the next, cleared in June 2022, moved them to 3.5.0, 2.3.0 and 1.2.0. Because the parts are separate, a maker can change one and argue, with evidence, that the others are untouched. The split also shapes what can go wrong: an analysis that runs on a server depends on the network, and the 2021 version added messages that tell an operator whether a submission was lost or is simply not ready yet.
Chapter 21 returns to the 2022 change, which replaced one model inside the analysis, the classifier that judges image quality, and changed which patients received an answer. An architecture lets a change be confined to one item. It does not guarantee that the change's effects stay there, which is why every change needs regression tests at the level of the whole system.
Draw every software item, including those from third parties, with its safety class, where it runs, and the mechanism that separates it from its neighbors. For each boundary that lowers a class, record the evidence that the mechanism works, such as memory-protection settings or the separate processor, and include a test that tries to break it. For each safety mechanism, record which failures it detects, which it cannot, and what hazards it could create itself.
- A watchdog shows the core idea: protect a critical function with a mechanism that does not share its weaknesses.
- IEC 62304 divides software into items and units; segregation is any mechanism that stops one item from harming another.
- Demonstrated segregation lets the strictest class apply only to the code that could cause serious harm, not to the whole program.
- IDx-DR's three separately versioned items show how architecture confines a change, though system-level tests must confirm its effects.
SOUP, OTS and the bill of materials.
+ The questionMost of the code in a modern device was written by someone else. How can a maker answer for it?
When the operating system expires
On 8 April 2014 Microsoft ended extended support for Windows XP, twelve years and three months after the system's lifecycle began. After that date Microsoft stopped routine security fixes for it. Medical devices built on Windows inherited that calendar, and many outlived it: an imaging scanner, a laboratory analyzer or a radiology workstation is expensive, is expected to last a decade or more, and was validated with one exact version of its operating system. Replacing that system is a design change that needs new testing and, for some devices, a new review.
The cost of waiting became visible on 12 May 2017, when the WannaCry ransomware spread across networks running unpatched or unsupported versions of Windows. The UK's National Audit Office found that at least 81 of the 236 trusts of the National Health Service in England, 34 percent, were affected, and identified 6,912 cancelled appointments, with more than 19,000 estimated. Bayer confirmed two reports from US customers of infected devices of its own, and Siemens Healthineers warned that some of its products might be affected. In 2020 a security vendor's survey of its customers' networks reported that 83 percent of the medical imaging devices it saw ran operating systems that were no longer supported, a share that had risen sharply since 2018 and that the vendor attributed to the end of support for Windows 7.
Suppose a device is launched in 2009 on Windows XP, sold until 2012, and used for ten years from the date of sale. Extended support ended in April 2014, five years after launch. A unit sold in 2012 stays in service until 2022, eight years after its operating system stopped receiving security fixes. Any vulnerability published in those eight years stays open on that unit unless the maker can fix it some other way. The FDA's 2026 cybersecurity guidance addresses this by asking makers to state, for each software component, its level of support and its end-of-support date.
SOUP and OTS
Two terms name the code a maker did not write, and they overlap without being the same. IEC 62304 calls it SOUP, software of unknown provenance: a software item that is already developed and generally available and was not developed to be part of the device, or a software item developed earlier for which adequate records of its development are not available. The FDA calls it OTS, off-the-shelf software: a generally available software component used by a device maker that cannot claim complete control of its life cycle.
The emphasis differs. SOUP is about provenance: a maker's own old code, written before anyone kept records, is SOUP even though nobody bought it. OTS is about control: a commercial operating system is OTS because its maker, not the device maker, decides what goes into it and when. Most third-party components are both. A compiler or a test tool that never ships in the device is not SOUP, because it is not part of the device; it is validated for its use under the quality-system rules of Chapter 9.
IEC 62304 asks the same things of every SOUP item. The maker states what functional and performance requirements the device places on it, what hardware and software it needs, and its exact identity and version. It evaluates the anomalies the component's own maker has published, to see whether any could cause a hazard, and controls its configuration so that the version tested is the version shipped. The FDA's guidance on off-the-shelf software, last revised in August 2023, asks for much the same and adds a plan for maintaining support. At its Enhanced documentation level, it also asks for assurance about how the component's developer works and what the device maker will do if that developer changes or stops its support.
How much of the code is borrowed
The borrowed share is usually most of the program. A vendor that audits software for company acquisitions reported in 2026 that 98 percent of the 947 codebases it examined contained open-source code, with a mean of 1,180 open-source components per application, and that 87 percent contained at least one component with a known vulnerability. The sample is software being sold in a transaction, not medical devices, but device software is built from the same libraries. An operating system, a network stack, a database, an image library, a cryptography package and a user-interface toolkit can each pull in dozens of further components of their own.
Models and data are now part of the list. A device can include a model pretrained by another company on images the device maker has never seen, or a public dataset used for tuning. Neither is code in the usual sense, but each is a component whose origin, version and known limits matter as much as a library's. The CycloneDX format for bills of materials added a machine-learning section in 2023 so that models and datasets can be listed with the code.
The bill of materials
A software bill of materials, or SBOM, is the machine-readable list of every component in a piece of software, with enough detail to identify each exactly. The US National Telecommunications and Information Administration set out seven minimum fields in July 2021: the supplier, the component's name, its version, other unique identifiers, its dependency relationships, the author of the SBOM data and a timestamp. The US cybersecurity agency CISA published updated minimum elements in mid-2026, adding a cryptographic hash of each component, its license, the tool that generated the SBOM and the context in which it was generated, and extending the scope explicitly to AI software and software delivered as a service. Two formats dominate: SPDX, published as the international standard ISO/IEC 5962 in 2021, and CycloneDX, an Ecma standard since 2024.
A component with a published vulnerability may be present but unreachable: the device may never call the affected function or never expose it to a network. A component with no published vulnerability may still be flawed. The SBOM answers which components are in the device; whether a given vulnerability can be exploited in that device is a separate statement, which the industry calls a VEX: not affected, affected, fixed, or under investigation.
For devices the SBOM is no longer optional in the US. Since 29 March 2023 the law has required every premarket submission for a "cyber device", one with software that can connect to the internet, a condition the FDA reads broadly, to include one, including commercial, open-source and off-the-shelf components, and the FDA's guidance asks for each component's support level and end-of-support date alongside it. Chapter 24 follows the SBOM after release, where it does its real work: when a new vulnerability is published in a component, the list tells a maker, and a hospital, which devices contain it, far faster than a search of old build records.
For each SOUP or OTS item, record its exact version, its supplier, the requirements the device places on it, the published anomalies reviewed and the conclusion for each, its support level and end-of-support date, and a plan for replacing it. Pin versions so that what was tested is what ships, generate the SBOM from the build rather than by hand, and review the anomaly lists again before each release.
- Windows XP's end of support in 2014 left devices running unpatched for years, as WannaCry showed in 2017.
- IEC 62304's SOUP is code of unknown provenance; the FDA's OTS is code whose life cycle the maker cannot control; most components are both.
- For each borrowed component a maker must state requirements, exact version and evaluated anomalies, and keep it under configuration control.
- An SBOM lists every component, now including models and datasets; whether a vulnerability is exploitable is a separate VEX statement.
Evidence at every level.
+ The questionWhich tests, at which level, give evidence a regulator accepts, and when is testing enough?
A sequence of steps nobody combined
In February 2023 Elekta began correcting its Monaco radiotherapy treatment planning system after finding that one sequence of user steps could make it display an inaccurate dose. A planner who added contours to a plan and then re-optimized it, without forcing a density value outside the patient's external outline, could see a dose distribution that did not match what the plan would deliver. The correction covered 2,020 installed systems running four builds of version 5.11, and the FDA recorded the root cause as software design. Each step on its own, drawing a contour, assigning densities, optimizing a plan, displaying a dose, was an ordinary use of the software. The fault lay in the combination.
Failures of this kind are why device software is tested at more than one level. A test of a single function can show that the function meets its own specification. Only tests that put functions together, in the sequences users actually follow, can show how they behave in combination, and only tests of the whole system in its intended environment can show whether the device does what its users need.
Levels of evidence
Verification is the FDA's word for objective evidence that the outputs of a stage of development meet the requirements set for that stage; in short, that the software was built as specified. It happens at several levels. Unit testing checks each smallest piece in isolation, often automatically, against its detailed design. Integration testing checks that units and items work together, including the boundaries Chapter 5 asked an architecture to defend. System testing checks the complete software, usually on the real hardware, against the software requirements. Around the tests sit other verification activities that find faults without running the code at all: reviews of requirements and design, inspection of the code by a second engineer, and static analysis tools that search the source for whole classes of mistakes, such as reading memory that was never written.
IEC 62304 asks for system testing at every class, unit and integration verification at class B and above, and records of each. The FDA's 2023 guidance on premarket software documentation asks, at the Basic level, for a summary of unit, integration and system testing and the protocol and report of the system-level tests; at the Enhanced level it asks for the unit and integration test protocols and reports as well. Both also ask for a list of the unresolved anomalies, the known defects the software will ship with, each with an assessment of its effect on safety and effectiveness. A device can be released with known defects, provided each has been judged and none leaves an unacceptable risk.
How much testing is enough
Chapter 3 showed that the answer cannot be "all of it". Coverage tools measure part of the answer by counting what a test campaign executed, and the FDA's 2002 guidance describes a ladder of coverage measures. Statement coverage, every line run at least once, it calls insufficient to provide confidence. Decision coverage runs each branch both ways. Condition coverage makes each condition inside a decision both true and false. Multiple-condition coverage tries every combination of the conditions in each decision. Full path coverage it calls generally not achievable. The FDA does not prescribe a level; the maker chooses one in proportion to risk and justifies it.
An infusion pump stops when "(line pressure is high AND the occlusion alarm is enabled) OR the door is open". Decision coverage needs two tests: one in which the pump stops and one in which it does not. Condition coverage can also be met with two tests, one with all three conditions true and one with all three false, which never shows what happens when pressure is high but the alarm is disabled. Multiple-condition coverage needs every combination of the three conditions, two to the power three, or 8 tests. For one decision the difference between 2 tests and 8 is small; across a program with thousands of decisions it is the difference between a campaign that can be run every night and one that cannot, which is why the strictest measures are kept for the code whose failure could cause serious harm.
Coverage measures say what ran, so they are paired with measures of what was checked: every requirement traced to at least one test that would fail if the requirement were broken, every hazard-related sequence from the risk analysis tested on purpose, and combinations of settings covered by the all-pairs method of Chapter 3. When a campaign meets those targets, passes, and leaves only judged anomalies, the team has the evidence a reviewer looks for. It has not proved the software correct, and the documentation says so by listing what was not tested and why.
Regression after every change
The FDA's count of the 1990s, in which 79 percent of software recalls followed a change, makes regression testing the most important habit in maintenance. Each change is followed by re-running the tests that cover what it touched and, for anything beyond a trivial change, the whole system-level suite, so that behavior verified before the change is verified again after it. A suite that runs automatically makes this cheap enough to do every time, which is the subject of Chapter 9.
The thread's device shows a regression test on a model. When Digital Diagnostics replaced the classifier that judges image quality in IDx-DR version 2.3, cleared in 2022, it re-ran the new version on the images from the 2017 trial and compared the answers with the original version's. Chapter 21 shows what that comparison found, and what it could and could not prove.
Verification and validation
Verification asks whether the software meets its specification; validation asks whether the specification was right. The FDA defines software validation as confirmation, by examination and objective evidence, that software specifications conform to user needs and intended uses. For a software device, validation reaches beyond testing code: it includes usability studies with representative users and, where the claim requires it, clinical evidence that the outputs are right for patients. IEC 62304 deliberately stops short of it; its scope excludes the validation and final release of the device, which belong to the maker's design controls.
A system test checks the software against requirements the team wrote. If a requirement was wrong, for example an image-quality threshold set for a camera the clinics do not use, every test can pass while the device fails its users. Validation needs evidence from outside the specification: real users, real conditions and, for a diagnostic claim, a comparison with an independent reference.
IDx-DR's 2018 evidence shows both. The trial of the opening scene validated the claim: novice operators with four hours of training, real patients, a reference standard. Alongside it, a smaller study checked repeatability. Twenty-four participants were each imaged ten times by three operators on two cameras of the same model; the system analyzed 235 of the image sets, gave identical answers for 23 of the 24 participants across every repeat, and agreed with itself 99.6 percent of the time. Repeatability is not accuracy, but a screening program whose answer changed with the operator would have failed before accuracy was even asked.
For each software item, state which levels of testing apply, which coverage measure is targeted and why it fits the item's class, how requirements and hazard sequences map to tests, and which test suites are re-run after which kinds of change. Keep the unresolved-anomaly list with each release and record, for every entry, why it leaves no unacceptable risk.
- Device software is verified at the unit, integration and system levels, with reviews and static analysis catching faults that tests miss.
- Coverage measures how much of the code ran; the level chosen should rise with the harm a failure could cause, and is paired with requirement and hazard coverage.
- A release may carry known defects only if each is listed and judged; regression testing after every change guards against the most common source of recalls.
- Verification shows the software meets its specification; validation, including usability and clinical evidence, shows the specification meets patients' needs.
Sprints with a paper trail.
+ The questionDoes IEC 62304 force a team into phases, or can it work in two-week sprints?
A standard without a sequence
IEC 62304 lists activities: planning, requirements analysis, architectural design, detailed design, implementation and unit verification, integration, system testing and release, with maintenance, risk management, configuration management and problem resolution running alongside. Read in that order, the list looks like a waterfall, in which each phase is finished and signed before the next begins. For years many device teams worked that way and assumed the standard required it. The standard asks for the activities to be done and recorded and for the maker's development plan to say how; it does not say in what order, or how many times.
Most software outside regulated industries is now built iteratively. A team works in short cycles of one to four weeks, each ending with working, tested software; requirements are refined as users see results; and the design grows with the product. The question for a device maker is whether that way of working can produce the evidence of Chapters 4 to 7: requirements traced to hazards and tests, a documented architecture, verification at each level, and a controlled release. A technical report from AAMI, the US standards association for medical technology, answers it. TIR45, first published in 2012 and revised in 2023, describes how agile practices map onto IEC 62304, and the FDA has recognized both editions in full, the second in May 2025.
Four layers of time
TIR45 describes the work in four nested layers. The project or product layer runs for months to a year or more and holds the high-level requirements and the overall architecture. A release runs for one to several months and ends with software that could be delivered. An increment, often called a sprint, runs for one to four weeks and ends with integrated, tested software. A story, one small piece of user-visible function, takes one to a few days.
Each IEC 62304 activity is placed in one or more layers. Planning happens at every layer, at its own scale. High-level requirements and the coarse architecture belong to the project; detailed requirements, detailed design and implementation with its unit tests belong to each story; integration and system testing happen at story, increment and release level; and the formal release activity sits at the project layer, carried out each time a release ships. A user story with its acceptance criteria becomes a design input, and the end of an increment or a release becomes a point for design review. The development plan states which of those reviews are formal and recorded.
Suppose a team plans a 12-month project in two-week increments, with a release to users every three months. That is 26 increments and 4 releases. If the team completes about ten stories per increment, it finishes about 260 stories in the year, and under a definition of done that includes them, each leaves updated requirements, design notes, unit tests, trace links and, where it touches a hazard, an updated risk analysis. The four release ends carry the formal system test, the anomaly review and the release approval. Instead of one large documentation effort at the end, the evidence is produced in about 260 small pieces, each reviewed while the work is fresh, and each release is assembled from records that already exist.
Software can move from a developer's machine to a test environment many times a day, and an increment can be demonstrated every two weeks. A release to patients is different: IEC 62304 requires that verification be complete, that known residual anomalies be documented and evaluated, that the released version and how it was built be recorded, and that the software be archived and reliably delivered. Agile speeds up everything before that point. It does not shrink the point itself.
Done means verified and recorded
The practice that makes agile work under a standard is a strict definition of done. A story is not finished when its code runs; it is finished when its requirement is written and reviewed, its design and code are reviewed, its unit and integration tests pass, its trace links are in place, the risk analysis reflects any new hazard or control, and any change to SOUP is recorded. If any of these is missing, the story goes back into the work, exactly as it would if a test failed.
Two further habits keep the records honest. The backlog of stories is not the requirements specification, but it feeds it: as stories are accepted, their requirements enter a controlled baseline that is approved at each release, so that the release is checked against a stable specification rather than a moving list. And traceability is kept live in the tools the team already uses, linking stories, requirements, code changes and tests as the work happens, so that the trace matrix of Chapter 4 is a report generated from the links rather than a document written afterwards.
What changed in 2023
The 2023 edition of TIR45 kept the core and updated the setting. According to AAMI, it adapts documentation to digital practice, with streamlined approval processes and signature requirements, and brings cybersecurity, risk management, design validation and usability engineering into agile work rather than treating them as separate phases. Its references were brought up to date with ISO 13485:2016, the amended IEC 62304 and ISO 14971:2019. The FDA accepts declarations of conformity to the 2012 edition only until 2 July 2028.
The same shift appears elsewhere. The second edition of ISPE's GAMP 5, the guide that regulated companies use to validate the computerized systems they buy and build, published in 2022, states that its life cycle is not inherently linear and supports iterative and incremental methods; Chapter 19 returns to it. The draft second edition of IEC 62304 itself, according to the German electrical standards association VDE in July 2026, places agile development in an informative annex.
Write the layers, their lengths and the activities placed in each into the software development plan, and state which reviews are formal. Make the definition of done include requirements, design, tests, trace links and risk updates for every story, and approve a requirements baseline at each release. An auditor who opens the plan should find the team's real way of working described there, not a waterfall it does not follow.
- IEC 62304 requires activities and records, not an order of phases, so iterative development can meet it if the plan says how.
- AAMI TIR45, recognized by the FDA, maps the standard's activities onto project, release, increment and story layers.
- A strict definition of done, a requirements baseline at each release and live traceability let evidence accumulate as the work is done.
- Increments can be frequent, but a release to users stays a controlled event with verification, anomaly review, archiving and approval.
Every build a record.
+ The questionWhen code is built and tested automatically on every change, what turns that pipeline into evidence, and where must a person still decide?
Deploying many times a day
In the 2024 survey of software delivery run by DORA, a research program at Google Cloud, the best-performing fifth of respondents, 19 percent, deployed changes on demand, often several times a day. A change took less than a day to go from code to production, 5 percent of their deployments failed, and they recovered from a failed deployment in less than an hour. The survey covers technology workers in general, not medical device makers, and the gap is deliberate: a device maker cannot put a new version in front of patients several times a day, and should not.
What the best general software teams do have is a pipeline: an automated sequence that takes every change a developer submits, builds the software from source, runs the tests, checks the code with analysis tools, and packages the result, with no step done by hand. Device makers increasingly use the same machinery, not to release faster to patients but to make every internal build a complete, repeatable record. The pipeline turns the regression testing of Chapter 7 from an occasional campaign into something that happens on every change.
At the elite team's change failure rate of 5 percent, three deployments a day produce 0.15 failed deployments a day, about one a week counting every day. In a test environment that is a cheap lesson: the failure is found within hours and the change is rolled back. For released device software the same arithmetic would mean a field failure each week. The device maker's answer is to run the frequent cycle internally, where failures are expected and contained, and to keep the release to patients as a separate, verified and approved event, as Chapter 8 described.
What makes a pipeline evidence
A pipeline produces evidence when its outputs can be trusted and traced. Configuration management, which IEC 62304 requires at every class, means that every item that goes into a build, source code, SOUP versions, build scripts, the compiler and its settings, is identified and controlled, so that the same inputs produce the same software. A reproducible build goes one step further: rebuilding a past release from its recorded inputs gives a bit-for-bit identical result, so the version tested and the version shipped can be shown to be the same.
Each stage then leaves a record tied to the exact build it checked. Unit and integration test results are stored with the commit that produced them and linked to the requirements they verify. Static analysis reports, coverage reports and the software bill of materials of Chapter 6 are generated by the pipeline rather than assembled by hand, so they describe the build that exists rather than the one someone remembers. The full system tests, which may take hours on real hardware, run nightly or before each candidate release. When a release is proposed, the evidence for it is not collected; it is retrieved.
Two decisions stay with people. Someone must judge each unresolved anomaly and decide that none leaves an unacceptable risk, and someone authorized must approve the release itself. The pipeline can check that every required record exists and every test passed. Whether the remaining defects are acceptable for patients is a judgment, and IEC 62304 and the regulators expect a named person to make it.
Trusting the tools
The pipeline's tools are software too: the build system, the test runner, the coverage tool, the static analyzer, the SBOM generator, the issue tracker that holds the trace links. None of them ships in the device, so none is SOUP, but each is quality-system software that must be validated for its intended use, whether bought, open-source or written in-house. A test runner that reports a pass for a test that never ran, or a coverage tool that miscounts, would manufacture false evidence about a device that might be dangerous.
ISO 13485:2016, the quality-management standard that the FDA's rules have incorporated since February 2026, requires a maker to validate software used in its quality system and in production before first use and after changes, in proportion to the risk of its use. ISO's technical report 80002-2, published in 2017, gives guidance for doing so. The FDA's guidance on computer software assurance, finalized on 24 September 2025 and reissued on 3 February 2026 under the title Computer Software Assurance for Production and Quality Management System Software, sets out a risk-based approach. A function is high process risk if its failure "may result in a quality problem that foreseeably compromises safety"; such functions get documented, scripted testing, while others can be assured by lighter, unscripted methods such as exploratory testing. The guidance replaces one section of the FDA's 2002 software validation guidance and explicitly does not cover the device software itself, which remains under design controls.
A test step that silently skips tests after a configuration change, a runner that treats a crash as a pass, or an SBOM generator that misses dependencies loaded at run time all produce records that look complete. Seeding known faults and confirming that the pipeline catches them is the most direct check, and it should be repeated whenever a tool or its configuration changes.
After release: fixes and maintenance
Release starts the longest part of a device's software life. IEC 62304 requires a maintenance plan and a problem-resolution process: field reports and internal findings become problem reports, each is investigated and classified for its effect on safety, changes go through the same controlled process as development, and trends across problems are analyzed. Chapter 2's lessons apply with full force here. In the FDA's count from the 1990s, 79 percent of software recalls followed changes made after release, and the Tandem app recall of 2024 was a change to software that already worked.
Hosted software changes the mechanics. When the analysis runs on a maker's server, as IDx-DR's does, a fix can reach every user at once and can be rolled back at once, which is an advantage for safety and a risk for control: a change can also reach every user at once. A switch that turns on a function for some users, often called a feature flag, is a change to the released device and goes through change control, however small the switch; whether it needs a regulator's review depends on its effect on claims and risk (Chapter 28). And much code in the field predates the processes now expected of it. The 2015 amendment to IEC 62304 added a route for this legacy software, legally marketed but without enough evidence that it was built to the standard: a risk analysis using field experience, a gap analysis against the standard's requirements, a plan to close the gaps that matter, and a documented rationale for continued use. Whether a given change needs a regulator's approval before release is a separate question, which Chapter 28 answers for the US, the EU and India.
List the tools in the pipeline and, for each, the records it produces and what a silent failure would do to the device's evidence. Give the tools that produce verification results, coverage, builds and SBOMs scripted testing with seeded faults; pin their versions and archive the build environment with each release; and use lighter, unscripted assurance for tools whose failure could not compromise a release.
- A pipeline builds and tests every change automatically; device makers use it to make each build a complete record, not to release faster to patients.
- Configuration management, reproducible builds and records tied to each build let a release's evidence be retrieved rather than assembled.
- Pipeline tools are off-the-shelf software that must be validated in proportion to risk, the approach of the FDA's computer software assurance guidance.
- After release, maintenance and problem resolution follow the same controls, and legacy software needs a gap analysis and a rationale for continued use.
+ Part III · Learning from data
How an AI model is built and scored.
In 2016 a team at Google trained a network to grade photographs of the retina for diabetic retinopathy without writing a single rule about what the disease looks like: no instruction to look for small hemorrhages, no threshold for the size of a lesion. Every rule the network applied had been fitted to 128,175 images that ophthalmologists had already graded. A program built that way can still be tested, but only by comparing its answers with right answers on data it has not seen, and every step of that comparison involves choices that decide what the result means. The four chapters of this part follow those choices: what a model is, how the data that builds it is kept apart from the data that judges it, what the right answers are and who decided them, and how a score becomes a decision with a measurable error.
Rules nobody wrote.
+ The questionWhat is a machine-learning model, and how is building one different from writing code?
Twenty-five million adjustable numbers
ResNet-50, one of the most widely used networks for analyzing images, has 25,557,032 adjustable numbers in the reference implementation published with the PyTorch software library. Each is a parameter: most are weights that multiply values inside the network before they are passed on, and the rest are offsets added along the way. A photograph goes in as a grid of pixel intensities; it passes through some fifty layers, each of which combines the values from the layer before using its weights; and a score comes out, for example the network's estimate that the image shows a particular disease. Change the weights and the same photograph produces a different score.
A machine-learning model is that whole arrangement: a fixed structure, chosen by engineers, with its parameters set from data rather than by hand. The structure is code like any other, written, reviewed and tested. The parameters are not written at all. They are found by training: the model scores a batch of examples whose right answers are known, a measure of how wrong it was, called the loss, is computed, and every parameter is nudged slightly in the direction that would have made the loss smaller. Repeated over millions of batches, the nudges settle into values that make the model's scores match the known answers well. The Google system of 2016, an ensemble of ten Inception-v3 networks, each of which its authors describe as having about 22 million parameters, was trained this way on 128,175 retinal images, each graded three to seven times by a panel of 54 US ophthalmologists and senior residents.
Fitting, and fitting too well
Training finds parameters that fit the examples. Whether they fit anything else is a separate question, and the reason is visible in the numbers above: the network had more parameters than there were images to fit. A model with that much freedom can, in principle, memorize its training set, giving the right answer for every image it has seen and an arbitrary one for any image it has not. Engineers limit this by the choice of structure, by starting from parameters already trained on more than a million ordinary photographs, by penalties that keep weights small and by stopping training before the fit becomes too close. None of these guarantees that the model will work on new images; only a test on new images can show it.
Take ten points along a straight line, y equals 2x for x from 0 to 9, each moved up or down by a small fixed amount, between about −1.4 and +1.3.4. A straight line, two parameters, fitted to the ten points misses them by 0.93 on average (root mean square). A polynomial with ten parameters can pass through every point exactly, so its error on the training points is zero. Now test both on nine new points taken halfway between the old ones, from the same line with fresh noise. The straight line misses them by 0.78 on average; the polynomial misses them by 1.67, more than twice as much, because between the training points it bends to follow the noise. The model with the better training score is the worse model.
The omics guide in this library meets the same problem in another form: a molecular signature discovered from many more measurements than patients fits its discovery data almost perfectly and fails on the next cohort, and only a locked signature tested on independent samples shows whether it is real. An image network differs in scale, not in principle. Its training score is a description of the past; its value is a prediction about data it has not seen.
The data is part of the design
Because the parameters come from the examples, whatever regularities the examples contain are candidates for the model to learn. Some are the disease; others are accidents of how the data was collected: which hospital an image came from, which camera took it, what text was stamped in its corner. A model has no way to tell a sign of disease from a sign of the clinic where diseased patients were photographed, unless the data makes the difference visible. Chapter 11 shows a network that learned exactly that.
This is the sense in which the invariant of this guide extends to AI. A conventional program's faults are written into its code; a model's faults can also be written into its data, by what the data contains, what it lacks and how its labels were assigned. The device is the code, the parameters and the data that set them, together with the processing that turns a raw image into the grid of numbers the model expects. Changing any of them changes the device.
Reviewers can read the code of a model's structure, but they cannot read its parameters as rules, and no list of the features a network relies on is complete. Evidence about what a model does therefore comes from its behavior on carefully chosen data, and the quality of that data, not the elegance of the code, sets the limit on what can be claimed.
Locked and learning models
A model whose parameters are fixed after training is locked: the same input always gives the same output, and it can be verified like any other software. IDx-DR's analysis was locked before the 2017 trial began, as the paper states. A model that keeps adjusting its parameters from new data in the field is continuously learning: its behavior on Tuesday may differ from Monday's, and the evidence gathered before release describes a model that no longer exists.
In 2019 the FDA wrote that the AI devices it had cleared or approved had typically "only included algorithms that are 'locked' prior to marketing." Updates happen, but as new versions: retrained, tested and released through the change controls of Chapter 21. When one of the earliest authorized plans for changing an AI device, for a program that guides cardiac ultrasound, was modified by a 510(k) in 2020, the summary stated that all algorithm modifications would be "trained, tuned, and locked prior to release" and excluded continuously learning algorithms. A draft EU manufacturing guideline of 2025 on AI in medicines production goes further, saying that dynamic models should not be used in critical applications at all. The question of how a model may change after authorization is the second half of this guide's central question, and Part V answers it.
Put the model's structure, its trained parameters, the preprocessing code and an identifier of the exact training data under configuration management as one released item, with a version that changes whenever any of them does. Record the data's origin, size and labeling method alongside, so that a later reviewer can tell which data produced which version of the model.
- A machine-learning model is a fixed structure whose millions of parameters are set by training on examples with known answers.
- A model can fit its training data closely and still predict poorly; only a test on data it has not seen shows whether it generalizes.
- The training data is part of the design, so a model can learn accidents of data collection as readily as signs of disease.
- Locked models give the same output for the same input; most authorized AI devices are locked and change only through new, tested versions.
Keeping the judge away from the builder.
+ The questionWhy must the data that judges a model be kept apart from the data that built it, and how does it leak back in?
A model that recognized the hospital
In 2018 a group at the Icahn School of Medicine at Mount Sinai in New York trained a network to detect pneumonia on chest radiographs, using images from two hospital systems, with a third held out for external testing, 158,323 in all, and then examined what it had learned. Pneumonia was far more common in one source than in another: 34.2 percent of the images from Mount Sinai Hospital showed it, against 1.2 percent of those from the NIH Clinical Center. A model trained on both sources together scored an area under the curve of 0.931 on test images from the same two sources, an excellent result. Scored on each hospital's test images separately, it reached only 0.805 for Mount Sinai and 0.733 for the NIH.
The gap had a simple cause. A second network could tell which hospital system a radiograph came from in 99.95 percent of the NIH test images and 99.98 percent of Mount Sinai's, helped by cues such as the laterality markers placed on the film, including a metal token. Within Mount Sinai, the portable scanners used on wards and an inverted color scheme on emergency-department images let a network identify even the department. Ranking images by nothing more than the pneumonia rate of their hospital, ignoring every pixel of lung, gave an area under the curve of 0.861 on the pooled test set, most of the way to the model's 0.931 on the same images. Part of what looked like pneumonia detection was hospital detection. The authors reported that the cues "only became apparent to us after manual image review."
Three kinds of data
A model is built and judged with three separate sets of data. The training set is what the parameters are fitted to. The tuning set, which machine-learning engineers usually call the validation set, is used to make choices the training itself does not make: how long to train, how large a model to use, which version of several to keep. The test set is held back until those choices are finished and is used once, to estimate how the finished model will perform. The regulatory word validation, which Chapter 7 used for proof that a device meets users' needs, means something different, and a submission that uses both senses has to say which it means.
The separation matters because every set that influences the model stops being a fair judge of it. A model scored on its training data is scored on examples it was fitted to; a model scored on the data used to pick it is scored on the data that selected it for doing well there. Only data that played no part in building the model estimates how it will do on the next patient. The international principles of good machine-learning practice, finalized by the International Medical Device Regulators Forum in January 2025, state it as their fourth principle: training datasets are independent of test sets.
How the judge leaks into the builder
Independence fails in ordinary ways. The same patient can appear on both sides of a split: two eyes, two visits, several slices of one scan, each a separate image but all carrying the same anatomy and often the same disease. A model that has seen a patient's left eye in training has partly seen the right eye in testing. The Mount Sinai group split its data by patient, 70 percent of patients for training, 10 for tuning and 20 for testing. The FDA wrote the same concern into the rules for retinal diagnostic software in 2018, requiring that analysis of performance make no unjustified assumption that repeated samples from one patient are independent.
Other leaks are quieter. Images can be normalized using statistics computed from the whole dataset, test images included. Duplicates and near-duplicates can sit in both sets. A source can be both common in training and over-represented in testing, so that its cues are rewarded twice, as the hospital cues were. And a team can look at the test results, adjust the model, and look again, until the test set has become a second tuning set. Each of these makes the reported score higher than the score on the next patient will be.
A test set used once gives an unbiased estimate for the model that was finished before the look. Every later decision informed by that score, a changed threshold, a retrained model, a different preprocessing step, makes the next score on the same set more flattering. Keeping a sealed test set for the final estimate, and a fresh one for each major version, keeps the estimate honest.
How large a test set
A test set must also be large enough that the score means something. A sensitivity is a proportion of diseased patients the model detects, and its uncertainty shrinks with the number of diseased patients in the test set, not with the total. A test set with plenty of healthy patients and few sick ones gives a precise specificity and a vague sensitivity.
A team expects a sensitivity of about 87 percent and wants its 95 percent confidence interval to extend no more than 5 percentage points either side. The usual approximation for a proportion gives the number of diseased patients needed: 1.96 squared, times 0.87, times 0.13, divided by 0.05 squared, which is about 174. If the disease is present in 24 percent of the people tested, about 725 participants are needed to include 174 with the disease. The IDx-DR trial was planned for at least 149 participants with the disease and 682 without, to give at least 85 percent power for its prespecified statistical tests, and it analyzed 198 and 621.
The calculation sets a floor, not a target. A test set drawn from one hospital, one camera or one population can be large and still unrepresentative, which is the subject of Part IV. Size makes the estimate precise; only the right composition makes it the estimate that matters.
Split the data by patient and, where possible, by site, before any training begins. Store the test set where the development team cannot use it for training or tuning, record who accessed it and when, and use it once for the final estimate of each version. Size it from the number of positive cases the claimed sensitivity needs, and record its composition so that a reviewer can compare it with the intended population.
- A pneumonia network scored far better on pooled data than on each hospital, because it had learned to recognize the hospital.
- Training, tuning and test sets must be separate; only data that played no part in building a model estimates its performance.
- Splitting by image instead of patient, shared preprocessing, duplicates and repeated looks at the test set all inflate the reported score.
- Precision depends on the number of diseased patients in the test set; about 174 are needed to pin an 87 percent sensitivity within 5 points.
What the right answer is, and who decided.
+ The questionWhen a model is scored, what is it scored against?
The answer key
For every participant in the IDx-DR trial, the right answer was decided at the Fundus Photograph Reading Center of the University of Wisconsin, from images the program never saw. Certified photographers took four overlapping stereo pairs of each retina, covering a far wider field than the program's two pictures per eye, and a scan of the macula by optical coherence tomography with at least 121 cross-sections. Three experienced readers graded the photographs on the severity scale of the Early Treatment Diabetic Retinopathy Study, without seeing the program's output or the other imaging, and the majority grade stood. A participant counted as having more than mild diabetic retinopathy if the worse eye reached level 35 or higher on that scale or showed macular edema.
That protocol is the trial's reference standard, often called ground truth: the best available judgment of the true state, against which the device is scored. The trial's headline numbers, the 87.4 percent sensitivity, the 89.5 percent specificity and the predictive values of Chapter 13, are statements of agreement with the photographic grade; adding the macular scans, in a secondary analysis, gave 85.9 and 90.7 percent, figures adjusted for the trial's enrichment. If the reference were wrong for some participants, the numbers would be wrong in ways the trial could not detect.
Experts disagree
Grading retinal photographs is skilled work, and skilled graders differ. A 2018 study by researchers at Google, with retina specialists from two clinical practices, measured how much. Three fellowship-trained retina specialists graded 1,813 photographs independently and then resolved every disagreement in adjudication sessions, producing a consensus grade for each image. Measured against that consensus, the three specialists working alone detected moderate or worse retinopathy in 74.6, 74.4 and 82.1 percent of the images that had it, while correctly clearing about 99 percent of those that did not. Three board-certified general ophthalmologists did similarly. A single expert, in other words, missed about one case in four that the adjudicated panel found.
The same study showed how much the reference matters to a model. Agreement between each grader and the consensus, measured on a weighted scale where 1 means perfect, ranged from 0.80 to 0.91 for the specialists. Tuning the network, an improved version of the Google system of Chapter 10, on a small set of adjudicated grades improved it substantially, and the improved network scored an area under the curve of 0.986 against the adjudicated reference. How good a model looks depends on which human answer it is compared with, and a model trained or tested against a single grader inherits that grader's misses.
If a model is trained on one grader's labels and tested against the same grader, a high score shows that the model has learned to agree with that person, including the cases that person gets wrong. Independent, adjudicated or multimodal reference standards cost more because they are the only way to measure accuracy rather than imitation.
The reference caps what can be measured
An imperfect reference makes a perfect device look imperfect. When the device is right and the reference is wrong, the disagreement is counted as the device's error; when both are wrong in the same direction, the error is invisible.
Take 1,000 people, of whom 240 truly have the disease, the 24 percent prevalence of the IDx-DR trial. Suppose the reference standard misses 5 percent of true disease, labeling 12 of the 240 as healthy, and wrongly labels 1 percent of healthy people as diseased, about 8 of the 760. Now score a model that is always right. The reference calls 236 people diseased, 228 truly diseased and 8 not; the model flags the 228, so its apparent sensitivity is 228 out of 236, 96.6 percent. The reference calls 764 people healthy, including the 12 it missed; the model clears 752 of them, so its apparent specificity is 98.4 percent. A flawless model loses more than three points of sensitivity to the reference's own errors.
The direction of the bias depends on how the reference and the device err. If both struggle with the same hard cases, small lesions or poor images, they may agree on the wrong answer and inflate the score. This is one reason the IDx-DR trial used a reference built from more and different information than the device received: wider photographs and stereo views, with a scan of the macula for a secondary analysis, read by people who never saw the device's answer. Disagreements then reflect the device's limits rather than shared blind spots.
A reference fit for the claim
Regulators' own wording has moved on this point. The good machine-learning practice principles published by the FDA, Health Canada and the UK's MHRA in 2021 asked that reference datasets be "based upon best available methods"; the international version finalized in January 2025 asks instead that reference standards be "fit-for-purpose". One reading of the change is that the best method may be impossible or unethical for every patient, and that the right reference depends on the claim. A stroke-triage program authorized in 2018 was scored against neuroradiologists' readings, with an additional neuroradiologist breaking disagreements. A dermatology dataset assembled at Stanford for testing skin-lesion models used biopsy results, so every label was confirmed by a pathologist. IDx-DR used a reading-center protocol because its claim was a screening decision about retinopathy severity, which is exactly what that protocol grades.
Whatever the choice, the reference must be defined before the test, applied the same way to every case, independent of the device's output, and described in the labeling, so that a reader of the results knows what the device was compared with. Chapter 17 shows how to find that description in a public summary and what it tells a reader about the limits of the numbers.
Write down the reference standard in the study protocol: who grades, with what information, using which scale, how disagreements are resolved, and how graders are kept from seeing the device's output. Measure and report the agreement among graders, plan adjudication for disagreements, and prefer a reference that draws on more or different information than the device uses.
- IDx-DR was scored against a reading-center reference built from wider stereo photographs and macular scans, read by three masked graders.
- Single experts disagree with an adjudicated consensus; retina specialists working alone missed about a quarter of moderate or worse cases.
- An imperfect reference lowers a perfect model's apparent accuracy and can hide errors that model and reference share.
- A reference standard must be fit for the claim, fixed before testing, independent of the device and described in the labeling.
A threshold turns a score into a decision.
+ The questionA model outputs a score; who chooses where yes begins, and what does moving that point trade?
Four counts
Of the 819 participants whose results could be analyzed in the IDx-DR trial, 198 had more than mild retinopathy by the reference standard and 621 did not. The program flagged 173 of the 198 and missed 25; it cleared 556 of the 621 and flagged 65 who did not have the disease. Every performance figure for the trial comes from those four counts. Sensitivity, the share of diseased participants flagged, is 173 out of 198, 87.4 percent. Specificity, the share of disease-free participants cleared, is 556 out of 621, 89.5 percent. The positive predictive value, the share of flagged participants who had the disease, is 173 out of 238, 72.7 percent; the negative predictive value, the share of cleared participants who did not, is 556 out of 581, 95.7 percent.
Each figure is an estimate from a sample and carries uncertainty. The exact 95 percent confidence interval for the observed sensitivity runs from about 81.9 to 91.7 percent; for specificity, from about 86.9 to 91.8 percent. The paper's headline figures, 87.2 and 90.7 percent, are slightly different because the trial deliberately recruited extra participants with poorly controlled diabetes, and the authors corrected for that enrichment with a statistical model. The trial's statistical test asked whether the lower end of each interval cleared a floor set in advance, 75 percent for sensitivity and 77.5 percent for specificity, while the alternatives of 85 and 82.5 percent, used to size the trial, are what the FDA's summary calls the prespecified thresholds; the estimates exceeded both. A study has to clear its floor with room to spare, and a small study, with wide intervals, cannot.
A threshold turns a score into a decision
Most models do not output a decision. They output a score, and a threshold turns it into one: above it, refer; below it, do not. Moving the threshold trades one error for the other. A lower threshold flags more patients, catching more disease and raising more false alarms; a higher one does the reverse. The receiver operating characteristic curve, or ROC curve, plots that trade across every possible threshold, sensitivity against the false-positive rate, and the area under the curve, between 0.5 for a coin toss and 1 for a perfect separation, summarizes how well the score separates the two groups regardless of where the threshold is set.
The Google network of Chapter 10 shows the choice in practice. On one validation set of 8,788 gradable images, its area under the curve was 0.991, and its authors reported two operating points on the same curve: a high-specificity point with 90.3 percent sensitivity and 98.1 percent specificity, and a high-sensitivity point with 97.5 percent sensitivity and 93.4 percent specificity. The model was the same in both cases. Which point to ship depends on what a false alarm and a missed case each cost in the setting of use, a clinical judgment rather than a property of the model, and it must be fixed before the test, because a threshold tuned on the test set inflates the result as Chapter 11 warned.
An area under the curve describes the whole curve, including thresholds nobody will use. Two models with the same area can have very different sensitivity at the specificity a clinic needs. A claim is only interpretable with the threshold, the sensitivity and specificity at it, their confidence intervals, the prevalence in the test population and the share of cases the device declined.
Prevalence changes what a positive means
Sensitivity and specificity describe the device; predictive values describe what its answers mean in a particular population, and they change with how common the disease is. The PCR and immunoassay guides in this library work through the same arithmetic for laboratory tests; it applies unchanged to software.
At IDx-DR's observed sensitivity of 87.4 percent and specificity of 89.5 percent, take a population in which 24 percent have the disease, as among the trial's analyzed participants. Of every 1,000 people, 240 are diseased and the program flags 210 of them; 760 are healthy and it flags 80 of them by mistake. A positive answer is right 210 times out of 290, about 72 percent, close to the trial's 72.7 percent. Now take a population in which 7 percent have the disease: 70 diseased, of whom 61 are flagged, and 930 healthy, of whom 98 are flagged. A positive answer is now right only 61 times out of 159, about 38 percent, while a negative answer is right about 99 percent of the time. The program is unchanged; the meaning of its answers is not.
Low prevalence and a weak score together produce a heavy burden of false alarms. In 2021 researchers at Michigan Medicine tested a sepsis prediction model built into a widely used electronic health record on 38,455 hospitalizations, 7 percent of which involved sepsis. The model's area under the curve was 0.63, against 0.76 to 0.83 reported in the developer's own documentation. At the alert threshold the hospital used, it flagged 18 percent of hospitalizations, caught 33 percent of sepsis cases and missed 67 percent, and only 12 percent of the hospitalizations it flagged involved sepsis: alerting once per patient, clinicians would evaluate about eight patients to find one with sepsis. The published counts reproduce every rate: 843 of 6,971 flagged hospitalizations had sepsis.
Answers the program declines to give
A diagnostic program can also decline to answer, and how those cases are counted changes the figures. IDx-DR returned "insufficient quality" for some participants even after dilation, and its labeling directs that such patients be referred. The trial's headline sensitivity and specificity exclude them: the program gave a usable answer for 96.1 percent of participants whose reference images could be graded. A complete account therefore reports how often the device answers, as well as its sensitivity and specificity when it does. Chapter 17 recomputes IDx-DR's sensitivity with the declined cases counted, and the immunoassay guide shows the same question for a blood test with a zone that gives no answer.
Decide with clinicians what a missed case and a false alarm each cost in the intended setting, choose the threshold on the tuning data accordingly, and freeze it before the test set is opened. Report sensitivity, specificity and their intervals at that threshold, the prevalence of the test population, predictive values for the populations where the device will be used, and the rate of declined answers.
- Sensitivity, specificity and predictive values all come from four counts; IDx-DR's trial gave 87.4 percent sensitivity and 89.5 percent specificity.
- A threshold turns a model's score into a decision, trading missed cases against false alarms along the ROC curve.
- Predictive values change with prevalence: the same program's positive answers are right about 72 percent of the time at 24 percent prevalence and 38 percent at 7.
- A full account also reports how often the device declines to answer, and the operating point behind any summary number.
+ Part IV · Patients the model never met
Performance where it is used.
In two studies of children and young adults with diabetes, for whom IDx-DR is not labeled, its specificity was about 79 percent, against about 90 percent in the adults of the 2017 trial: roughly twice as many false referrals among patients without the disease. The comparison is not like for like, since the youth studies used their own reference readings rather than the adult trial's reading-center protocol, and the company's founder was an author of the later one. But the patients were different, and so was the result. A model's accuracy is a measurement made on particular people, images, references and conditions, and it moves when any of them moves. The four chapters of this part follow that movement: what happens when a model meets a different setting, how it can be accurate on average and wrong for a group, what changes when a clinician and a model decide together, and how to read a published performance claim for what it does and does not cover.
Tested on yesterday's patients.
+ The questionWhy does a model that passed its test do worse in another clinic, on another camera, or a year later?
Eleven clinics in Thailand
Thailand had, by one press account, about 4.5 million people with diabetes and around 200 retinal specialists to examine their eyes when Google and the Thai Ministry of Public Health began deploying a deep-learning screening system in primary care. Over eight months, researchers made regular visits to 11 clinics in the provinces of Pathum Thani and Chiang Mai, watching nurses photograph patients' eyes and interviewing them about the system. The study, presented at a human-computer interaction conference in 2020, did not measure grading accuracy. It found a model that often would not grade at all.
The system had been built to refuse images below a quality threshold, a sensible safety choice in the laboratory. In the clinics, lighting differed from room to room and images came out blurred or with dark areas, so the system marked them ungradable. A press account of the study reported that more than a fifth of images were rejected. The protocol at first sent patients with rejected images to a specialist. The researchers changed it so that eye specialists reviewed the ungradable images together with the patient's records, instead of referring everyone automatically. The authors, most of them Google staff, described the core finding as a tension between the model's requirements for image quality and the images an under-resourced setting could produce.
Google's system in a prospective cohort
A separate study in Thailand measured accuracy directly. Between December 2018 and March 2020, nine primary-care sites in the national screening program used the system in real time, with regional retina specialists over-reading every image as a safety measure. Of 7,940 people screened, 7,651 were analyzed and 31.5 percent were referred; the accuracy figures come from a smaller subset whose size the abstract does not give. For vision-threatening retinopathy, the system's sensitivity was 91.4 percent against 84.8 percent for the specialists, and its specificity was 95.4 percent against their 95.5 percent, according to the 2022 paper, whose authors include Google staff and whose funding came from Google and a Thai hospital.
Taken together, the two studies show one common shape of real-world shift. In the cohort study the model's accuracy matched or beat the specialists' over-reads. The losses came outside the grading model, in the quality gate and the workflow around it: lighting, image quality and the referrals that followed a rejected image. A test on curated images measured one link of a longer chain, and the links it did not measure decided much of what patients experienced.
Kinds of shift
A model is fitted and tested on one distribution of cases: patients, images and labels in particular proportions. Dataset shift is any difference between that distribution and the one the model meets in use, and it comes in recognizable kinds. Input shift, also called covariate shift, changes what the model sees: a different camera, different lighting, a different scanner protocol, an older or younger population. Prevalence shift changes how common the disease is, which Chapter 13 showed changes what a positive answer means. Concept shift changes what counts as the right answer: a new grading guideline, a different reference. And acquisition shift hides inside the processing chain: a software update on a scanner, a new image format, a change in compression, invisible to a person looking at the image and possibly visible to a model.
In the same study as the pneumonia network of Chapter 11, a network trained only on Mount Sinai's images met a shift between hospital systems: its area under the curve fell from 0.802 on Mount Sinai's test data to 0.717 on the NIH's. Time produces shifts too. Populations change, treatments change the appearance of disease, and clinical practice changes which patients are tested at all. A model tested once describes the cases it was tested on, at the time it was tested.
In the adult trial, IDx-DR cleared 89.5 percent of participants without the disease, so among every 1,000 such adults it would refer about 105 unnecessarily. In the youth studies its specificity was about 79 percent, so among every 1,000 young people without the disease it would refer about 210 unnecessarily, twice as many. Sensitivity in one of those studies was reported as about 86 percent, close to the adult figure. A screening program that used the device outside its labeled age range would see its specialist clinics fill with false referrals, a cost that no adult study could have shown.
External and prospective tests
Two kinds of test reach further than a held-out split. An external test uses data from sites, devices or periods that contributed nothing to development, and it shows how the model travels. A prospective test runs the device forward in its intended setting, on patients enrolled for the purpose, with the workflow and the operators it will really have. The IDx-DR trial was prospective in primary care, which is why its figures carry more weight than a retrospective test on archived images would. The good machine-learning practice principles ask for testing that demonstrates performance under clinically relevant conditions, and both kinds answer that request.
A model that holds its accuracy at a second hospital has passed one more test, not the test of every hospital. Sites differ in equipment, protocols, populations and workflow at once, and a single external result cannot separate those effects. Several external sites, chosen to span the intended use, say far more than one large one.
Neither removes the problem, because the next site is always new. What an external and prospective test can do is show that performance survives the kinds of difference the intended use contains, and name the conditions it was tested under, so that a user can see when a site falls outside them. Chapter 19 describes the check a hospital can run before it relies on a device, and Chapter 20 the monitoring that continues afterwards.
Record, for every test set, the sites, devices and their software versions, image formats, acquisition protocols, operators, population characteristics and dates. State them in the labeling as the conditions under which performance was shown, and give users a way to compare their own setting with them before use.
- In Thai clinics a retinal model rejected many images taken in real lighting, and the workflow around those rejections decided much of what patients experienced.
- In a prospective Thai cohort, the same kind of system matched or beat retina specialists' accuracy, so much of the loss lay outside the model.
- Dataset shift changes inputs, prevalence, the right answer or the acquisition chain; IDx-DR's specificity fell from about 90 to 79 percent outside its labeled age range.
- External and prospective tests show how far performance travels, and the conditions they cover belong in the claim.
Accurate on average, wrong for some.
+ The questionCan a model be accurate overall and still fail one group of patients, and how would anyone find out?
A test set built to compare skin tones
Three published models for telling malignant skin lesions from benign ones had reported areas under the curve between 0.88 and 0.94 on their own test sets. In 2022 a group at Stanford tested them on a new collection, the Diverse Dermatology Images set: 656 images from 570 patients seen at Stanford clinics between 2010 and 2020, every diagnosis confirmed by a biopsy report, with images of the lightest and darkest skin tones matched by diagnosis, age, sex and date. On the whole set, the three models scored between 0.56 and 0.67. Split by skin tone, the drop was steeper for darker skin. In the figures of the group's preprint, the best of the three scored 0.72 on the lightest skin types and 0.57 on the darkest; another scored 0.61 and 0.50, no better than chance.
The published paper adds two findings that sharpen the lesson. Dermatologists reading the same images, scored against the biopsy results, were also less accurate on darker skin, which matters because dermatologists' judgments supply the labels for most training sets. And fine-tuning two of the models on images from the new set closed the gap between light and dark skin, which suggests that the original failure came largely from what the models had been trained on.
Averages hide groups
A single accuracy figure for a whole test set is an average over its members, weighted by how many of each kind it contains. If one group is small, the average can be high while that group's performance is poor, and nothing in the headline reveals it. A 2021 study of chest-radiograph classifiers trained on three large public datasets, together more than 700,000 images, measured how often each called a sick patient's film normal. The rate was higher for female patients, for patients under 20, for Black and Hispanic patients and for patients insured through Medicaid, and higher still where these characteristics combined, for example in Hispanic women compared with white women.
The remedy is to look. The IDx-DR trial reported its population, 28.6 percent of participants African American, 16.1 percent Hispanic, 63.4 percent White and 1.6 percent Asian, with a median age of 59, and analyzed sensitivity by age, sex, race, ethnicity, blood-sugar control, lens status and site. None had a significant effect on sensitivity; specificity was somewhat higher in participants over 65. Such an analysis cannot prove that no group does worse, but it can show that no large difference was detected among the groups present in meaningful numbers, and it tells a reader which groups were too small to say.
A subgroup result of 90 percent from a few dozen patients is compatible with performance far below the overall figure. Reporting the point estimate alone invites the conclusion that the group is served well. The interval, and the number of patients behind it, show whether the evidence is strong enough to conclude anything.
How many patients a subgroup needs
The precision of a subgroup estimate depends on the number of patients in that subgroup, which is usually a fraction of the whole test set. A study sized for a precise overall sensitivity may give a vague one for every group within it.
A model detects disease in 45 of 50 diseased patients in a subgroup: 90 percent. The 95 percent confidence interval, by the Wilson method that the FDA's guidance on diagnostic statistics uses in its examples, runs from 78.6 to 95.7 percent. With 450 detected out of 500, the estimate is still 90 percent, but the interval narrows to 87.1 to 92.3 percent. With 9 out of 10, it spans 59.6 to 98.2 percent and says almost nothing. To show that a subgroup's sensitivity is within a few points of the overall figure, the subgroup needs hundreds of diseased patients of its own, which is why subgroups are planned when the study is designed, not discovered afterwards.
Looking at many subgroups brings its own trap. A study that compares twenty subgroups will usually find one or two that differ by chance alone, so subgroups are named in the protocol before the data are seen, and a difference found afterwards is treated as a question for the next study rather than a conclusion. When a planned subgroup does fall short, a maker has three honest responses: add data from that group and retrain, as fine-tuning on the dermatology images did; narrow the claim, as IDx-DR's labeling does by covering only the adults its trial enrolled; or tell users plainly where performance is unproven. Each is a change to the device or its labeling, with the evidence that change requires.
Representative data, by rule
Regulators have written representativeness into their expectations. The good machine-learning practice principles of 2021 asked that clinical study participants and data sets be representative of the intended patient population, and the international version of 2025 asks that clinical evaluation use datasets representative of that population. When the FDA created the regulation for stroke-triage software in 2018, its special controls required results showing effective triage across relevant subgroups. India's 2026 guidance on medical device software goes further for AI: it asks makers to disclose the demographic and geographic composition of training, validation and test data, to justify models trained or validated outside India, and to assess bias, generalizability and robustness across relevant Indian sub-populations. The EU's AI Act adds data-governance duties for high-risk systems, which Chapter 29 places on its dated timeline.
Representative does not mean proportional. A subgroup that is rare in the population but at high risk may need to be over-sampled so that its performance can be measured at all, and the study then corrects for the over-sampling when it reports overall figures, as the IDx-DR trial did for its enrichment. What a regulator asks to see is that the groups the intended use contains were present, in numbers that support a conclusion, and that their results are reported.
List the subgroups the intended population contains, including skin tone, sex, age, ethnicity, disease severity, device model and site, and size the study so that each group whose performance matters has enough diseased and healthy patients for a useful interval. Report every planned subgroup with its interval, and say plainly which groups were too small to assess.
- Dermatology models that scored 0.88 to 0.94 on their own test sets fell to 0.50–0.57, at or near chance, on the darkest skin in a balanced, biopsy-confirmed set.
- An overall figure can hide poor performance in a small group; chest-radiograph classifiers missed disease more often in several under-served groups.
- A subgroup estimate is only as precise as the number of patients in it: 90 percent from 50 patients spans about 79 to 96 percent.
- Regulators in the US, the EU and India now expect data representative of the intended population and results reported by subgroup.
The reader and the machine.
+ The questionWhen AI assists a clinician rather than deciding alone, what has to be measured: the model, the person, or the pair?
When the suggestion is wrong
An assistant that is right most of the time can still make an expert worse. In a study published in 2023, 27 radiologists read 40 test mammograms, each shown with a suggested category on the standard breast-imaging scale that the readers were told came from an AI system. The suggestions had in fact been set by the investigators, and 12 of the 40 were deliberately wrong. When the suggestion was right, readers of every level of experience rated about 80 percent of mammograms correctly. When it was wrong, inexperienced readers rated 19.8 percent correctly, moderately experienced readers 24.8 percent and very experienced readers 45.5 percent.
The pattern has a name: automation bias, the tendency to accept a machine's output in place of one's own judgment, which studies suggest is strongest when the person is uncertain. It turns the model's errors into the clinician's errors, and it is strongest among the people an assistant is often meant to help most, those with less experience.
In the study's design, the suggestion was right in 28 of 40 cases, 70 percent. A reader who is right 79.7 percent of the time with a correct suggestion and 19.8 percent with a wrong one ends at 0.7 times 79.7 plus 0.3 times 19.8, about 62 percent overall. A very experienced reader, right 82.3 percent and 45.5 percent of the time, ends at about 71 percent. The model's error rate, multiplied by how far each reader follows a wrong suggestion, sets most of the gap between them. The study's 30 percent rate of wrong suggestions was chosen to measure the effect, not to mimic any product; a better model would make wrong suggestions rarer, and each one may be harder to spot.
Three roles for a model
Software that analyzes clinical data plays one of three roles, and each changes what has to be measured. An autonomous device gives the answer itself: IDx-DR's output is the screening decision, and no clinician interprets the images first, so the device's own accuracy is what reaches the patient. An assistive device gives a reader extra information, such as marks on suspected lesions, and the reader decides; the performance that matters is the reader's with the device compared with the reader's without it. A triage or notification device works alongside the usual workflow: the stroke-triage program the FDA authorized in February 2018 was described as a "notification-only, parallel workflow tool" that alerts a specialist to a suspected blocked artery while "trained radiologists read all images per standard of care, regardless of the performance" of the device.
The role also decides where automation bias can do harm. An autonomous device has no human to bias, but no human to catch its errors either. A parallel triage tool is designed so that its misses do not remove the usual read, although its alerts may still change what gets read first. An assistive device places its output directly in front of the reader at the moment of decision, which is where the mammography study showed the risk. The FDA's revised guidance on clinical decision support, issued in January 2026, places the same concern in its criteria: software that a clinician cannot independently review, or that supports a time-critical decision where the clinician is likely to rely on it without review, is a device software function.
Measuring the pair
For an assistive device, the evidence comes from a reader study. The FDA's guidance on computer-assisted detection in radiology, last revised in September 2022, describes the usual design: many readers each read many cases, with and without the device, and a "fully-crossed" design, in which every reader reads every case in both conditions, "offers the greatest statistical power for a given number of cases." Reading sessions are separated "by at least four weeks to avoid memory bias," the primary measure is usually the area under the ROC curve for readers with the aid against readers without it, and the number of readers must be justified rather than fixed by rule. Case sets may be enriched with diseased cases, the guidance says, but enrichment can change how readers behave and biases predictive values, the problem of Chapter 13 in another form.
The good machine-learning practice principles name the target directly. The 2021 version asked for focus on the performance of the human-AI team; the international version of 2025 asks that a device be assessed with a focus on human-AI interaction in the intended use environment. A model's standalone figures remain necessary, because they show what it does on its own, but for an assistive device they do not show what patients receive.
A strong model and a capable reader can combine badly if the reader follows the model's errors and overrides its correct calls. Only a study of readers working with the device, on cases that include the model's failures at a realistic rate, shows whether the pair beats the reader alone. Standalone accuracy cannot answer that question.
Telling users what they need to know
Clinicians can resist a wrong suggestion only if they know when the model is likely to be wrong. In June 2024 the FDA, Health Canada and the UK's MHRA published guiding principles on transparency for machine-learning devices, asking that makers share the device's intended use, performance, limitations and, where it can be done, the logic behind its outputs, with the right users at the right time and in a form designed for them. For IDx-DR, the labeling states the patients, the camera and the conditions under which performance was shown, and tells users that an "insufficient quality" result after dilation may itself be a sign of disease that needs referral. Chapter 17 reads that kind of information as an outside reader would, and asks what it leaves out.
Decide whether the device is autonomous, assistive or parallel, and design the display for that role: show the model's confidence or the evidence behind a suggestion where it helps review, and avoid placing a suggestion where it anchors the reader before an independent look. Measure the reader with and without the device in a study whose cases include the model's errors at a realistic rate, and train users on the failure modes the study finds.
- In a 2023 mammography study, wrong suggestions labeled as AI cut readers' accuracy from about 80 percent to between 20 and 46 percent.
- Autonomous, assistive and parallel triage devices place the model differently, and each role changes what has to be measured.
- Reader studies, ideally fully crossed with a washout of at least four weeks, measure the clinician with and without the device.
- Transparency about intended use, performance and limits gives users what they need to recognize a wrong output.
What a decision summary says, and leaves out.
+ The questionWhat should a reader look for in a software device's published performance, and what will it not tell them?
A public document with the evidence in it
The FDA publishes a decision summary for every device it grants through De Novo, and a shorter summary for most devices it clears through a 510(k). IDx-DR's is free to download. It gives the indications for use and the limitations, describes the software and its architecture, states the special controls that every later device of the same type must meet, summarizes the software documentation and the clinical study, lists the risks and their mitigations, and ends with the agency's judgment that the probable benefits outweigh the probable risks. For an outside reader, a hospital evaluating a product or an engineer studying a competitor, it is the most complete regulatory account of the evidence.
It also rewards careful reading. Its heading gives "DATE OF DE NOVO: January 12, 2018", which is the date the request was received; the FDA's database shows that the request was granted on 11 April 2018. Its table of results gives an observed sensitivity of 87.4 percent with a 95 percent confidence interval of 81.9 to 92.9 percent, but an exact interval computed from the trial's counts, 173 of 198, runs from 81.9 to 91.7 percent, which is what the device's 2022 510(k) summary prints. Such slips do not change the conclusion. They do show that a summary is a document written by people, to be read with a calculator at hand.
Rebuilding a table's percentages and intervals from its counts takes minutes and catches transcription errors, mismatched denominators and intervals of the wrong kind. When the counts are not given, that absence is itself worth noting, because a reader then cannot check how the cases the device declined were treated.
Questions to put to any claim
A performance claim can be read against a short list. Who was tested: the intended population, how participants were recruited, and whether the study was enriched. Where and by whom: the setting, the operators and their training. Against what: the reference standard and who applied it, masked or not. Which endpoints, and whether they were fixed before the study. How many: the counts behind each percentage, and the intervals. For which groups: subgroup results and the groups too small to assess. On what inputs: the camera or scanner, its software, the image format, the version of the device's own software. What happened to the cases it declined. And who ran the study: its funding, and whether the authors work for the maker.
IDx-DR's record answers most of these well. The population was adults with diabetes not previously diagnosed with retinopathy, partly enriched with participants whose diabetes was poorly controlled. The setting was ten primary-care practices with operators who had never imaged an eye, and the reference was a masked reading-center protocol. The endpoints were set in advance with input from the FDA, the counts are published, results by subgroup are reported, and the camera and software version are named. The paper also discloses that its first author was a shareholder, director and employee of IDx, the funder, which by the company's account he founded; another author held shares, and the statistician was paid by the company. What the record cannot say is how the device performs outside those conditions, which is why Chapter 14's youth studies and Chapter 20's monitoring matter.
Counting the cases it declined
The trial's headline sensitivity and specificity cover the 819 participants for whom both the reference standard and the program gave a usable result. Another 33 participants had gradable reference images but received "insufficient quality" from the program, and 10 of those 33 had the disease. How they are counted changes the result.
Leaving the 33 out, as the headline figures do, the program flagged 173 of 198 diseased participants: 87.4 percent. Counting the 10 diseased participants among the 33 as misses, it flagged 173 of 208: 83.2 percent, with an exact 95 percent interval of 77.4 to 88.0 percent, still above the trial's floor of 75 percent. Counting them as referrals, as the labeling directs for a persistent insufficient-quality result, the program sent 183 of 208 diseased participants onward: 88.0 percent. For the 23 declined participants without the disease, the same choice moves specificity from 89.5 percent to 86.3 percent if they are counted as false referrals. A full claim states which convention it uses and gives the others, as the immunoassay guide in this library recommends for a blood test with an indeterminate zone.
What a summary leaves out
Some things are absent by design. A decision summary rarely describes the training data, the model's structure or how its threshold was chosen; it reports the clinical evidence for the locked device, not how the device was built. It describes performance at authorization, not in use: IDx-DR's summary cannot report how the program has done in the clinics that bought it. And it describes the device as authorized, for the inputs named in it; it says nothing about cameras or populations outside the label.
Other gaps come from the difference between a regulator's words and a seller's. The FDA's press release of 2018 called IDx-DR the first device to give "a screening decision without the need for a clinician to also interpret the image or results"; the maker's materials speak of a diagnosis and of a device "De Novo-cleared", though a De Novo is granted, not cleared. A developer's own performance figure can differ widely from an independent one: Chapter 13's sepsis model was reported by its developer at an area under the curve of 0.76 to 0.83 and measured by an outside team at 0.63. The rule that follows is simple. A claim is read in the regulator's record and the peer-reviewed paper, with the maker's own statements labeled as such, and an independent evaluation is worth more than any of them.
Before buying, deploying or competing with a software device, fill in a single page: intended population, setting, operators, reference standard, prespecified endpoints, counts and intervals, subgroups, inputs and versions, the handling of declined cases, the study's funding and authors, and any independent evaluation. Mark each answer with its source, and treat any blank line as a question for the maker or a reason for a local test.
- A decision summary is the most complete regulatory account of a device's evidence, and it should be read with a calculator, as its dates and intervals show.
- A claim is read against who was tested, where, against what reference, with which endpoints, counts, subgroups, inputs and sponsors.
- How declined cases are counted changes the result: IDx-DR's sensitivity is 87.4, 83.2 or 88.0 percent depending on the convention.
- Summaries leave out training data and real-world performance, and a maker's wording and figures need labeling and independent checks.
+ Part V · Into the hospital, and after
Deploying, watching and changing a model.
On 10 June 2021 the FDA cleared a new version of IDx-DR that could receive images in DICOM, the format hospitals use for medical images, put its instructions on screen step by step, let clinics configure file names and keep images locally, and told operators during the exam whether an image was good enough to send. None of these changes touched what the program claimed to detect, and the clearance relied on no new clinical data. They were about fitting into clinics. A year later a second clearance replaced one of the models inside the analysis. The four chapters of this part follow a model from authorization into use: connecting it to a hospital's systems, checking it on the hospital's own patients, watching it once it runs, and changing it without losing the evidence. By the end of the part, the central question has its answer.
Fitting the software into a hospital.
+ The questionAn authorized AI device arrives at a hospital. What has to be connected, configured and agreed before it is used on the first patient?
The label on every image
Every image a hospital's scanners and cameras produce in the DICOM format carries a header: a list of tagged fields describing where the image came from. The field tagged 0008,0060 records the modality, such as CT or ophthalmic photography; 0008,0070 the manufacturer of the equipment; 0008,1090 the model name; 0018,1020 the software versions running on it; 0008,1030 a description of the study; and 0010,0020 the patient's identifier. The DICOM standard, revised several times a year, defines thousands of such fields; a 2026 release was current in October 2026.
For an AI device the header is the first line of defense. A device authorized for images from one camera can read the manufacturer and model fields and refuse anything else, and it can record the software version of the equipment that produced each image, so that a later change in performance can be traced to an update on the scanner. Many fields are filled in by local protocols or typed by staff, however, and the same examination can be described differently in two hospitals. A device that routes images by free-text descriptions will meet descriptions its makers never saw.
Where the software sits
An AI device rarely stands alone. An imaging model receives studies from the picture archiving and communication system, the PACS, after the scanner sends them there; it returns its results to the PACS, to the radiologist's worklist, to the electronic health record or to a specialist's phone; and it may run on a server in the hospital or in the maker's cloud. Each of those connections is an interface with its own standard. DICOM carries images and image-based results. HL7 version 2, which its standards body says is used by 95 percent of US healthcare organizations, carries orders and reports as messages. FHIR, HL7's newer standard built from web resources that can be addressed individually, has been at release 5 since March 2023; release 6 was in its second normative ballot in July 2026.
Standards for the AI step itself are younger. The IHE initiative, which writes profiles that tell vendors how to combine standards for a task, has published one for AI results, which defines how analysis results are encoded in DICOM objects and displayed, and one for AI workflow, which defines how a request for analysis is sent, managed and performed. In October 2026 both were still at trial implementation, the stage before final text. A hospital connecting several AI products therefore often builds part of each integration itself, or buys a platform that does it.
Inputs inside and outside the specification
The authorization defines the inputs. IDx-DR is indicated for use with the Topcon NW400, and its 2021 summary sets a minimum image resolution of 22 pixels per degree. Inside those limits the trial's evidence applies; outside them it does not, whatever the images look like. A device has two choices when an input falls outside its specification: refuse it, with a message the user can act on, or analyze it anyway and risk an answer the evidence never covered. IDx-DR refuses, with its "insufficient quality" output and, since 2021, with feedback during the exam so that the operator can retake the image.
Inputs also drift after installation. A hospital changes a scanner protocol to reduce dose, a vendor updates the scanner's reconstruction software, the PACS begins compressing images to save storage, or a new interface converts images to a different format on the way to the model. Each change can be invisible to a person reading the images and plain to a model, the acquisition shift of Chapter 14. Chapter 19 describes a hospital study in which a change of image format, more than any change in the patients, took a model from excellent to useless.
A new scanner protocol, a reconstruction update, compression in the archive or a format conversion in an interface engine can move a model's inputs outside the conditions it was tested under, without any change to the device itself. Agreeing in advance who notifies whom of such changes is part of installing the device.
The stroke-triage program authorized in 2018 was compared with usual care, retrospectively, in 44 cases with a confirmed blocked artery that it had flagged correctly: the specialist was notified a median of 5.6 minutes after the CT angiogram, against 51.5 minutes when notification came through the radiology report. In a 2023 randomized trial of the same company's software at four Houston stroke centers, involving 243 treated patients, the time from arrival to the start of clot removal fell by an estimated 11.2 minutes from a median of 100, about 11 percent. The two studies measured different intervals in different hospitals, so the numbers cannot be subtracted, but together they show the shape of the problem: the model shortens one link, and transfer of the images, the alert reaching the right person, the team assembling and the procedure room being ready decide how much of the gain reaches the patient. The trial found no significant difference in patients' functional outcome, an exploratory measure.
Who answers for what
The maker answers for the device within its specified environment, and the EU's medical device regulation makes it state that environment: manufacturers "shall set out minimum requirements concerning hardware, IT networks characteristics and IT security measures" needed to run the software as intended. The hospital answers for the network, the interfaces, the other systems and the way the device is used. Between them sits the integration, which neither controls alone.
The standard for that middle ground is IEC 80001-1, whose 2021 edition sets requirements for organizations applying risk management before, during and after connecting a medical device or health software to their IT infrastructure, covering safety, effectiveness and security. It is already marked for revision, and ISO intends to replace it with a new standard in the 81001 series. An earlier technical report in the same family describes responsibility agreements: written statements of which party, the hospital, the IT supplier or the device maker, does what across the life of the connection. The questions such an agreement must settle are concrete: who tests the interfaces, who approves changes on each side, who is told when a scanner or the archive is updated, and what clinicians do when the AI is unavailable.
Before the first patient, list every interface the device depends on and test each with real local data, and confirm that the local images and their headers fall inside the device's specification. Measure the time from acquisition to result against the clinical need. Define what clinicians do when the device is down or declines an input, and sign a responsibility agreement that names who notifies whom of changes on either side.
- DICOM headers identify the equipment and software behind every image, letting a device refuse inputs outside its specification and trace changes.
- AI devices depend on interfaces to PACS, health records and cloud services; the IHE profiles for AI were still trial implementations in October 2026.
- A model's evidence covers only inputs inside its specification, and hospital-side changes can move inputs outside it without touching the device.
- The maker states the required environment, the hospital manages the connection under IEC 80001-1, and a responsibility agreement divides the rest.
Site acceptance, local performance and computerized system validation.
+ The questionThe maker validated the software; why does the site have to validate it again, and how much is enough?
Two validations, two questions
A device that arrives at a hospital has been validated once already, by its maker, and validated again would seem redundant. It is not, because the two validations answer different questions. The maker's validation shows that the device meets its intended use in its specified environment, on the patients and inputs of its studies. The hospital's question is whether that evidence carries over to this hospital: its scanners and their settings, its interfaces, its patients and its way of working. Only the hospital has the data to answer it.
Professional societies now say so plainly. In a joint statement published in January 2024, the radiology societies of the United States, Canada, Europe, Australia and New Zealand called testing on local data, with local systems and workflows, "essential", and recommended that each site "perform a statistically rigorous evaluation of performance on their own local data." Where that is not feasible, they advised comparing local data with the vendor's test data and, where the two differ, proceeding "with great caution, if at all."
A trial nobody saw
In August 2020 a team at the Hospital for Sick Children in Toronto began running a deep-learning model on live clinical data with its outputs hidden from clinicians, an approach the team calls a silent trial. The model predicted, from renal ultrasound images, which children with swelling of the kidney would need surgery. On a random 20 percent held out from its development data, 1,643 kidneys from 294 patients in all, its area under the curve had been 0.90. In the silent trial, on 523 kidneys from 150 patients seen between August and December 2020, it was 0.50: no better than chance.
The team traced the collapse to three differences. The live patients were younger, the share of obstructed right kidneys was higher, and the images reached the model in a different form: the development data had been processed JPEG files, while the live data arrived as unprocessed PNG files that looked different to the model even after the same preprocessing. Adjusting for age and side barely helped, raising the area to 0.51. Reprocessing the live images to match the original pipeline raised it to 0.84 to 0.85. After retraining on the original and silent-trial data together, a second silent trial on 711 kidneys from 202 patients gave 0.91 to 0.92. The decisive fault was not clinical at all. It lay in the integration chain of Chapter 18, and only running the model on the hospital's own data, before anyone acted on it, revealed it.
Acceptance and a local check
A site's own validation has two layers. The first is acceptance: confirming that the device is installed as specified, that each interface passes the right data, that results reach the right screen for the right patient, and that the fallback works when the device is unavailable. The second is a local performance check: running the device on a sample of the site's own cases, silently or retrospectively, and comparing its outputs with a local reference. The multisociety statement suggests focusing that review on the predictive values clinicians will experience, identifying the cases where the device would change care, and categorizing its false positives and negatives before deciding whether to deploy. It recommends repeating the review whenever the AI software or the equipment used with it changes.
A hospital wants evidence that a device's sensitivity on its patients is above 75 percent, and expects it to be about 87 percent, as in the device's studies. The standard calculation for comparing a proportion with a fixed value, at one-sided 5 percent significance and 80 percent power, gives about 69 diseased cases. If 10 percent of the hospital's screened patients have the disease, that means about 690 consecutive cases. With 60 detections out of 69, 87 percent, the exact one-sided lower confidence bound is about 78 percent, above the target. A check run on a few dozen convenient cases cannot reach that conclusion, which is why the societies call for statistical rigor rather than a demonstration.
The results of the check become the baseline for the monitoring of Chapter 20: the distribution of inputs, the rate of declined cases, and the local sensitivity, specificity and predictive values against which later drift is measured.
Computerized system validation
Some sites must go further, because the rules they work under require it. Laboratories, blood services, organizations running clinical trials and pharmaceutical manufacturers operate under good-practice rules, collectively called GxP, that require any computerized system affecting product quality, patient safety or data integrity to be validated by its user. The method most of them follow is ISPE's GAMP 5, whose second edition was published in July 2022. It is risk-based: the effort scales with the system's impact, complexity and novelty, and its authors describe it as placing patient safety and product quality ahead of compliance for its own sake and supporting iterative and agile methods.
GAMP sorts software into categories: infrastructure software such as operating systems and databases, and then a continuum from products used without configuration through configured products to custom software. The more a system is configured or customized for the site, the more the site must verify itself; for a standard product, the user leans on the supplier's evidence after assessing the supplier. The work follows a chain. User requirements say what the site needs, and a risk assessment shows where failure would matter. The system is then specified and configured, and verified, traditionally as installation, operational and performance qualification and increasingly as risk-based scripted and unscripted testing. A traceability matrix links requirements to tests, release is controlled, and the system is reviewed periodically for as long as it is used. Data integrity runs through all of it: records that are attributable, complete and protected, with audit trails, as the FDA's rule on electronic records and signatures, 21 CFR Part 11, has required since 1997.
Supplier documentation lets a site avoid repeating tests that do not depend on its environment. It cannot cover the site's interfaces, configuration, data or patients, and a package accepted without assessing the supplier, or without local performance data, leaves the site's own question unanswered.
AI has added to the method. The 2022 edition of GAMP 5 included an appendix on machine learning with a life cycle for an ML subsystem, and in July 2025 ISPE published a separate GAMP guide to AI, covering data governance, model testing, monitoring and change management for AI-enabled systems used in GxP work. In July 2025 the European Commission also put out for consultation a new annex on AI to the EU's good manufacturing practice rules. It would admit only static, deterministic models in critical manufacturing uses, would not apply to generative AI and large language models, and would require acceptance criteria at least as high as the performance of the process the model replaces. As of October 2026 the annex was still a draft. It governs medicines manufacturing, not medical devices, but its test-data rules read like a summary of Chapters 11 to 13.
The builder's standard and the user's
IEC 62304 and GAMP 5 describe the same software from opposite sides. IEC 62304 is written for the organization that builds device software, and its records show how the device was made. GAMP 5 is written for the organization that uses a computerized system, and its records show that the system does what that user needs, in that user's configuration. A laboratory working under good-practice rules and installing an AI-enabled analyzer, or a contract research organization using an AI reading tool in a trial, needs both: the maker's evidence, and its own validation built on it. Makers can shorten the second by supplying a validation package with their product: requirements and traceability, test scripts the site can re-run, and installation and operational qualification documents.
Write the site's requirements and acceptance criteria, including the local performance check and its sample size, before purchase, and ask the supplier for its validation package and its specified environment. Where GxP rules apply, scope the work with GAMP 5's categories and risk assessment; in every case, run the device silently or retrospectively on local cases before clinicians rely on it, and keep the results as the monitoring baseline.
- The maker's validation covers the intended use in a specified environment; only the site can show that the evidence carries over to its systems and patients.
- A silent trial in Toronto took a model from 0.90 to 0.50, to 0.85 after fixing an image-format mismatch in the integration and to 0.92 after retraining.
- A local check needs enough diseased cases to support a conclusion, about 69 to show a sensitivity above 75 percent, and becomes the monitoring baseline.
- Where GxP rules apply, sites validate AI systems with GAMP 5's risk-based method, now extended to AI, building on the supplier's evidence.
When the true answer arrives late.
+ The questionHow can a maker, or a hospital, tell a model is still working when the right answer arrives months later, or never?
Truth arrives late
A model can be wrong for months before anyone could know. When a screening program tells a patient to come back in a year, nobody checks whether that answer was right until the patient returns, if they do. When it refers a patient, the specialist's examination confirms or overturns the answer weeks later, in another clinic's records. For most outputs of most diagnostic models, the reference standard of Chapter 12 is never applied at all in routine use. The direct measure of performance, agreement with the right answer, is available late, partially and only for some patients.
The early months of the COVID-19 pandemic showed how fast the ground can move meanwhile. A study of 24 hospitals in four US health systems found that daily alerts from a widely used sepsis prediction model rose by 43 percent while the hospitals' patient census fell by 35 percent, ahead of the surge in COVID admissions. The patients had changed, and the model's inputs with them. At Michigan Medicine, one of the four systems, the flood of alerts led to the alerts being paused, according to the university's account of the study. Nobody needed to wait for confirmed diagnoses to see that something had shifted: the volume of the model's own output showed it.
What can be watched at once
Monitoring therefore starts with what is available immediately. The inputs can be checked as they arrive: the equipment and software versions in the image headers of Chapter 18, image quality measures, patient age and other characteristics the device records. The outputs can be counted: the share of positive results, the share of declined cases, the distribution of scores. None of these says directly whether the model is right. Each can show, within days, that the conditions it was validated under no longer hold, which is the trigger to look harder.
Suppose a screening device with 87.4 percent sensitivity and 89.5 percent specificity runs at a site where 10 percent of patients have the disease. It should return a positive result for about 18.2 percent of patients: 8.7 points from the diseased and 9.5 from false alarms among the healthy. At 200 screens a week, that is about 36 positives a week, with a typical week-to-week spread of about 5. If a camera fault lowers specificity to 80 percent, the expected positive rate rises to about 26.7 percent, about 53 positives a week, more than three spreads above the baseline. A simple chart of weekly positives would flag the change within a few weeks, months before referral outcomes could confirm that the extra positives were false.
Some drift hides from counts. A 2017 study at US Veterans Affairs hospitals followed seven models predicting acute kidney injury, regression and machine-learning models alike, for nine years after their development data. All kept their ability to rank patients by risk; their discrimination was maintained. Their calibration drifted: every model came to overstate the risk as the rate of kidney injury changed, and the overprediction grew steadily for the regression models while staying small for the random forest and the neural network. A model used with a fixed threshold on predicted risk would have alerted more and more often for the same patients, with a stable area under the curve throughout.
Discrimination measures whether a model ranks sick patients above healthy ones; calibration measures whether its scores mean what they say. Monitoring that tracks only the area under the curve can miss a drift that moves every decision made at a fixed threshold, which is the drift that changes alert rates and referrals.
Closing the loop with outcomes
The direct measures follow. Referred patients' specialist findings can be collected and matched to the device's outputs, giving a running estimate of the positive predictive value. A random sample of negative results can be sent for a reference reading, the only way to estimate missed cases before patients return with them. Complaints, adverse event reports and service records feed the maker's postmarket surveillance. The multisociety radiology statement of 2024 recommends re-evaluating each AI tool on updated local data "at specified intervals, but at least annually," and after any new version, noting that "naturally occurring data drift will cause AI model performance to degrade over time," and it asks institutions to define in advance how a problem will be escalated and resolved.
Who watches
Responsibility is shared and not yet settled. The maker has postmarket obligations everywhere a device is sold, and the good machine-learning practice principles end with the expectation that deployed models are monitored for performance and that the risks of retraining are managed. The hospital sees the inputs, the workflow and the outcomes, and only it can link the device's outputs to its own patients' later records. In September 2025 the FDA opened a public discussion on how the real-world performance of AI-enabled devices should be measured and evaluated, asking about metrics, monitoring methods, data sources and the triggers for a response; it described the document as proposing no policy, and no follow-up had been published by October 2026. The EU's AI Act will add monitoring duties for hospitals that deploy high-risk AI, on the timetable set out in Chapter 29.
What a monitoring plan needs is clear even where the division of labor is not: a baseline from the site validation of Chapter 19, measures of inputs and outputs watched continuously, a reference sample and outcome matching on a schedule, thresholds that trigger investigation, and named people on each side who act on them. When monitoring shows that performance has moved, the response is a change to the device, its labeling or its use, and Chapter 21 sets out how a model can be changed without losing the evidence behind it.
Define the input checks, the output and declined-case rates to chart, their baseline from the local validation and the limits that trigger investigation. Schedule a reference reading of a random sample of negatives and the matching of positives to specialist findings, assign who on the maker's and the site's side reviews each, and state what happens when a limit is crossed: investigation, restricted use, suspension or retraining.
- For most outputs of a diagnostic model the right answer arrives late or never, so accuracy cannot be monitored directly in real time.
- Inputs, output rates and declined-case rates show within days when conditions have shifted, as sepsis alerts did early in the pandemic.
- Calibration can drift while discrimination holds, so monitoring must watch decisions at the threshold, not only the area under the curve.
- Reference reads of sampled negatives, outcome matching and agreed triggers close the loop; maker and site share the work.
Agreeing in advance what may change.
+ The questionHow can a model be retrained after authorization without a new review every time, and what keeps the retrained model as good as the one that was authorized?
From a discussion paper to a final guidance
On 2 April 2019 the FDA published a discussion paper, explicitly not a draft guidance, proposing a new way to regulate changes to machine-learning software. Its premise was that such software improves by retraining, and that requiring a new premarket review for every improvement would either freeze models at their first version or require a new submission for each improvement. Its proposal was that a maker could describe in advance a "region of potential changes", which it called the pre-specifications, together with an algorithm change protocol: the methods it would use to make, test and control those changes. If the plan was authorized with the device, changes inside it could be made without a new submission.
The idea took five years to become final guidance. The FDA, Health Canada and the UK's MHRA published five guiding principles in October 2023: a plan should be focused and bounded, risk-based, evidence-based and transparent, and should take a total product lifecycle perspective. The FDA's final guidance on predetermined change control plans for AI-enabled device software followed on 4 December 2024 and was reissued on 18 August 2025. A broader draft guidance on such plans for all devices, issued in August 2024, was still a draft in October 2026.
What a plan contains
The final guidance asks for three things. A description of modifications specifies the planned changes and the characteristics and performance the modified device will have: for example, retraining on new data from the same kinds of sites, or adjusting a threshold within stated limits. A modification protocol describes how each change will be developed, validated and implemented: how new data will be collected and kept separate, how the model will be retrained, how its performance will be evaluated, the acceptance criteria it must meet, and how users will be told. An impact assessment documents the benefits and risks of carrying out the plan, and how those risks will be controlled.
The plan has firm limits. Modifications must stay within the device's intended use, and generally within its indications; a change of what the device is for needs a new submission. Changing the plan itself generally needs a new submission too. The device's labeling should say that it contains machine learning and has an authorized plan, and the public summary of the authorization should describe the plan, so that users and other makers can see what may change without further review. The plan can be authorized through a 510(k), a De Novo or a premarket approval.
A plan that lists permitted changes but sets loose or self-adjusting acceptance criteria gives the regulator nothing to rely on. The criteria that matter, in this guide's view, are fixed before any modification is made: performance on a sequestered test set at least as good as the authorized version, overall and in each named subgroup, with the same reference standard. A modified model that misses them does not ship.
The thread's change
IDx-DR's own record shows a change made before the guidance existed, and what evidence it used. The 2018 De Novo already required, as a special control, "a protocol... that describes the level of change in device technical specifications that could significantly affect the safety or effectiveness of the device," an early, device-specific form of the same idea. In June 2022 the FDA cleared version 2.3, whose changes included a new classifier for judging image quality inside the analysis. To show the change did no harm, the maker re-ran the new version on the images from the 2017 trial, 850 of which could be evaluated, and compared its answers with the original version's.
Against the reading-center reference, version 2.3 flagged 171 of 195 diseased participants it analyzed, 87.7 percent, against version 2.0's 173 of 198, 87.4 percent; it cleared 553 of 614 participants without disease, 90.1 percent, against 556 of 621, 89.5 percent. On the participants each version analyzed, both figures rose slightly. Counted over the same 850 participants, however, version 2.3 detected 171 diseased participants against 173 and cleared 553 without disease against 556: the rise came from declining more people. The new quality classifier changed which participants received an answer: after all resubmissions, version 2.3 gave a usable result for 809 of 850 participants, 95.2 percent, against version 2.0's 819, 96.4 percent; the 96.1 percent of the opening scene counted 852 participants with a reference grade. Ten more participants out of 850 were declined, 1.2 points. In a program screening 10,000 people a year, that would mean, at the trial's case mix, about 120 more patients sent for re-imaging or referral for image quality alone. The comparison was a regression test on the original patients; it showed that the new version did not do worse on them, and it could say nothing about patients the trial never included.
Plans in use
The plans the FDA has authorized are now numerous enough to describe. One of the earliest was in the De Novo authorization of Caption Guidance, a program that guides ultrasound users to acquire cardiac images, in February 2020; when the plan was modified by a 510(k) that September, its summary stated that all algorithm modifications would be "trained, tuned, and locked prior to release", and continuously learning algorithms were excluded. A peer-reviewed study published in July 2026, using FDA data to April 2026, counted 170 devices across all review panels cleared with a change control plan, 34 of them radiology AI devices, 22 of those cleared in 2025; it found that the public summaries described neither ongoing performance monitoring nor predefined triggers for retraining. The insulin-dosing app UpDoc, cleared in December 2025, lists five categories of permitted change, from new insulin formulations to alternative ways of entering data, none of which alters the cleared dosing logic.
A plan changes what the maker may do without a new review. It does not change what a hospital needs to know. Each new version alters the device the site validated in Chapter 19, and a site that relies on its baseline needs to be told what changed, to re-run its local check when the change could move performance, and to reset its monitoring limits from Chapter 20.
The answer so far
Parts II to V together answer the central question. For conventional software, the evidence that it will be right for the next patient is a lifecycle: requirements derived from risk and traced to tests, an architecture that confines the dangerous code, an account of every borrowed component, verification at each level and regression after every change. For an AI model, it adds a test kept apart from training and drawn from the setting of use, against a reference fit for the claim, reported with its uncertainty and by subgroup, then checked at each site and watched in use. What keeps that evidence true after an update is change control: every change judged against the claims it could move, and, for a model, a plan agreed in advance that bounds the changes and fixes the tests each new version must pass. Part VI adds the threat that does not wait for an update.
Decide before the first submission which changes the model is likely to need: retraining on new data, new input devices, threshold adjustments. Describe them, the data they will use, the sequestered test sets and the acceptance criteria overall and by subgroup, and how users and sites will be told of each new version. Keep the plan narrow enough that each permitted change can be tested against criteria fixed today.
- The FDA's 2019 discussion paper led to final guidance in December 2024: a predetermined change control plan lets authorized changes proceed without a new submission.
- A plan holds a description of modifications, a modification protocol with fixed acceptance criteria, and an impact assessment, all within the intended use.
- IDx-DR's 2022 version re-ran the 2017 images: accuracy on analyzed participants rose slightly because 1.2 points more of them received no answer.
- Each new version changes the device a site validated, so sites must be told, re-check performance where needed and reset their monitoring.
+ Part VI · Defending the device
Cybersecurity.
On 28 August 2017 Abbott wrote to physicians that a firmware update for several of its pacemakers, formerly made by St. Jude Medical, was ready to close vulnerabilities in the devices' radio link. The update could not be sent remotely. Each patient had to visit a clinic, where the update took about three minutes through the programmer's wand while the pacemaker ran in a backup mode; press reports of the FDA's communication put the number of affected devices in the US at about 465,000. Nothing in those pacemakers had failed. According to a cardiology society's summary of the FDA's communication, the flaws could allow intrusions, some of which could affect how the device operated. The three chapters of this part treat that kind of failure: what threats a device faces, which defenses must be designed in before release, and how weaknesses found after release are judged and fixed.
An attacker is an input.
+ The questionHow can a device that meets every requirement still be made to harm a patient?
A pump that could not be patched
On 27 June 2019 the FDA warned patients and clinicians that certain Medtronic MiniMed insulin pumps were being recalled because an unauthorized person "could potentially connect wirelessly to a nearby MiniMed insulin pump and change the pump's settings." The pumps, the MiniMed 508 and several models of the Paradigm series, communicated by radio with remote controllers and glucose meters, and the protocol they used lacked proper authentication: a pump could not tell its own patient's devices from an attacker's transmitter within range. Medtronic had identified about 4,000 patients in the US who were potentially using vulnerable pumps.
The remedy was unusual. The FDA's announcement said Medtronic was "unable to adequately update" the pumps "with any software or patch," and the company offered patients replacement pumps with stronger built-in protection. The FDA said it was not aware of any confirmed patient harm. The US cybersecurity agency's advisory on the same day described the weakness as improper access control, exploitable only by an attacker nearby and requiring high skill, with no known public exploit. The recall record for the MiniMed 508, posted in March 2020, was still open in October 2026.
Safety risk and security risk
The pumps met their requirements. They delivered insulin as programmed, alarmed as specified and passed their verification. What they lacked was a requirement nobody had written in the form that mattered: refuse commands from anyone but the patient's own devices. The fault was built in on the day they shipped, like every software fault in this guide, and the input that reached it was not an unlucky combination of events but a deliberate one.
That difference changes how risk is judged. Safety risk management, Chapter 4's ISO 14971 chain, estimates how likely a sequence of events is. An attacker is not a sequence of events with a probability; an attacker chooses the inputs, searches for the rare combination and repeats it at will. The FDA's 2026 guidance on cybersecurity in premarket submissions therefore describes security risk assessment as non-probabilistic: it judges a vulnerability by its exploitability, how easily an adversary could use it, and by the harm that would follow, and it treats known vulnerabilities as reasonably foreseeable. Where a security weakness could lead to patient harm, the two assessments meet, and a security control becomes a safety control.
The 2019 advisory scored the pump weakness 7.1, High, on version 3 of the Common Vulnerability Scoring System, with the vector AV:A/AC:H/PR:N/UI:N/S:U/C:L/I:H/A:H. Read left to right: the attack vector is adjacent, meaning within radio range rather than over the internet; attack complexity is high; no privileges and no user interaction are needed; the scope is unchanged; the impact on confidentiality is low, but on integrity and availability high, because settings and delivery could be changed or stopped. A 2018 advisory on the same pumps' remote controllers had scored two weaknesses 4.8 and 5.3, Medium, because each affected only one of confidentiality or integrity, and one also needed user interaction. The score summarizes technical severity. It says nothing about how many patients use the device or what a wrong dose of insulin does, which is why a maker's assessment adds the clinical harm on top.
The attack surface
Every way data or commands can enter a device is part of its attack surface: radio and Bluetooth links, Wi-Fi and Ethernet, USB and serial ports, the connection to a cloud service, a service interface used only by technicians. The law that has governed connected devices in the US since March 2023 applies to a "cyber device", which, among other conditions, includes software and can connect to the internet. The FDA reads that broadly, "whether intentionally or unintentionally, through any means": Wi-Fi, cellular, Bluetooth, magnetic inductive links, USB, Ethernet and serial ports all count, and so does a USB connection used briefly for service.
A threat model lists those entry points, the assets behind them, such as the dose settings, the patient's data, the device's software itself, and the adversaries who might want them, and asks for each path what an attacker could do and what would stop them. MITRE and the Medical Device Innovation Consortium, with FDA funding, published a playbook for threat modeling medical devices in 2021, and its guidance expects the threat model to be part of the design, updated as the architecture of Chapter 5 changes. IDx-DR's 2018 special controls already required a cybersecurity vulnerability and management process, because its images and answers cross the internet between the client and the server.
Estimates that an attack is unlikely rest on assumptions about who would try, with what skill and for what gain, and those assumptions age. Research, published exploits and automated tools lower the skill an attack needs. Judging a weakness by how easily it could be used, and by the harm that would follow, ages better than judging it by how likely an attack seems today.
When security becomes safety
Some attacks threaten patients directly, as a changed insulin dose would. Others threaten data, and through it trust and care. In January 2025 the FDA warned that the software of a patient monitor sold as the Contec CMS8000, and under another name, contained a backdoor and could send patient data outside the care setting. It advised hospitals that relied on the monitors' remote monitoring to unplug and stop using them, and others to use them for local monitoring only. The fix released in July 2025 removed the monitors' networking entirely: the cure for an unsafe connection was to have none.
The pump and the monitor show the two hard truths of device security. A defense that was not designed in may be impossible to add later, and the safest response to some weaknesses is to give up a function. Chapter 23 describes the defenses that have to be in the design from the start.
Start from the architecture diagram of Chapter 5 and mark every interface, including service and debug ports and cloud connections. For each, list what could be read, changed or stopped through it, by whom and with what effort, and the control that prevents it. Judge each path by exploitability and clinical harm rather than by an estimate of how likely an attack is, and update the model with every design change.
- MiniMed pumps recalled in 2019 accepted radio commands without proper authentication and could not be patched, so patients were offered replacements.
- Security risk is judged by exploitability and harm, not probability, because an attacker chooses the inputs; known vulnerabilities count as foreseeable.
- Every interface is part of the attack surface; US law and the FDA's broad reading make almost any connectable device a cyber device.
- A threat model built from the architecture names each path, its harm and its control, and some weaknesses can be removed only with the function.
Built to be defended.
+ The questionWhich defenses have to be designed in before release, because they cannot be added later?
A signature checked at every start
A firmware image is one file: the complete program a device will run, compiled, tested and released under a single version number. Before it leaves the maker, a release system computes a fingerprint of that file, a short number called a hash that changes completely if a single bit of the file changes, and signs the fingerprint with a private key kept on a guarded machine. The signature travels with the file. Inside the device, in memory that cannot be rewritten, sit the matching public key and a small piece of start-up code. Every time the device powers on, that code computes the fingerprint again, checks it against the signature, and runs the firmware only if the two agree.
This chain is called secure boot, and the protected code and key at its start are the root of trust. Each stage checks the next before handing over control, so a program that someone has altered, however slightly, is refused before it runs. NIST's guidelines on firmware resiliency, published in May 2018 for computing platforms generally, describe three duties that such a design serves: protecting the firmware against unauthorized changes, detecting changes that occur anyway, and "recovering from attacks rapidly and securely," usually by returning to a copy known to be good.
The design depends on something present in the hardware from the first unit shipped: protected memory for the key and the start-up code. A device built without it cannot acquire it through an update, because the update itself would have to be trusted by code that cannot tell a genuine update from a forged one.
Why a checksum is not a lock
Embedded software has long carried checksums, most often a cyclic redundancy check, or CRC, which detects a file corrupted by a faulty memory chip or a noisy cable. The FDA's premarket cybersecurity guidance, revised on 3 February 2026, says plainly that such checks "do not provide integrity or authentication protections in a security environment," and tells makers not to rely on them as security controls. The reason is simple. A CRC is computed by a published method from the file alone, so anyone who changes the file can compute a new one that matches. A signature can be produced only with the private key, so a matching signature proves both that the file is unchanged and that it came from the holder of the key.
The same distinction runs through every defense in this chapter. A safety mechanism protects against chance: a corrupted bit, a failed sensor, an unlikely input. A security control must hold against someone who knows how the mechanism works and is trying to defeat it, the adversary of Chapter 22. Many of the controls look alike from outside. Their design differs in what they assume about the person on the other side.
A lifecycle with security in it
The standard for building security into health software is IEC 81001-5-1, published in December 2021. It was written to sit on top of a device maker's existing process rather than beside it: its scope is based on IEC 62304, its clauses follow IEC 62304's order, and it takes its security activities from IEC 62443-4-1, the equivalent standard for industrial control products. Where IEC 62304 has a software risk management process, it has a security risk management process, and it adds security tasks such as threat modeling, security requirements, security testing and the handling of vulnerabilities after release. A team already following IEC 62304 extends its procedures rather than writing a second set. The standard has been flagged for revision since September 2025.
The FDA calls such a process a secure product development framework, and its 2026 guidance names IEC 81001-5-1 and ANSI/ISA 62443-4-1, the US adoption of IEC 62443-4-1, as examples that can satisfy it. For the risk assessment itself it points to two documents from AAMI: a technical report, TIR57, published in 2016, and the standard ANSI/AAMI SW96, published in 2023 and aligned with the ISO 14971 process of Chapter 4.
Eight kinds of control
The guidance's first appendix sorts the defenses a device may need into eight categories: authentication, authorization, cryptography, the integrity of code, data and execution, confidentiality, event detection and logging, resiliency and recovery, and firmware and software updates. The list reads like a summary of what the MiniMed pumps lacked. Their radio protocol did not authenticate the sender, so a pump could not tell its patient's remote from an attacker's transmitter. A 2018 advisory on the pumps' remote controllers had also described a replay attack, in which recorded traffic could be sent again to trigger a bolus when the remote options were enabled.
The guidance answers each of those faults in a sentence. It asks makers to "use cryptographically strong authentication, where the authentication functionality resides on the device," and not to use passwords "that are hardcoded, default, easily guessed, or easily compromised." It asks for anti-replay measures, such as a number used only once in each exchange, for commands that could do harm. It asks that updates be authenticated before they are installed, and that hardware-based security be used where possible. The architecture the maker submits is expected to show the device in four views: the whole system with its connections, the paths by which one compromise could harm many patients, how updates travel from the maker to every unit, and the security of each use case. The testing it lists includes penetration testing, fuzz testing, tests with malformed and unexpected inputs, analysis of the attack surface and of how weaknesses can be chained, and software composition analysis of the compiled binaries.
Protected memory for keys, a processor able to check signatures quickly, room for a second firmware image and a radio protocol with authentication are decisions taken early in design. A device shipped without them can sometimes be protected by its surroundings, never fully repaired, and the MiniMed recall shows that the remedy may be a new device.
Designing for the patch
A device that can be updated has a security control the MiniMed pumps lacked: a way to change its own code after release. The Abbott pacemakers of 2017 had one, though only through a programmer's wand in a clinic, and Abbott's letter to physicians shows that the update path itself carries risk. It recommended against replacing the devices to avoid the vulnerability, and asked physicians to weigh the update's risks for each patient, with particular care for patients who depend on their pacemaker for every heartbeat.
Abbott's letter gave three failure rates from its earlier firmware updates: the device reloading its previous firmware after an incomplete update, 0.161 percent; loss of the programmed settings, 0.023 percent; complete loss of device function, 0.003 percent. Per 100,000 updates, that is about 161 devices left on the old version but working, 23 needing to be reprogrammed, and 3 that stop working. For a patient whose heart does not beat without pacing, the last outcome is the one that matters, which is why the letter advised considering the update for such patients where temporary pacing was available. The most common failure was also the safest, because the device kept its previous firmware and returned to it when the new one did not load completely. That fallback is a design decision, and it turned most failed updates into a device still working on its old version.
The example shows the shape of the decision a maker faces after release. A patch closes a vulnerability that, for these pacemakers, had produced no reports of a compromised device; it opens a small, measurable risk of its own. Good design makes the trade easier: an update path that authenticates the update, writes it to a spare slot, checks it, and keeps the old version until the new one has started correctly. Chapter 24 follows what happens when the vulnerability arrives after release, in code the maker did not write.
Decide at the start which hardware the device needs for security: protected key storage, signature checking at start-up, room for a fallback image. Add authentication, anti-replay and authorization to every interface in the threat model, plan how updates will reach every unit and how a failed one recovers, and list the security tests, from fuzzing to penetration testing, in the verification plan.
- Secure boot checks a signed firmware image against a key held in protected hardware at every start, so altered code never runs.
- A CRC detects accidents but not attackers; only a signature made with a private key proves a file is genuine.
- IEC 81001-5-1 adds security risk management and security tasks to an IEC 62304 process, and the FDA accepts it as a development framework.
- Controls such as authentication and an update path with a fallback must be designed in; a device built without them may have to be replaced.
The threat changes while the code does not.
+ The questionOnce a device ships, how are new weaknesses in its code and its components found, judged and fixed?
Eleven flaws in someone else's code
Eleven vulnerabilities, six of them serious enough to let an attacker on the network run code of their own choosing, were disclosed in July 2019 in IPnet, a network stack: the part of an operating system that sends and receives data over a network. The security company that found them, Armis, named them URGENT/11 and reported that the affected code had been part of the VxWorks real-time operating system since version 6.5, and of several other operating systems that had used the same stack. On 1 October 2019 the FDA warned patients, clinicians and makers that medical devices and hospital networks using those operating systems could be affected. Its announcement named six operating systems from six suppliers, said that device makers' notices so far included an imaging system, an infusion pump and an anesthesia machine, and said the agency knew of no adverse events.
No device maker had written IPnet. It reached their products inside operating systems they had chosen and, under Chapter 6's rules, treated as SOUP. The code in those devices had not changed since release. What changed was public knowledge about it, and with it the risk. This is the general pattern of security after release: the threat changes while the code does not, and most of the code at issue came from someone else.
Matching the list to the news
The first task when a vulnerability is published is finding out which products contain the affected component, and in which version. This is the work the software bill of materials of Chapter 6 exists for. A maker that keeps a machine-readable SBOM for every released version can search all of them for the component within hours; a maker that must ask each development team, or inspect old build files, may take weeks. The FDA's 2026 premarket guidance expects makers to monitor sources such as NIST's National Vulnerability Database, to assess the known vulnerabilities in each component, and to give particular weight to those in the US cybersecurity agency's catalog of vulnerabilities already being exploited in real attacks.
Presence is only the first question. A device may contain the vulnerable component without using the vulnerable function, or without exposing it to any interface an attacker can reach. The answer for each vulnerability and each product is recorded as a VEX statement, a Vulnerability Exploitability eXchange, which the US Commerce Department's telecommunications agency, NTIA, described in 2021 as an assertion about the status of a vulnerability in a specific product, with four values: not affected, affected, fixed, or under investigation.
In June 2020, after vulnerabilities in another widely used network stack, made by Treck, were disclosed, the infusion pump maker B. Braun had received 24 patches from Treck for its Outlook 400ES pump and had determined that 20 did not apply, according to a trade report. Four of 24 is 17 percent. Each of the 20 needed a reasoned statement of why the pump was not affected; each of the 4 needed an assessment of exploitability and harm, and a fix. The scale of that work grows with the bill of materials. If a product carries the mean of 1,180 open-source components that a code-audit vendor reported (Chapter 6), and 1 percent of them receive a new advisory in a year, about 12 advisories arrive for that product each year; at B. Braun's ratio, about 2 would need a fix, and every one would need a written conclusion. Without an SBOM, the first step of each, knowing whether the component is present at all, is the slowest.
Judging what a weakness means for patients
A vulnerability that applies is judged on two scales. The first is technical severity, usually expressed with the Common Vulnerability Scoring System that Chapter 22 read for the MiniMed pumps. Its fourth version, published on 1 November 2023, added a supplemental Safety metric that records whether exploitation could cause injury, using the consequence categories of the functional safety standard IEC 61508. The specification is explicit that supplemental metrics do not change the score. A rubric for applying the scoring system to medical devices, written by MITRE under an FDA contract, was qualified by the FDA as a medical device development tool in October 2020.
The second scale is clinical. The FDA's postmarket cybersecurity guidance of December 2016 divides risks into controlled, where the residual risk of patient harm from a vulnerability is acceptable, and uncontrolled, where it is not. For an uncontrolled risk, the agency said it did not intend to enforce its usual reporting rules for corrections if four conditions held. There were no known serious adverse events or deaths; the maker told its customers within 30 days of learning of the risk, with interim measures and a plan; it fixed the risk, validated the fix and distributed it within 60 days; and it took part in an information-sharing organization for the sector. Since March 2023 the law has set the pace for connected devices: known unacceptable vulnerabilities fixed on a "reasonably justified regular cycle," and critical ones that could cause uncontrolled risks "as soon as possible out of cycle."
Coordinated disclosure
Most device vulnerabilities are found by people outside the company: academic researchers, security firms, hospital engineers. Coordinated vulnerability disclosure is the practice by which they report a weakness privately to the maker, the maker acknowledges and investigates it, and the two agree when the details are published, ideally with a fix or mitigation available. A government coordinator often stands between them; the MiniMed advisories of 2018 and 2019 were published by the US government's coordinator, now part of CISA, after researchers reported the flaws and Medtronic analyzed which models shared them. Two international standards describe the maker's side, ISO/IEC 29147 for receiving and publishing reports and ISO/IEC 30111 for handling them internally, and both were flagged for revision in 2025. Since 2023 US law has required makers of connected devices to submit a plan for monitoring and addressing vulnerabilities after release, "including coordinated vulnerability disclosure and related procedures."
The hospital's share
A fix protects no patient until it is installed, and for many devices installation happens in the hospital. Each update has to be scheduled around clinical use, tested against the hospital's own configuration where the integration of Chapter 18 could be affected, and applied unit by unit, sometimes by the maker's service engineers and sometimes by the hospital's clinical engineering staff. Armis, which found URGENT/11, estimated in December 2020 that 97 percent of the operational-technology devices affected had not been patched; its figure covers far more than medical devices, but it shows how slowly fixes reach equipment that is not a personal computer.
Publishing a patch moves the risk from the maker's queue to the hospital's. Until a unit is updated, its protection depends on what surrounds it: network segmentation, restricted access, disabled services and monitoring. Hospitals need the maker's SBOM and VEX statements to know which units are exposed, and makers need the hospitals' installed versions to know which patients remain at risk.
The EU's medical device guidance on cybersecurity, issued in 2019 and revised in 2020, describes security as a joint responsibility of makers, integrators, operators and users, while noting that only makers carry the legal duties. It also says a device should not depend on security controls in its environment. In practice, the defense of a device in the field is shared: the maker designs it to be patched, watches for new weaknesses and issues fixes; the hospital keeps an inventory of what it runs, installs fixes and protects what it cannot yet fix.
Keep the SBOM of every released version searchable, monitor vulnerability databases and the catalog of exploited vulnerabilities daily, and record a VEX statement for each match. Judge each applicable vulnerability by exploitability and clinical harm, set a target date in the regular cycle or out of it, publish a disclosure policy with a contact address, and tell hospitals which versions are affected and what to do until they can patch.
- URGENT/11 put eleven vulnerabilities into devices through an operating system's network stack that no device maker had written.
- A searchable SBOM finds affected products within hours, and a VEX statement records whether each one is actually exposed.
- Severity scores measure the weakness; FDA guidance and US law set fix timelines by whether the risk to patients is controlled.
- Researchers report through coordinated disclosure, and fixes protect patients only when hospitals install them; both sides share the defense.
+ Part VII · The field in 2026
Who builds software devices.
The FDA's public list of AI-enabled medical devices was last updated on 4 September 2026. It is a table rather than a register: each row is one marketing authorization, and the agency says the list "is not a comprehensive resource." The page itself gives no total. Independent analysts counted 1,451 entries decided by the end of 2025 and 1,524 including decisions up to 30 March 2026, and in November 2025 the director of the FDA's device center told an advisory committee that the agency had authorized more than 1,200 AI-enabled devices. The single chapter of this part uses the list and the makers' own records to describe the field as it stood in October 2026: which devices reach patients, who makes them, and where generative AI has arrived. It is the most dated chapter in the guide and will be the first to age.
Software and AI devices in 2026.
+ The questionWhich software and AI devices reach patients in 2026, who makes them, and where does generative AI stand?
Mostly images
About three in four entries on the list are radiology devices. One analyst's count, reported in July 2026, put radiology at 1,104 of the 1,451 entries decided by the end of 2025, 76 percent, and at about the same share of the devices authorized in 2025 alone. The rest come from other specialties, ophthalmology among them, where IDx-DR sits.
The reasons are not hard to see from the earlier chapters. Images are digital from the moment they are made, they carry the standard headers of Chapter 18, and hospitals have kept them for decades in archives alongside the reports that describe them, which gives makers training data with something close to a label attached. A reading task also has a natural reference, the expert reader of Chapter 12, against which a model's answers can be scored.
Regulation has reinforced the pattern. When the FDA granted a De Novo to the stroke-triage program ContaCT on 13 February 2018, it created a new classification for radiological software that flags suspected findings and notifies a specialist. Later devices of the same type could then be cleared through the faster 510(k) route by showing they were substantially equivalent to one already on the market, and later devices were. A first authorization in a category opens the road for the products that follow.
Who makes them
The same analyst's counts of radiology authorizations by company, through the end of 2025, put four makers of scanners and imaging systems at the top: GE HealthCare with 120, Siemens Healthineers with 89, Philips with 50 and Canon with 45, followed by United Imaging with 38. Aidoc, a company that makes only AI software, came next with 31, and DeepHealth with 28. Much of the equipment makers' AI runs inside or beside their own scanners and reaches hospitals with the hardware. The software specialists sell triage and detection programs that run beside the archive, often through a platform that hosts several makers' tools behind one integration.
A third group rarely appears on the list but carries much of this guide: makers of pumps, pacemakers, ventilators and monitors, whose software drives therapy rather than reading images. Most of their software is conventional code, judged by the lifecycle of Part II and defended as in Part VI. The MiniMed and Abbott cases of Chapters 22 and 23 came from this group, not from AI.
GE HealthCare said in July 2025 that it had 100 listed authorizations; the analyst's count put it at 120 radiology authorizations by the end of 2025. Both can be right, because they count at different dates and by different rules. Units matter more. Qure.ai, a company based in Mumbai, said in February 2026 that it held 26 FDA-cleared indications across 9 products, about 2.9 indications per product, while its own regulatory page listed 7 cleared products. Aidoc announced in January 2026 a clearance covering 14 acute findings on CT from a single model, 11 of them newly cleared. If that was one submission, it is one entry counted by submission and 14 counted by indication. Two analysts counting radiology's share of 2025 authorizations got 75 percent and 71.5 percent, 211 of 295, a gap of 3.5 points from counting method and cutoff. A number of "AI devices" means little until it says whether it counts submissions, products or indications, and at what date.
Company figures for clearances, indications and "firsts" are written for investors and customers. Each can be checked against the FDA's databases, where every authorization has a number, a date and a decision summary, and against the list's own entries. Where the two disagree, the regulator's record is the evidence.
Makers in India
Indian makers reach patients first through other regulators. Qure.ai's head-CT triage program qER was cleared by the FDA on 17 June 2020, under the classification created for ContaCT, and the company's regulatory page lists CE certificates under the EU's device regulation for several products. Its page also reports an approval from India's regulator without giving a number. India's own requirements for software devices were set out in one place only in July 2026, in the regulator's final guidance on medical device software, described in Chapter 26. An Indian manufacturer announced in August 2026 what it called India's first manufacturing license for diagnostic software, a claim about in vitro diagnostic software that is not specific to AI. This guide found no reliable source naming the first AI software device licensed by India's regulator.
Generative AI at the door
Generative AI, models that produce text, images or speech rather than a score, reached the FDA's advisory process before it reached the list. The agency's Digital Health Advisory Committee met for the first time in November 2024 to discuss the whole product lifecycle for generative AI devices, and again in November 2025 on generative AI for mental health. At the second meeting the device center's director said that none of the digital mental-health devices authorized so far involved generative AI. The list page, as updated in September 2026, says the FDA will explore ways to identify and tag devices that incorporate foundation models, "from large language models (LLMs) to multimodal architectures"; no entry was tagged.
A device whose maker says it uses large language models was cleared six months before its maker announced it. UpDoc, an insulin-management program for adults with type 2 diabetes, was cleared through a 510(k) on 23 December 2025 as a drug-dose calculator, class II. Its public summary describes conversational data collection by voice or chat and insulin instructions computed from parameters set by the patient's clinician, and reports non-clinical testing only. It does not use the words "large language model" or "generative AI." The company announced on 25 June 2026 that this was the first FDA clearance of software using patient-facing large language models; the FDA has not confirmed the claim, and when a reporter asked whether generative AI made treatment decisions, the chief executive would not say. The clearance Aidoc announced in January 2026, built on what the company calls a foundation model, is a different case: a large image model, not a language model, doing a triage task of a kind already regulated.
The evidence question for generative AI is the central question of this guide in a new place. A language model that collects a patient's answers and passes them to fixed, cleared rules is an interface, and the evidence can concentrate on whether it captures what the patient meant. A model that writes the recommendation itself is the decision, and it would need the test sets, references and subgroup results of Parts III and IV for an output that can vary with every phrasing. The public record should say which of the two a device is. For UpDoc, the summary's description, with doses computed from clinician-set parameters, points to the first.
For any software device, find its FDA submission number, decision date and summary, and note the product code and the device it was compared with. Check which indications the authorization covers, whether it includes a change control plan, and what testing it reports. Treat a maker's counts and "firsts" as claims until the record supports them, and look for where any language model sits relative to the clinical decision.
- About three in four entries on the FDA's list of AI devices are radiology devices, helped by digital archives, natural references and early classifications.
- Scanner makers lead the counts, AI software specialists follow, and the makers of therapy devices carry most conventional device software.
- Counts of AI devices depend on whether they count submissions, products or indications; the regulator's record settles disagreements.
- UpDoc, cleared in December 2025, is claimed by its maker as the first device with a patient-facing language model; the FDA record does not say so.
+ Part VIII · Quality and regulation, as of October 2026
What a software device must show.
On 27 July 2026 a regulation the EU calls the Digital Omnibus on AI entered into force, three days after its publication. Among other changes to the AI Act of 2024, it moved the date from which AI in products such as medical devices must meet the act's high-risk requirements, from 2 August 2027 to 2 August 2028. It nearly did more: the European Parliament had proposed moving medical devices to a lighter category of the act's product list, and the final compromise left them where they were. Rules for software devices change like this, by amendment, guidance and revised date, while the mechanism beneath them holds. The four chapters of this part therefore put the lasting mechanism first and the clauses last: what makes software a device and sets its class, what evidence it must file, which updates need a regulator, and what the rules on security and AI add, including for the hospital that deploys it. Every status in them is as of October 2026.
Intended use draws the line.
+ The questionWhich software is regulated as a medical device in the US, the EU and India, and in what class?
Two guidances on one day
On 6 January 2026 the FDA reissued two guidances that together mark where its oversight of software stops. The first, on general wellness products, said that certain non-invasive products estimating blood pressure, oxygen saturation, blood glucose or heart rate variability can be wellness products, outside device regulation, when their outputs are "intended solely for wellness uses." The conditions are strict. Such a product may suggest that a user see a health professional when a reading falls outside a wellness range, but it may not name a disease, call a value abnormal, use clinical thresholds or give alerts for managing a disease, and it may not show values that look clinical unless they have been validated.
The second, on clinical decision support software, revised the criteria for software that advises clinicians without being a device. It added enforcement discretion for software that gives a single recommendation where only one is clinically appropriate, and it moved its treatment of time-critical decisions to the criterion about whether a clinician can independently review the basis of a recommendation, the change Chapter 16 described. The FDA corrected the decision-support guidance on 29 January 2026. Neither document changed the law. Both moved the practical line by explaining how the agency reads it, which is how the boundary of software regulation usually moves.
The US line
In US law a device is an instrument or related article intended for the diagnosis of disease or other conditions, or for the cure, mitigation, treatment or prevention of disease, and software can meet that definition. The 21st Century Cures Act of December 2016 carved five kinds of software function out of that definition. Software for the administrative support of a health care facility is excluded, as is software for maintaining a healthy lifestyle unrelated to any disease, software serving as an electronic patient record, and software that transfers, stores, converts or displays laboratory or device data without interpreting it. The fifth exclusion is clinical decision support, which must meet four criteria at once.
The software must not acquire, process or analyze a medical image, a signal from an in vitro diagnostic test, or a pattern or signal from a signal acquisition system. It must display, analyze or print medical information. It must support or provide recommendations to a health care professional. And it must let that professional independently review the basis for its recommendations, so that the professional does not rely primarily on them. IDx-DR fails the first criterion, because it analyzes images, and the fourth, because no professional reviews the images before the result is given. It is a device on two counts.
A function that is a device is then classed by risk, I, II or III, through a classification regulation that names the device type, as IDx-DR's 2018 De Novo created one for retinal diagnostic software in class II. The class sets the route to market: most class II software reaches it through a 510(k) or a De Novo, class III through premarket approval.
The EU line
The EU's medical device regulation classifies software by its Rule 11. Software that provides information used to take decisions for diagnosis or therapy is class IIa; it is class IIb if those decisions could cause a serious deterioration of a person's health or a surgical intervention, and class III if they could cause death or an irreversible deterioration. Software that monitors physiological processes is IIa, or IIb where it monitors vital parameters whose variations could result in immediate danger. All other software is class I. The rule's practical effect is that almost any software informing a clinical decision is at least class IIa, and so needs a notified body, an independent organization designated to assess conformity, before it can carry the CE mark.
The Medical Device Coordination Group's guidance on qualifying and classifying software, first endorsed in October 2019, was revised in June 2025. According to a regulatory consultancy's review, the revision made no substantive change, but it used the term "medical device artificial intelligence" for the first time, with a reference to the AI Act, and added examples, including software that informs a prognosis or prediction.
Every scheme in this chapter asks what could happen to a patient if the software's information or action were wrong. A maker that classifies by how the software usually performs, or by the clinician who is expected to catch its errors, will usually reach a lower class than the regulator.
India's grid
India's regulator published its final guidance on medical device software on 21 July 2026, after a draft in October 2025. It drops the terms software as a medical device and software in a medical device, and distinguishes standalone software from software that is part of, or drives, a hardware device. Software that drives or influences hardware takes the hardware's class. Standalone software is classed on a grid that follows the 2014 international framework of Chapter 1, with India's classes A to D.
Take autonomous screening for diabetic retinopathy from retinal photographs, used by primary-care staff, as in the opening scene. In the US it is a device, because it analyzes images and no clinician reviews the basis of its answer; its type is class II. Under the EU's Rule 11 it provides information used for a diagnostic decision, so it is at least IIa; whether it is IIb turns on whether a missed referral could cause a serious deterioration of health, an argument the maker must make and the notified body must accept. On India's grid, a situation that is serious and software that diagnoses give class C; a critical situation would give D, and software that only informed a clinician in the same serious situation would be class A. India's guidance adds that software used by non-clinical users in a serious situation without specialist support may be treated as critical, which would move this function to D if lay operators ran it at home. One function, three schemes: class II, IIa or IIb, and C or D, each decided by the claim and the user.
The rest of India's scheme follows the same logic. Where several classes could apply, the highest wins, and the regulator keeps a list of software it has classified. Manufacturing licenses for classes A and B are issued by state authorities and for classes C and D by the central authority. The guidance states that it reflects current practice under India's Medical Devices Rules of 2017 rather than creating a new control. Internationally, the medical device regulators' forum published a further document in January 2025 on characterizing software risk, which suggests assuming, where possible, that the software will fail, so that the risk assessment rests on the consequences rather than on an estimate of probability.
State in one paragraph what the software is for, who uses it, in which clinical situation, and what its output decides. Classify it under each market's scheme from that paragraph, record the reasoning, and check it against the regulators' examples. Keep the intended use under change control, because a new claim or a new user can change the class without a line of code changing.
- The FDA moved its boundary in January 2026 by guidance: wellness estimates and single-recommendation decision support, under unchanged law.
- US law excludes five kinds of software; decision support qualifies only if a clinician can review its basis and it does not analyze images or signals.
- EU Rule 11 makes almost any decision-informing software class IIa or higher; India classes standalone software A to D on a grid.
- The same screening function is class II in the US, IIa or IIb in the EU and C or D in India, decided by the claim and the user.
Documents, validation and clinical evidence.
+ The questionWhat does each regulator ask to see before a software or AI device is sold, and what do the validation rules ask of the sites and makers that use software under GMP?
Two levels of documentation
A documentation level, in the FDA's guidance on premarket submissions for device software, final since 14 June 2023, is the amount of software evidence a submission must contain, set by what a failure could do. The level is Enhanced where a failure or latent flaw in the software "could present a hazardous situation with a probable risk of death or serious injury," judged before any risk control is counted; otherwise it is Basic. The guidance replaced one from 2005 and applies to every device with software, AI or not.
Most of the file is the same at both levels: the evaluation of the level itself, a description of the software, the risk management file, the software requirements, an architecture diagram, the version history and the list of unresolved anomalies. The Enhanced level adds the detailed design specification, the configuration management and maintenance plans, and the protocols and reports of unit and integration testing, where Basic asks for a summary of testing and the system-level protocol and report. At the Basic level a maker can describe its development practices by declaring conformity to IEC 62304, which the FDA has recognized in full since 2019. Chapters 4 and 7 traced where each of these records comes from; the guidance is the list of which ones leave the company.
Behind the submission sits the quality system. Since 2 February 2026 the FDA's quality rule for devices, renamed the Quality Management System Regulation, has incorporated the international standard ISO 13485:2016 by reference, and the agency's inspectors have stopped using their old inspection technique. The surgical energy guide in this library treats IEC 62304 and the documentation levels as items on a generator maker's compliance list; this chapter has followed them from the software's side.
What the FDA asks of AI
For AI, the FDA's main document is still a draft. Its guidance on lifecycle management and marketing submissions for AI-enabled device software functions was published in January 2025 and remained a draft in October 2026. It proposes sections a submission should contain beyond the general software file: a description of the device and its user interface, labeling, risk assessment, data management, the model's description and development, validation, plans for monitoring performance in the field, cybersecurity, and a public summary.
Its expectations read like Parts III and IV of this guide. Test data should be independent of training data, for example drawn from different clinical sites, and sequestered from the developers. Validation data should represent the intended population, and relying on a single collection site is generally not appropriate; the draft gives at least three geographically diverse US sites as an example of what may be. Performance should be reported across subgroups such as sex, age, race, ethnicity, disease severity, site and acquisition equipment, at each operating point. It suggests, without requiring, a model card: a short structured summary of the model, its data and its performance, placed in the labeling. The final guidance on change control plans of Chapter 21 completes the set.
The EU and India
In the EU, a software device's evidence goes into the technical documentation that a notified body reviews. For clinical evidence, the coordination group's guidance of March 2020 sets out three components. The first is a valid clinical association: the software's output is associated with the clinical condition it targets. The second is technical performance: the software generates its intended output accurately and reliably from its input data. The third is clinical performance: the output is clinically relevant in accordance with the intended purpose. All three are kept current by post-market evaluation. The guidance predates the AI wave and does not mention AI. IEC 62304 is not a harmonized standard under the EU's device regulation, and neither is the security standard of Chapter 23: in the list as consolidated on 17 June 2026, neither appears. A maker can still use them as the state of the art, but they give no presumption of conformity.
India's final guidance of July 2026 uses the same three components. For AI it adds dataset requirements of its own: the composition of training, validation and test data, including demographic distribution, geographic origin and clinical diversity; whether data are real-world, from public databases or synthetic; and a justification where models were trained or validated outside India. Performance shown on non-Indian data should be supplemented with validation in representative Indian populations. The guidance lists IEC 62304, IEC 81001-5-1, ISO 14971 and ISO 13485 among the standards that may apply.
The FDA's AI-enabled device guidance, its draft on change control plans for all devices and the EU's draft manufacturing annex on AI all describe where regulators are heading, and makers sensibly design toward them. A submission is still judged against the final documents in force on the day it is filed, so a file needs to show which version of each it followed.
Standards in motion
Standards change more slowly than guidance, and their status is easy to misreport. IEC 62304's first edition dates from 2006, with an amendment in 2015. A second edition, widened to health software in general, has been in preparation for years, and its status can be read directly from the international project record.
ISO and IEC track every project with a stage code: 30 is the committee stage, where drafts circulate among experts; 40 the enquiry, a draft circulated to national bodies for vote; 50 the approval of a final draft; 60.60 publication. The project record for the second edition of IEC 62304, as mirrored by a national standards body, showed stage 30.20, a committee draft ballot initiated, from 7 August 2026, with the next milestone due on 27 November 2026. A consultancy that follows the committee reported almost 1,500 comments on the first committee draft and expected the final draft no earlier than 2028. A vendor's web page, by contrast, said the final draft stage had begun on 22 May 2026 and that publication was scheduled for 12 August 2026. A standard cannot be at stage 30 and stage 50 at once. The record wins: in October 2026 the 2015 consolidated first edition is the one in force, and anything said about the second edition, such as the proposed replacement of the three safety classes by two levels of process rigor, reported by the German association VDE, describes a draft.
Validation rules for those who use software
The rules for sites and makers that use software, rather than sell it, run in parallel. The FDA's rule on electronic records and signatures, 21 CFR Part 11, has applied since 1997. Its guidance on computer software assurance, final in September 2025 and revised on 3 February 2026 under a new title that refers to quality management system software, sets out the risk-based testing of Chapter 9 for software used in production and in the quality system; it does not cover the device software itself. In the EU, a revised Annex 11 on computerized systems and a new Annex 22 on AI, the draft described in Chapter 19, went out for consultation from July to October 2025. On 8 October 2026 the EU's list of good manufacturing practice guidelines still showed the Annex 11 of 2011 and no Annex 22.
Determine the documentation level before design starts, and keep each required record, from requirements to unresolved anomalies, current with every release so that a submission is an export rather than a project. For AI, keep the data management records, the sequestered test sets and the subgroup results the FDA, EU and Indian documents ask for, and record which version of each guidance and standard the file follows.
- The FDA asks for Basic or Enhanced software documentation, set by whether a failure could probably cause death or serious injury before controls.
- The FDA's AI guidance is still a draft; it asks for independent, sequestered and representative test data and subgroup results, and suggests a model card.
- The EU and India ask for clinical association, technical performance and clinical performance; India adds Indian-population data.
- IEC 62304's second edition is a committee draft, and Part 11, CSA and the 2011 Annex 11 govern software used under GMP, with a revised Annex 11 and a new Annex 22 still in draft.
Changes, jurisdiction by jurisdiction.
+ The questionWhich software changes can a maker release on its own, and which need a regulator first?
One release, three answers
The same two-week release can be routine in one jurisdiction and a new submission in another. Suppose a team finishes version 2.4 of an AI screening program. It retrains the model on more data from the same kinds of clinics, moves the operating threshold within the range already validated, applies a security patch to the operating system and upgrades an image-processing library to a new major version. The code was built, verified and released through the pipeline of Chapter 9, and every record is in place. Whether it may now reach patients depends on where they are.
The three regulators ask the same underlying question, the one this guide has asked of every change: could it move a claim the evidence supports, or create a risk the evidence did not cover? They encode the answer differently. The FDA asks a series of questions that lead to a likely answer; the EU sorts changes into significant and not; India divides them into major and minor. Each also offers a way to agree some changes in advance.
The FDA's questions
The FDA's guidance on deciding when to submit a 510(k) for a software change, final since 25 October 2017, applies to cleared devices and to those authorized through De Novo. It works as a flowchart. A change made solely to strengthen cybersecurity, with no other effect on the device, is likely not to need a new submission, and neither is one made solely to return the device to the specification of its most recently cleared version. A change that introduces a new risk, or modifies an existing one, that could result in significant harm and is not effectively mitigated likely does need one. So does a new or modified risk control for a hazard that could cause significant harm. The last question is whether the change could significantly affect clinical functionality or performance specifications; if so, a submission is likely needed. Changes to what the device is indicated for go through the FDA's general guidance on device changes.
The flowchart leaves the decision to the maker, which must document its reasoning and is accountable for it in an inspection. For a model, the performance question usually decides: retraining is meant to change performance, so a retrained model is likely to need a new submission unless an authorized change control plan already covers it. That is why the plans of Chapter 21 matter. UpDoc's, for example, lists five categories of change the maker may make without a new submission, none of which alters the cleared dosing logic.
The EU: significant change and substantial modification
Under the EU's medical device regulation, a maker tells the notified body that certified a device of planned changes, and changes that could affect conformity need that body's approval. The most detailed published list of what counts for software is in the coordination group's guidance on significant changes. That guidance was written for devices still certified under the older directives, but its software chart sets out the reasoning plainly. A change is significant if it brings a new or modified architecture or database structure, or a change of an algorithm, or if a closed-loop algorithm replaces a user's input. So is a new operating system or any new component, or a major change to one. A new diagnostic or therapeutic feature, a new channel of interoperability, a new user interface or a new presentation of medical data are significant too. Safety-neutral bug fixes, security updates, changes to appearance and gains in efficiency are not significant, provided they do not affect diagnosis or treatment. One significant answer makes the whole change significant.
The AI Act adds its own test. A substantial modification is a change to an AI system after it has been placed on the market that was not foreseen in its initial conformity assessment and that affects its compliance or modifies its intended purpose. A high-risk system that undergoes one needs a new conformity assessment. Changes that the provider predetermined at the initial assessment, and documented, are not substantial modifications: the EU's form of a change control plan. For medical device AI, those provisions apply from 2 August 2028, and the EU's joint guidance on the device rules and the AI Act, from June 2025, notes that retraining may trigger reassessment under both and that a plan agreed in advance can reduce how often.
A model retrained on new data has new parameters and, by design, new performance. Every regulator in this chapter treats that as a change to evaluate against the device's claims, and none treats it as maintenance. A plan agreed before authorization is the main way to make retraining routine.
India: major and minor
India's final guidance on medical device software, published in July 2026, divides software changes into two lists. Major changes need the licensing authority's approval before they are made. They include design changes that affect specifications, indications or performance; changes to software design or system requirements, new clinical claims or new types of input data; and changes of intended use. Label changes other than font, color or layout are major, and so are version changes that affect intended use, safety, effectiveness or risk controls. Minor changes are notified: bug fixes and security patches that do not affect intended use, safety or performance, version changes with no such effect, and re-tuning of performance within validated ranges. Software is identified by its version, revision level and build or release date. Under India's rules a minor change must be notified within 30 days.
For in vitro diagnostic devices, an addendum of 13 March 2026 to the regulator's questions and answers states that each new version of approved software needs a post-approval change application, the route the immunoassay guide in this library describes for analyzers. For AI, the July guidance adds an optional algorithm change protocol with four parts: a data management plan, a plan for evaluating and monitoring performance, a retraining plan where retraining is intended, and a plan for software updates and rollback.
Take the four changes of the opening section. In the US, the operating-system security patch falls under the flowchart's first question and is likely not to need a submission; the threshold move within the validated range and the library upgrade need a documented assessment of whether they could significantly affect performance; the retraining is likely to need a 510(k) unless an authorized plan covers it. In the EU, the security patch is not significant, but the retraining is a change of an algorithm and the library's new major version is a major change of a component. By the reasoning of the coordination group's chart, at least two of the four are significant, and the threshold move may be a third if it changes the algorithm's output; under the regulation, the notified body must approve any of them that could affect conformity. In India, the patch and the threshold move are minor and notified within 30 days, the library upgrade is minor if it affects nothing the guidance lists, and the retraining is major if it changes performance specifications. One release yields one likely submission in the US, at least two significant changes in the EU and one major change in India, before the AI Act applies to any of them.
Patches and borrowed code
Security patches and updates to third-party components are where the schemes come closest. All three treat a patch that only closes a vulnerability as something a maker can release without prior approval, after documenting its assessment and, in India, notifying the regulator, which keeps the regular and out-of-cycle fixes of Chapter 24 possible. They part company on larger component changes: a new major version of an operating system or library is significant under the EU's chart, while the FDA and India ask what it could do to safety and performance. A maker that keeps its SBOM, its regression evidence and its assessment of each component change in one record can answer all three from the same file.
For each change in a release, record which claims and risks it could affect, the verification run, and the regulatory decision for each market: documented assessment or submission in the US, significant or not in the EU, major or minor in India. Group changes that need a regulator into planned releases, keep security patches separate so they can ship fast, and write a change control plan for the changes a model will need.
- The FDA's 2017 software-change flowchart lets security-only and return-to-specification changes ship, but changes that affect risk or performance usually need a 510(k).
- The EU treats algorithm, architecture and major component changes as significant; the AI Act adds substantial modification, with predetermined changes excluded.
- India lists major software changes needing approval and minor ones to be notified, and offers an optional algorithm change protocol.
- One release can need one submission in the US, notified-body review in the EU and approval in India; plans agreed in advance make change routine.
Section 524B, the AI Act and the deployer's duties.
+ The questionWhat do the cybersecurity and AI-specific rules add, what do they ask of the hospital that deploys the device, and how does one device meet two EU regulations at once?
Section 524B
Ninety days after an appropriations law was enacted in the US on 29 December 2022, a new section of the Food, Drug, and Cosmetic Act took effect, on 29 March 2023. Section 524B applies to any cyber device: one that includes software validated, installed or authorized by its maker, can connect to the internet, and has characteristics that could be vulnerable to cybersecurity threats. A maker submitting such a device must do four things. It must provide a plan to monitor, identify and address vulnerabilities after release, including coordinated disclosure. It must design, develop and maintain processes giving reasonable assurance that the device and related systems are secure, and make updates and patches available on a regular cycle and, for critical vulnerabilities, out of cycle. It must provide a software bill of materials covering commercial, open-source and off-the-shelf components. And it must meet any further requirements the FDA sets by regulation.
The FDA gave makers a transition: until 1 October 2023 it worked with sponsors whose submissions lacked the new information, and after that date it could refuse to accept them. Its premarket cybersecurity guidance, the source of Chapters 22 to 24, was revised on 3 February 2026 to align with the new quality management system rule, without new technical expectations, according to two law and certification firms' reviews. In March 2026 a standards body announced that the FDA had recognized its consensus report on cybersecurity considerations unique to machine-learning devices.
One device, two EU regulations
In the EU, cybersecurity is part of the medical device regulation itself. Its general safety and performance requirements ask that software be developed according to the state of the art, taking account of the development lifecycle, risk management including information security, verification and validation. They also require makers to set out the minimum requirements for hardware, IT networks and IT security measures, including protection against unauthorized access, needed to run the software as intended. The coordination group's cybersecurity guidance of December 2019, revised in July 2020, adds that a device should not rely on security controls in its operating environment. The EU's Cyber Resilience Act, which applies to most products with digital elements from 11 December 2027, excludes devices covered by the medical device and diagnostic regulations, so for a hospital it matters for other software it buys, not for its medical devices.
AI adds a second regulation on top. Under the AI Act, an AI system is high-risk if it is, or is a safety component of, a product covered by the EU laws listed in its first annex and that product must undergo assessment by a third party. Medical devices are on that list, and almost all decision-informing software needs a notified body under Rule 11, so most medical device AI is high-risk. The high-risk requirements cover risk management, data governance, technical documentation, automatic logging, information for those who deploy the system, human oversight, and accuracy, robustness and cybersecurity. The Digital Omnibus of July 2026 kept medical devices in that annex, moved the date these requirements apply to medical device AI to 2 August 2028, and, according to the consolidated text, clarified that AI used solely for purposes other than safety does not qualify as a safety component.
The two regulations meet in one assessment. A device that is also high-risk AI follows the conformity assessment of the medical device regulation, and the notified body checks the AI Act's requirements within it. The joint guidance of the device coordination group and the EU's AI board, published in June 2025, describes one technical documentation, an integrated quality system and integrated risk management, rather than two parallel files.
Segmentation, firewalls and access control at the site are valuable layers, and the EU's minimum IT requirements belong in the instructions for use. A device designed to be safe only behind them, though, transfers its security to an organization that did not design it and may change its network tomorrow.
The deployer's duties
The AI Act gives duties to the organization that uses a high-risk system under its own authority, which it calls the deployer: for a hospital AI device, the hospital. It must use the system in accordance with its instructions and assign human oversight to people with the necessary competence, training and authority. Input data under its control must be relevant and sufficiently representative for the intended purpose. It must monitor the system's operation, inform the provider without undue delay of risks, and immediately of serious incidents, and suspend use where it sees a risk. It must keep the logs the system generates automatically, where they are under its control, for at least six months, and inform workers before the system is used in their workplace. The Omnibus also rewrote the act's article on AI literacy: providers and deployers must take measures to support it, without having to guarantee any level of literacy in an individual.
These are, almost line for line, the site practices of Chapters 18 to 20 made into law: integration within the specification, a local check that the inputs match the conditions of the evidence, monitoring with named people, and a channel back to the maker. For hospitals using AI that is a medical device, they apply from 2 August 2028.
India
India's guidance of July 2026 brings security into the same file as safety. According to trade reports of its text, it asks makers to keep a bill of materials that includes an SBOM, covering third-party, open-source and commercial components, and to link it to monitoring and mitigating known vulnerabilities. Cybersecurity is treated as a risk area, vulnerabilities are to be monitored after deployment, and devices are to keep operating safely during incidents, including partial loss of connectivity, denial of service and loss of integrity. It lists IEC 81001-5-1 among the standards that may apply. Its AI dataset rules were set out in Chapter 27.
Patient data are governed by the Digital Personal Data Protection Act of 2023 and its rules, notified in November 2025 and phased in. The rules on the data protection board applied at once; the registration of consent managers applies from 13 November 2026; and most obligations on those who process data, including security safeguards and notifying breaches, apply from 13 May 2027. The penalty for failing to keep reasonable security safeguards can reach 250 crore rupees.
From 8 October 2026, the next dates that reach a software device or the site that runs it fall over two years. India's consent-manager rules apply in about one month, on 13 November 2026, and its main data protection duties, including breach notification, in about seven months, on 13 May 2027. The Cyber Resilience Act's full application follows on 11 December 2027, about 14 months away, for the hospital's non-device software. The AI Act's high-risk requirements for medical device AI, including the hospital's duties as deployer, apply on 2 August 2028, about 22 months away. A device authorized today and updated every quarter will see about seven releases before the last of these dates, and each release will be assessed against the rules in force on its own release day.
For each device, list the rules that apply in each market and the date each takes effect: section 524B for cyber devices, the device regulation and the AI Act together in the EU, India's software guidance and data protection rules. Assign each duty to the maker or the site, write the site's duties into the responsibility agreement of Chapter 18, and review the list whenever a regulation changes its dates.
The four-hour operators, again
The operators of the opening scene had four hours of training and a locked program. What their trial showed in 2017 was narrow and solid: that program, on that camera, run by people like them, gave the right screening decision for adults like those enrolled often enough to beat bars agreed in advance. It could not show what the program would do in a clinic whose systems sent images differently, after a camera or model changed, or against someone trying to make it fail.
The same device filed in October 2026 would carry evidence for each of those gaps. Because its images and answers cross the internet, it is a cyber device, so its file would hold a threat model, an SBOM, security test results and a disclosure plan. Its software documentation would be at the level its potential for harm sets. Following the FDA's draft for AI, it would show test data drawn from several sites and kept from its developers, with results by subgroup and by camera, and very likely a change control plan for the retraining it will need. In India it would be expected to add validation in Indian populations; in the EU, from August 2028, logging and human-oversight measures assessed by its notified body. The hospital that installed it would test its interfaces, check its inputs against the specification, run it silently on local patients before relying on it, watch its positive and declined rates, and tell the maker when something moved.
None of that changes the invariant. Every way the software can fail was built in on the day it shipped or arrived with an update, and for the model the training data is part of what was built. The evidence that it will be right for the next patient is a measurement of what was built, taken on patients chosen to resemble the next one; what keeps it true after an update is the discipline that measures again. Four hours were enough to take the pictures. Everything else in this guide is what it takes to keep trusting the answer.
- Since March 2023, US cyber devices must file a vulnerability plan, show secure processes with patching, and provide an SBOM.
- In the EU, a medical device that is high-risk AI meets the device regulation and the AI Act in one assessment, with AI duties from 2 August 2028.
- Hospitals deploying such AI become deployers, with duties for oversight, input data, monitoring and logs that mirror good site practice.
- India's 2026 guidance brings SBOMs and cyber resilience into the device file, and its data protection duties phase in through May 2027.
Lessons.
Fourteen working rules follow from the chapters, for anyone who builds, buys, installs, regulates or writes about software and AI in medical devices. Each names the chapters that explain why it holds.
- Write the intended use before anything else. The claim turns code into a device, sets its class in each market and fixes what the evidence must show; a new claim or a new user can change all three without a line of code changing (Chapters 1 and 26).
- Look for the fault and the input that reaches it, not a wear-out rate. Software does not age: every failure was built in on the day it shipped or arrived with an update, and it shows itself only when some input reaches it (Chapter 2).
- Build confidence from the process, because testing alone cannot supply it. Paths outnumber any test campaign and very low failure rates cannot be demonstrated by running the software, so the evidence is requirements derived from risk and traced to tests, an architecture that keeps the dangerous part small, and rigor scaled to harm (Chapters 3 to 5 and 7).
- Account for every component someone else wrote. Record each one's exact version, published anomalies, support status and replacement plan, generate the SBOM from the build, and treat pretrained models and public datasets as components too (Chapter 6).
- Let every sprint and every pipeline run leave its records. IEC 62304 prescribes activities and records, not phases, so agile increments and automated builds can produce its evidence, provided a definition of done includes verification and a person still decides each release (Chapters 8 and 9).
- Keep the data that judges a model away from the people and data that built it. Separate test sets by patient and, where possible, by site, sequester them from developers, and watch for the leakage paths that let a model learn the hospital instead of the disease (Chapters 10 and 11).
- Ask who decided the right answer and where yes begins. The reference standard limits what measured accuracy can mean, and the operating point, the endpoints and the treatment of declined cases must be fixed before the test and reported with their intervals (Chapters 12, 13 and 17).
- Test where the device will be used, and report by subgroup. Accuracy is a measurement on particular patients, equipment and settings; it moves when any of them moves, and an average can hide a group the model fails (Chapters 14 and 15).
- Measure the reader and the model together when a clinician uses the output. An assistant that is usually right can still make a reader worse, so the evidence for an assistive device is the pair's performance, with automation bias in view (Chapter 16).
- Treat installation as part of the evidence. Test every interface with local data, confirm that local inputs fall inside the specification, run the device silently on local patients before anyone relies on it, and write down who notifies whom of changes on either side (Chapters 18 and 19).
- Watch inputs and output rates from the first day, and close the loop on a schedule. Shifts show in days in what the device receives and returns; accuracy shows only in reference reads of sampled negatives and matched outcomes, with named people acting on agreed limits (Chapter 20).
- Agree in advance what a model may change, and fix the tests each version must pass. A change control plan is the main route to routine retraining: authorized plans in the US, an optional algorithm change protocol in India and predetermined changes under the EU's AI Act from August 2028; outside such a plan, retraining is a change a regulator must see (Chapters 21 and 28).
- Design the device to be attacked and to be patched. Authenticate every interface, sign the code and keep a fallback image, build the threat model from the architecture, put the SBOM to work when vulnerabilities are published, and publish a way to report them (Chapters 22 to 24).
- Read the record, and date every status. Makers' counts and "firsts" are claims until a regulator's database supports them, and every rule, guidance and standard in this field is quoted as of a day; the mechanism lasts longer than the clause (Chapters 25 to 29).
Glossary.
Terms are defined as they are used in this guide.
- 510(k)
- A US premarket submission by which the FDA clears a device shown to be substantially equivalent to one already on the market. Once a De Novo has created a device type, later devices of that type can follow this route; IDx-DR's new versions were cleared this way in June 2021 and June 2022.
- AAMI TIR45
- A technical information report from AAMI, the Association for the Advancement of Medical Instrumentation, that maps agile development, in short cycles each ending with tested software, onto IEC 62304 in four layers: project, release, increment (a sprint of one to four weeks) and story (a few days of user-visible function). The FDA recognizes its 2012 and 2023 editions in full.
- AI Act
- The EU's 2024 law on artificial intelligence, under which most medical device AI is high-risk, because medical devices are on its product list and almost all decision-informing software needs a notified body. Its high-risk requirements, from risk management and data governance to logging and human oversight, apply to medical device AI from 2 August 2028, as moved by the Digital Omnibus of July 2026.
- AI model
- In this guide, an algorithm whose decision rules were fitted to example data rather than written line by line. Its accuracy is a measurement on a test set and holds only for patients like those in it, against the reference chosen, for the exact model that was locked.
- Algorithm
- The procedure the software follows. When its decision rules are fitted to example data rather than written line by line, the guide calls it an AI model.
- Algorithm change protocol
- In the FDA's 2019 discussion paper, the methods a maker would use to make, test and control changes to a machine-learning model within an agreed region of potential changes. India's July 2026 guidance offers an optional AI protocol of the same name in four parts: data management, performance evaluation and monitoring, retraining, and updates with rollback.
- All-pairs testing
- Testing every pair of parameter settings rather than every combination, which grows slowly with the number of settings. In 98 percent of 109 FDA recall reports that NIST researchers could trace, it would have revealed the failure; the HAMILTON-C6 ventilator fault needed four conditions at once, so a pairwise campaign would not have been sure to find it.
- Attack surface
- Every way data or commands can enter a device: radio and Bluetooth links, Wi-Fi and Ethernet, USB and serial ports, a cloud connection, a service interface used only by technicians. The MiniMed pumps' unauthenticated radio link was the part of theirs that mattered.
- AUC
- Area under the curve: the area under a receiver operating characteristic curve, from 0.5 for a coin toss to 1 for perfect separation, which summarizes how well a score separates diseased from healthy cases regardless of where the threshold is set. It hides the operating point and can stay stable while calibration drifts; the Google retinal network of 2016 scored 0.991.
- Automation bias
- The tendency to accept a machine's output in place of one's own judgment, which studies suggest is strongest when the person is uncertain. In a 2023 study, wrong suggestions labeled as AI cut radiologists' accuracy on mammograms from about 80 percent to between 20 and 46 percent, with the least experienced readers affected most.
- Autonomous, assistive and triage devices
- The three roles of software that analyzes clinical data. An autonomous device gives the answer itself, as IDx-DR does; an assistive device gives a reader extra information and the reader decides, so the reader's performance with and without it is what counts; a triage device works in parallel with the usual workflow, alerting a specialist while radiologists still read every image.
- Baseline
- The results of a site's local performance check, kept as the reference for later monitoring: the distribution of inputs, the rate of declined cases, and the local sensitivity, specificity and predictive values. In agile work the word also names the controlled set of requirements approved at each release.
- Calibration
- Whether a model's scores mean what they say, such as predicted risks matching the true rate; discrimination is whether it ranks sick patients above healthy ones. Seven kidney-injury models at US Veterans Affairs hospitals kept their discrimination for nine years while their calibration drifted, so a fixed threshold would have alerted ever more often.
- CDSCO
- Central Drugs Standard Control Organisation: India's device regulator. Its final guidance on medical device software, of 21 July 2026, classes standalone software A to D, divides software changes into major and minor, and asks makers to justify AI models trained or validated outside India; licenses for classes A and B come from state authorities, for C and D from the central authority.
- CE mark
- The marking under which a device is sold in the EU. Under Rule 11 almost any software that informs a clinical decision is at least class IIa, so it needs a notified body's assessment before it can carry the mark.
- Change control
- The discipline that judges every change to a device against the claims it could move and the risks it could create before it reaches patients; for an AI model it adds a plan agreed with the regulator in advance that bounds what may change and fixes the tests each new version must pass. It is the guide's answer to what keeps the evidence true after an update.
- Clinical decision support
- Software that advises clinicians, which US law excludes from the device definition only if it meets four criteria at once: it does not analyze a medical image or a diagnostic signal or pattern; it displays, analyzes or prints medical information; it supports recommendations to a health care professional; and that professional can independently review their basis. IDx-DR fails the first and the fourth.
- Computerized system validation (CSV)
- Validation by its user of a computerized system that affects product quality, patient safety or data integrity, which GxP rules require of laboratories, blood services, clinical-trial organizations and drug manufacturers. Most follow GAMP 5, scaling the effort to the system's impact, complexity and novelty and building on the supplier's evidence.
- Confidence interval
- The range that expresses the uncertainty of an estimate made from a sample, at a stated confidence, usually 95 percent; it narrows as the number of cases behind the estimate grows. IDx-DR's observed sensitivity of 87.4 percent has an exact interval of 81.9 to 91.7 percent, and a sensitivity of 90 percent spans 78.6 to 95.7 from 50 patients but 87.1 to 92.3 from 500.
- Configuration management
- Identifying and controlling every item that goes into a build, from source code and SOUP versions to build scripts and the compiler with its settings, so that the same inputs produce the same software and the version tested is the version shipped. IEC 62304 requires it at every class.
- Coordinated vulnerability disclosure
- The practice by which someone who finds a weakness reports it privately to the maker, the maker investigates, and the two agree when the details are published, ideally with a fix; a government coordinator such as the US cybersecurity agency CISA often stands between them. US law has required it in connected-device makers' vulnerability plans since 2023.
- Coverage
- A measure of how much of the code a test campaign executed. The FDA's 2002 guidance describes a ladder from statement coverage, every line run once, which it calls insufficient on its own, through decision, condition and multiple-condition coverage, to full path coverage, which it calls generally not achievable; coverage shows what ran, not whether it was right.
- CRC
- Cyclic redundancy check: a checksum computed from a file by a published method, which detects corruption by a faulty memory chip or a noisy cable. Anyone who changes the file can compute a matching one, so the FDA tells makers not to rely on it as a security control; only a signature made with the maker's private key proves a file is genuine.
- CSA
- Computer software assurance: the FDA's risk-based approach to validating software used in production and in the quality system, finalized in September 2025 and reissued in February 2026. Functions whose failure could foreseeably compromise safety get documented, scripted testing and the rest lighter, unscripted methods; it does not cover the device software itself.
- CVSS
- Common Vulnerability Scoring System: the usual measure of a vulnerability's technical severity, built from factors such as the attack vector, its complexity and the impact on confidentiality, integrity and availability. The 2019 MiniMed weakness scored 7.1, High, on version 3; the score says nothing about how many patients use a device or what harm follows.
- Cyber device
- Under section 524B of the US Food, Drug, and Cosmetic Act, a device that includes software validated, installed or authorized by its maker, can connect to the internet, and has characteristics that could be vulnerable to cybersecurity threats. The FDA reads connection broadly, down to a USB port used briefly for service; IDx-DR, whose images and answers cross the internet, would be one.
- Cyber Resilience Act
- The EU law, CRA for short, on the cybersecurity of products with digital elements, which applies in full from 11 December 2027. It excludes devices covered by the medical device and diagnostic regulations, so for a hospital it governs other software it buys, not its medical devices.
- Dataset shift
- Any difference between the distribution of cases a model was fitted and tested on, meaning patients, images and labels in particular proportions, and the one it meets in use. It can change the inputs (a new camera or population), the prevalence, the right answer itself, called concept shift (a new grading guideline), or the acquisition chain (a scanner update, compression, a new format), and time produces it too.
- De Novo
- The US route by which the FDA grants authorization to a new kind of device and creates a classification regulation for its type, so that later devices of the type can be cleared through a 510(k). A De Novo is granted, not cleared: IDx-DR's, on 11 April 2018, created the class II type of retinal diagnostic software device.
- Decision summary
- The FDA's published account of the review behind a De Novo, with a shorter summary for most 510(k) clearances, covering the indications and limitations, the software, the special controls, the clinical study and the risks with their mitigations. IDx-DR's is the most complete public record of its evidence, though its date and one interval need checking against the counts.
- Declined case
- A case for which a device gives no usable answer, such as IDx-DR's result of insufficient image quality, which its labeling directs be referred. Headline figures usually leave such cases out: counting IDx-DR's 10 declined diseased participants as misses lowers its sensitivity from 87.4 to 83.2 percent, and counting them as referrals raises it to 88.0.
- Definition of done
- The conditions a story must meet before it counts as finished in agile work: requirement written and reviewed, design and code reviewed, unit and integration tests passed, trace links in place, risk analysis updated and any SOUP change recorded. Applied strictly, it lets the evidence for a release accumulate in small pieces as the work is done.
- Deployer
- Under the EU's AI Act, the organization that uses a high-risk AI system under its own authority; for a hospital AI device, the hospital. From 2 August 2028 it must follow the instructions, assign competent human oversight, keep its input data relevant and representative, monitor the system, report risks and serious incidents to the provider, and keep the logs for at least six months.
- Device software function
- Software that meets the legal definition of a medical device, whether it runs inside a pump or on a server by itself. What makes it one is its intended use, not anything in its code.
- DICOM
- Digital Imaging and Communications in Medicine: the standard format hospitals use for medical images and image-based results. Each image carries a header of tagged fields, such as the modality, the manufacturer, the model name and the equipment's software versions, which a device can read to refuse inputs outside its specification; IDx-DR accepted DICOM images from its 2021 version.
- Documentation level
- The FDA's measure, since June 2023, of how much software evidence a submission must contain: Enhanced where a failure could present a hazardous situation with a probable risk of death or serious injury, judged before any risk control is counted, and Basic otherwise. Enhanced adds the detailed design and the unit and integration test protocols and reports.
- End of support
- The date after which a software supplier no longer publishes security fixes. Extended support for Windows XP ended on 8 April 2014; devices that outlive their operating system's support stay exposed, as WannaCry showed in 2017 on unpatched or unsupported Windows systems; the FDA asks makers to state each component's support level and end-of-support date.
- Enrichment
- Recruiting more of one kind of participant or case than the population holds, such as the extra participants with poorly controlled diabetes in the IDx-DR trial, whose headline figures were corrected for it. In reader studies, enrichment with diseased cases can change how readers behave and biases predictive values.
- Evidence
- In this guide, records that someone outside the team could check: requirements, test results, study data, monitoring reports.
- Exploitability
- How easily an adversary could use a vulnerability. Because an attacker chooses the inputs, the FDA's cybersecurity guidance judges security risk by exploitability and by the harm that would follow, not by probability, and treats known vulnerabilities as reasonably foreseeable.
- External and prospective tests
- An external test uses data from sites, devices or periods that contributed nothing to a model's development and shows how it travels; a prospective test runs the device forward in its intended setting, on patients enrolled for the purpose, with its real workflow and operators. The IDx-DR trial was prospective, in primary care.
- Fault
- A flaw in a program's logic or data, present from the moment it was written, that causes a failure when an input reaches it, in every copy of that version. The HAMILTON-C6 ventilator fault waited inside three software versions for four ordinary events to coincide.
- FHIR
- Fast Healthcare Interoperability Resources: HL7's newer standard, built from web resources that can be addressed individually. Release 5 has been current since March 2023, and release 6 was in its second normative ballot in July 2026.
- Foundation model
- A class of large model, ranging in the FDA's words from large language models to multimodal architectures; in September 2026 the agency said it would explore tagging devices that incorporate one, and none was yet tagged. Aidoc's January 2026 clearance of 14 acute CT findings from one model rests on what the company calls a foundation model, a large image model.
- GAMP 5
- Good Automated Manufacturing Practice: the guide from the industry body ISPE that regulated companies follow to validate the computerized systems they buy and build, in its second edition of July 2022. It is risk-based, sorts software into categories from infrastructure to custom software, supports iterative and agile methods, and was joined in July 2025 by a separate GAMP guide to AI.
- Generative AI
- AI models that produce text, images or speech rather than a score. Where such a model sits relative to the clinical decision decides the evidence a device needs: one that only collects a patient's answers for fixed, cleared rules is an interface, while one that writes the recommendation is the decision.
- Good machine-learning practice
- Principles for developing AI devices, published by the FDA, Health Canada and the UK's MHRA in 2021 and finalized by the International Medical Device Regulators Forum in January 2025. Among them: test sets independent of training data, reference standards fit for purpose, data representative of the intended population, assessment of the human-AI team, and monitoring of deployed models.
- GxP
- Good practice: the collective name for the good-practice rules, such as good manufacturing practice, under which laboratories, blood services, clinical-trial organizations and drug manufacturers work. They require any computerized system that affects product quality, patient safety or data integrity to be validated by its user.
- Hazard, hazardous situation and harm
- The chain that ISO 14971 traces: a hazard is a potential source of harm, a hazardous situation is one in which a sequence of events exposes someone to it, and harm is what may follow. Each chain is judged by the severity of the harm and the probability of the sequence; for software, regulators suggest assuming the failure will happen.
- HL7
- Health Level Seven: a standards body, and its version 2 standard, which carries orders and reports between hospital systems as messages; the body says version 2 is used by 95 percent of US healthcare organizations.
- IEC 62304
- The international standard for the life cycle of medical device software, published on 9 May 2006, amended in June 2015 and recognized in full by the FDA since January 2019. It prescribes activities and their records, not their order, and sets three software safety classes; a second edition widened to health software was a committee draft in October 2026, not expected before 2028.
- IEC 80001-1
- The standard, in its 2021 edition, for organizations applying risk management before, during and after connecting a medical device or health software to their IT infrastructure, covering safety, effectiveness and security. It is marked for revision and is to be replaced by a standard in the 81001 series.
- IEC 81001-5-1
- The standard, published in December 2021, for building security into health software. It follows IEC 62304's scope and clause order and takes its security activities from IEC 62443-4-1, adding security risk management, threat modeling, security testing and vulnerability handling to an existing process; the FDA names it as one way to meet its secure product development framework.
- IHE
- Integrating the Healthcare Enterprise: an initiative that writes profiles telling vendors how to combine standards for a task. Its profiles for AI results and AI workflow, which define how results are encoded in DICOM and how analysis requests are sent and managed, were still at trial implementation in October 2026.
- Intended use
- What the maker says software is for, in its labeling and claims. It, not anything in the code, makes software a medical device, sets its class in each market and fixes what the evidence must show; IDx-DR's covers adults with diabetes not diagnosed with retinopathy, imaged on the Topcon NW400.
- ISO 13485
- The international standard for the quality management systems of device makers, in its 2016 edition. The FDA's device quality rule has incorporated it by reference since 2 February 2026, and it requires makers to validate software used in the quality system and in production, in proportion to risk.
- ISO 14971
- The international standard for the risk management of medical devices, whose 2019 edition was confirmed in 2025. It traces hazards through hazardous situations to harms and gives each unacceptable risk a control.
- Large language model
- A generative AI model that produces text, abbreviated LLM by the FDA. UpDoc's maker announced in June 2026 that its insulin program, cleared in December 2025, was the first FDA clearance of software using patient-facing large language models, a claim the FDA has not confirmed; its public summary points to an interface rather than the decision.
- Leakage
- Any route by which the data that judges a model influences how it is built, inflating the reported score: the same patient on both sides of a split, preprocessing statistics computed with the test images included, duplicates in both sets, or repeated looks at the test set until it becomes a second tuning set.
- Legacy software
- Software legally on the market but without enough evidence that it was built to IEC 62304. The standard's 2015 amendment lets it stay in use after a risk analysis using field experience, a gap analysis against the standard, a plan to close the gaps that matter and a documented rationale.
- Local performance check
- Running a device on a sample of a site's own cases, silently or retrospectively, and comparing its outputs with a local reference, after acceptance has confirmed the installation, interfaces and fallback. Showing a sensitivity above 75 percent when about 87 percent is expected takes about 69 diseased cases, or 690 consecutive cases where 10 percent have the disease.
- Locked model
- A model whose parameters are fixed after training, so that the same input always gives the same output and it can be verified like any other software; a continuously learning model, by contrast, keeps adjusting them from new data in use. IDx-DR's analysis was locked before the 2017 trial, and almost every authorized AI device is locked.
- Machine-learning model
- A fixed structure chosen by engineers whose parameters are set by training on data rather than by hand. The structure is code that can be reviewed; the parameters cannot be read as rules, so evidence about the model comes from its behavior on carefully chosen data.
- MDR
- Medical Device Regulation: the EU's law for medical devices, which classifies software by Rule 11, so that almost any software informing a clinical decision needs a notified body, and makes makers state the minimum hardware, network and IT security their software needs. Its counterpart for diagnostic tests is the in vitro diagnostic regulation, IVDR.
- Medical Device Coordination Group (MDCG)
- The EU body whose guidance interprets the device regulations: on qualifying and classifying software (2019, revised June 2025), clinical evidence for software (March 2020), significant changes, cybersecurity (2019, revised 2020) and, jointly with the EU's AI board in June 2025, meeting the device rules and the AI Act in one assessment.
- Model card
- A short structured summary of a model, its data and its performance, placed in the labeling; the FDA's draft guidance on AI-enabled devices suggests one without requiring it.
- Modification protocol
- The part of a predetermined change control plan that says how each planned change will be developed, validated and implemented: how new data are collected and kept separate, how the model is retrained and evaluated, the acceptance criteria it must meet and how users are told. Its acceptance criteria, fixed before any change is made, carry the plan's weight.
- More than mild diabetic retinopathy
- The condition IDx-DR screens for: in the 2017 trial, a worse eye at level 35 or higher on the severity scale of the Early Treatment Diabetic Retinopathy Study (ETDRS), or macular edema, as graded at the Wisconsin reading center. It was present in 24 percent of the analyzable participants.
- Notified body
- An independent organization designated in the EU to assess a device's conformity before it can carry the CE mark. Under Rule 11 almost all software that informs a clinical decision needs one, and for a device that is also high-risk AI, the notified body checks the AI Act's requirements within the same assessment.
- OTS
- Off-the-shelf software: in the FDA's term, a generally available software component used by a device maker that cannot claim complete control of its life cycle. A commercial operating system is OTS, and so is a compiler or test tool that never ships in the device but must be validated for its use.
- Overfitting
- Fitting a model so closely to its training examples that it follows their noise and predicts new data worse. In the guide's example, a ten-parameter polynomial passes through all ten training points yet misses new points by 1.67 on average, where a straight line misses them by 0.78.
- PACS
- Picture archiving and communication system: the hospital system that stores images after the scanner sends them, from which an imaging model receives studies and to which it often returns its results.
- Parameter
- One of the adjustable numbers inside a model, a weight that multiplies one value before it is passed on; ResNet-50, a widely used image network of some fifty layers, has 25,557,032. Parameters are not written but found by training.
- Part 11
- 21 CFR Part 11, the FDA's rule on electronic records and signatures, in force since 1997, which requires records that are attributable, complete and protected, with audit trails.
- Pipeline
- An automated sequence that takes every submitted change, builds the software from source, runs the tests and analysis tools and packages the result, with no step done by hand. For a device maker it makes every internal build a complete record; a person still judges the unresolved anomalies and approves the release.
- Predetermined change control plan (PCCP)
- A plan, authorized with a device, that lets specified changes be made without a new submission. The FDA's final guidance for AI-enabled software, of December 2024, asks for a description of the planned modifications, a modification protocol and an impact assessment of their benefits and risks, all within the intended use.
- Predictive value
- What a device's answer means in a given population: the positive predictive value (PPV) is the share of positive answers that are right, the negative predictive value (NPV) the share of negative answers that are right. Both change with prevalence: IDx-DR's positive answers are right about 72 percent of the time at 24 percent prevalence and about 39 percent at 7 percent.
- Premarket approval (PMA)
- The US route to market for class III devices. A predetermined change control plan can be authorized through it, as through a 510(k) or a De Novo.
- Prevalence
- How common a condition is in the population tested: about 24 percent for more than mild retinopathy among the IDx-DR trial's analyzed participants, 7 percent for sepsis in the Michigan hospitalizations. It leaves sensitivity and specificity unchanged but changes predictive values, and a change in it is one kind of dataset shift.
- Quality Management System Regulation (QMSR)
- The FDA's quality rule for devices under its name since 2 February 2026, when it began to incorporate ISO 13485:2016 by reference and the agency's inspectors stopped using their old inspection technique.
- Random and systematic failure
- A random failure is one of a hardware part, which wears and fails at its own moment, so that a rate per hour describes a population; a systematic failure occurs every time the same conditions recur, in every copy of the same version, which is how software fails. Two copies of a program fed the same input fail together, so redundancy does not protect against a software fault.
- Reader study
- A study of clinicians reading cases with and without an assistive device. In the fully crossed design of the FDA's guidance, every reader reads every case in both conditions, sessions are separated by at least four weeks to avoid memory bias, and the primary measure is usually the area under the ROC curve with the aid against without it.
- Reference standard
- The best available judgment of each case's true state, against which a device is scored, often called ground truth. For IDx-DR it was the majority grade of three masked readers at the Wisconsin reading center, from four stereo photograph pairs of each retina and a macular scan; an imperfect reference lowers a perfect model's apparent accuracy and can hide errors both share.
- Regression testing
- Re-running the tests that cover what a change touched, and for anything beyond a trivial change the whole system-level suite, so that behavior verified before the change is verified again after it. Because 79 percent of the software recalls the FDA counted in 1992 to 1998 followed changes after release, it is the most important habit in maintenance.
- Release
- In agile work, a period of one to several months ending with software that could be delivered. A release to patients is a regulated event: IEC 62304 requires it to be verified as complete, its known anomalies evaluated, the released version archived and reproducible, and the release approved.
- Reproducible build
- A build in which rebuilding a past release from its recorded inputs gives a bit-for-bit identical result, so that the version tested and the version shipped can be shown to be the same.
- Responsibility agreement
- A written statement of which party, the hospital, the IT supplier or the device maker, does what across the life of a device's connection to a hospital's systems: who tests the interfaces, who approves changes on each side, who is told when a scanner or the archive is updated, and what clinicians do when the AI is unavailable.
- ROC curve
- Receiver operating characteristic curve: a plot of sensitivity against the false-positive rate at every possible threshold of a model's score, showing the trade between missed cases and false alarms. The Google retinal network's curve carried a high-specificity point, 90.3 percent sensitivity at 98.1 percent specificity, and a high-sensitivity point, 97.5 percent at 93.4 percent.
- Root of trust
- The protected start-up code and key, held in memory that cannot be rewritten, at the start of a secure boot chain. It must be in the hardware from the first unit shipped, because a device built without it cannot gain one by update.
- Rule 11
- The EU device regulation's rule for classifying software: IIa if it informs diagnostic or therapeutic decisions, IIb if those decisions could cause a serious deterioration of health or a surgical intervention, III if they could cause death or an irreversible deterioration. Monitoring software is IIa, or IIb for vital parameters whose variations could mean immediate danger; all other software is class I.
- Rule of three
- When no failures are seen in a number of independent, realistic trials, the upper bound of the 95 percent confidence interval for the failure probability is about three divided by that number. Showing a rate of 1 failure in a billion hours this way would take about three billion failure-free hours, some 342,000 years.
- SBOM
- Software bill of materials: a machine-readable list of every component in a piece of software, with supplier, name, version and dependencies, now extended to models and datasets. Required for US cyber devices since March 2023, it tells a maker or a hospital which devices contain a newly vulnerable component within hours rather than weeks.
- Section 524B
- The section of the US Food, Drug, and Cosmetic Act, in effect since 29 March 2023, that requires the maker of a cyber device to file a plan for monitoring and addressing vulnerabilities after release, including coordinated disclosure; to keep processes that make the device secure, with patches on a regular cycle and out of cycle for critical vulnerabilities; and to provide an SBOM.
- Secure boot
- A start-up chain in which each stage checks the next before handing over control. At every power-on, protected code recomputes the firmware's hash, a fingerprint that changes completely if a single bit changes, checks it against the maker's signature made with a private key, and runs the firmware only if the two agree, so altered code never runs.
- Secure product development framework
- The FDA's name for a development process with security built in, from threat modeling and security requirements to security testing and the handling of vulnerabilities after release; its 2026 guidance names IEC 81001-5-1 and IEC 62443-4-1 as examples that can satisfy it.
- Segregation
- In IEC 62304, any mechanism that prevents one software item from negatively affecting another: a separate processor, operating-system memory protection, or checks on data crossing a boundary. Demonstrated segregation lets the strictest class apply only to the dangerous code, in the guide's pump example 8,000 of 200,000 lines.
- Sensitivity
- The share of people with the condition whom a device flags: 173 of 198, 87.4 percent, in the IDx-DR trial. Its precision depends on the number of diseased patients tested, not on the total.
- Significant change
- Under the EU's device rules, a change that brings in the notified body. The coordination group's software chart counts a new or modified architecture, database structure or algorithm, a new operating system or component or a major change to one, and a new feature, interoperability channel or user interface as significant, but not safety-neutral bug fixes, security updates or cosmetic changes.
- Silent trial
- Running a model on live clinical data with its outputs hidden from clinicians, so that its local performance is measured before anyone acts on it. At the Hospital for Sick Children in Toronto one took a kidney-ultrasound model from an area under the curve of 0.90 to 0.50; reprocessing the live images, which arrived as unprocessed PNG files rather than processed JPEGs, restored most of the loss.
- Software as a medical device
- Software that is a medical device on its own, running on ordinary computers, phones or cloud servers rather than inside hardware; IDx-DR's analysis program is an example. India's 2026 guidance drops the term in favor of standalone software.
- Software requirement
- A statement of what the software must do, or must never do, written so that a test, an inspection or an analysis can show whether it holds. A display that is easy to read is not one; a dose shown with its unit in characters at least 5 mm high until the user confirms it is.
- Software safety class
- IEC 62304's class A, B or C, set by the worst harm software could contribute to once risk controls outside it are counted: A where no unacceptable risk remains, B for non-serious injury, C for death or serious injury. Until classified, software is treated as class C; the class sets how much process the standard demands, and only controls outside the software can lower it.
- Software system, item and unit
- IEC 62304's division of a program: the software system is divided into software items, any identifiable parts, and these in turn until they reach software units, items not divided further. IDx-DR has three separately versioned items: the client, the service that passes images and results, and the analysis.
- SOUP
- Software of unknown provenance: in IEC 62304, a software item already developed and generally available that was not developed for the device, or one developed earlier without adequate records. For each, a maker states the requirements the device places on it and its exact version, evaluates its published anomalies and keeps it under configuration control.
- Special controls
- The requirements the FDA sets for a device type when it creates one, which every later device of that type must meet. IDx-DR's included a protocol on which changes could significantly affect safety or effectiveness and a cybersecurity vulnerability and management process.
- Specificity
- The share of people without the condition whom a device clears: 556 of 621, 89.5 percent, in the IDx-DR trial, and about 79 percent in studies of young people outside its label.
- Subgroup analysis
- Reporting a device's performance separately for groups within the intended population, such as by sex, age, ethnicity, skin tone, site or equipment, each with its confidence interval, with the groups named in the protocol before the data are seen. Dermatology models that had scored 0.88 to 0.94 on their own test sets fell as low as chance on the darkest skin.
- Substantial modification
- Under the EU's AI Act, a change to an AI system after it is placed on the market that was not foreseen in its initial conformity assessment and affects its compliance or modifies its intended purpose; a high-risk system that undergoes one needs a new conformity assessment. Changes predetermined and documented at the initial assessment are not substantial modifications.
- Threat model
- A list of a device's entry points, the assets behind them, such as dose settings, patient data and the software itself, and the adversaries who might want them, with what an attacker could do on each path and what would stop it. The FDA expects it to be part of the design, updated as the architecture changes.
- Threshold
- The value that turns a model's score into a decision, refer above it and not below. Moving it trades missed cases against false alarms; the point chosen, the operating point, is fixed from the clinical costs before the test set is opened.
- Traceability
- The links from each hazard to the requirements that control it, from each requirement to the design that implements it and the tests that verify it, and from each test to its result, so that a reviewer can follow any risk to its control and its evidence.
- Training
- Finding a model's parameters from examples whose right answers are known: the model scores a batch, a measure of its error called the loss is computed, and every parameter is nudged in the direction that would have made the loss smaller, over millions of batches.
- Training, tuning and test sets
- The three separate bodies of data behind a model: the training set its parameters are fitted to, the tuning set used for choices training does not make (which engineers usually call the validation set), and the test set, held back and used once to estimate performance on the next patient. Only the test set, kept independent and split by patient, judges the model fairly.
- Unit, integration and system testing
- The three levels of verification by test: unit testing checks each smallest piece in isolation against its detailed design, integration testing checks that units and items work together across their boundaries, and system testing checks the complete software, usually on the real hardware, against the software requirements.
- Unresolved anomaly
- A known defect that software ships with. Submissions list each one with an assessment of its effect on safety and effectiveness, and a release may carry it only if a person has judged that it leaves no unacceptable risk.
- Valid clinical association
- The first of three components of clinical evidence for software in the EU and India: the software's output is associated with the clinical condition it targets. The others are technical performance, producing the intended output accurately and reliably from the input, and clinical performance, an output clinically relevant to the intended purpose.
- Validation
- Confirmation, by examination and objective evidence, that software specifications conform to user needs and intended uses; it asks whether the specification was right and reaches to usability and clinical evidence. In machine learning the same word names the tuning set, and a submission that uses both senses must say which it means.
- Verification
- Objective evidence that the outputs of a stage of development meet the requirements set for that stage, that is, that the software was built as specified. It runs at unit, integration and system level, with reviews, code inspection and static analysis finding faults without running the code.
- VEX
- Vulnerability Exploitability eXchange: a statement of whether a published vulnerability affects a specific product, with one of four values, not affected, affected, fixed or under investigation. An SBOM says which components a device contains; a VEX says whether a weakness in one of them can be exploited there.
- Watchdog
- A timer, often a separate circuit, that the software must reset at regular intervals; if the program hangs, the timer runs out and forces the device into a safe state. A program that runs on time while computing a wrong dose keeps resetting it, so a watchdog cannot catch a wrong output.
Sources.
Each source is listed under the chapter whose text first relies on it, with the opening scene and the front matter first. Company documents, papers by a company's founders or staff, and secondary reports are marked as such. Regulatory status, standards editions and product facts are as of October 2026.
- Opening — Abràmoff M.D., Lavin P.T., Birch M., Shah N., Folk J.C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digit Med 1, 39 (2018). doi:10.1038/s41746-018-0040-6 (first author founded the company; trial funded by IDx)
- Opening — US Food and Drug Administration. De Novo DEN180001, IDx-DR (IDx, LLC): database record and decision summary; received 12 Jan 2018, granted 11 Apr 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN180001; accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf
- Opening — Digital Diagnostics. IDx rebrands to lead AI health care revolution (5 Oct 2020). digitaldiagnostics.com/idx-rebrands-to-lead-ai-health-care-revolution/ (manufacturer)
- Opening — US Food and Drug Administration. FDA permits marketing of artificial intelligence-based device to detect certain diabetes-related eye problems. Press release, 11 Apr 2018. drugdiscoverytrends.com/fda-permits-marketing-of-ai-based-device-to-detect-certain-diabetes-related-eye-problems/ (read via the drugdiscoverytrends.com reprint; the fda.gov page returned 404)
- Opening — US Food and Drug Administration. General Principles of Software Validation; Final Guidance for Industry and FDA Staff. 11 Jan 2002 (Section 6 since superseded by the computer software assurance guidance). fda.gov/media/73141/download
- Ch. 1 — US Code of Federal Regulations (eCFR). 21 CFR 886.1100, Retinal diagnostic software device (codified at 87 FR 3205, 21 Jan 2022). ecfr.gov/current/title-21/chapter-I/subchapter-H/part-886/subpart-B/section-886.1100
- Ch. 1 — Central Drugs Standard Control Organisation, India. Guidance document on Medical Device Software under MDR-2017. Doc No. CDSCO/MD/GD/MDSW/01/2026. 21 Jul 2026. cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/Guidance-document-on-Medical-Device-Software-under-MDR-2017.pdf (text read to section 12.5.2)
- Ch. 1 — US Food and Drug Administration. Changes to Existing Medical Software Policies Resulting from Section 3060 of the 21st Century Cures Act. Final guidance, docket FDA-2017-D-6294. 27 Sep 2019. fda.gov/media/109622/download
- Ch. 1 — US Food and Drug Administration. General Wellness: Policy for Low Risk Devices. Final guidance, docket FDA-2014-N-1039. 6 Jan 2026 (supersedes the version of 27 Sep 2019). fda.gov/media/90652/download
- Ch. 1 — US Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final guidance, docket FDA-2022-D-2628. 4 Dec 2024, reissued 18 Aug 2025. fda.gov/media/166704/download
- Ch. 1 — Akin Gump Strauss Hauer & Feld. IMDRF releases international framework for regulating device software (Oct 2014), on IMDRF/SaMD WG/N12 FINAL:2014, Software as a Medical Device: Possible Framework for Risk Categorization and Corresponding Considerations (18 Sep 2014). akingump.com/en/insights/alerts/imdrf-releases-international-framework-for-regulating-device (secondary; IMDRF page not opened)
- Ch. 2 — US Food and Drug Administration. Ventilator Software Correction: Hamilton Medical Issues Correction for HAMILTON-C6 Medical Ventilators to Address Risk of Failed Ventilation Restart. Recall notice, posted 11 Jul 2024. fda.gov/medical-devices/medical-device-recalls-and-early-alerts/ventilator-software-correction-hamilton-medical-issues-correction-hamilton-c6-medical-ventilators
- Ch. 2 — US Food and Drug Administration. Medical Device Recalls: Class 1 recall Z-2020-2024, ventilator HAMILTON-C6 (Hamilton Medical), software versions 1.1.4–1.1.6; initiated 15 May 2024, posted 18 Jun 2024. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?id=207956
- Ch. 2 — Wallace D.R., Kuhn D.R. Failure modes in medical device software: an analysis of 15 years of recall data. Int J Reliab Qual Saf Eng 8, 351–371 (2001). doi:10.1142/S021853930100058X (read via the NIST copy, tsapps.nist.gov/publication/get_pdf.cfm?pub_id=917180)
- Ch. 2 — Wallace D.R., Kuhn D.R. Lessons from 342 medical device failures (NIST; conference version, c. 1999) (read via inf.ed.ac.uk/teaching/courses/seoc/2006_2007/resources/CS_342failures.pdf)
- Ch. 2 — US Food and Drug Administration. Tandem Diabetes Care, Inc. Recalls Version 2.7 of the Apple iOS t:connect Mobile App… (used with the t:slim X2 insulin pump). Recall notice, posted 28 Aug 2024. fda.gov/medical-devices/medical-device-recalls-and-early-alerts/tandem-diabetes-care-inc-recalls-version-27-apple-ios-tconnect-mobile-app-used-conjunction-tslim-x2
- Ch. 2 — US Food and Drug Administration. Medical Device Recalls: Class 1 recall Z-1609-2024, t:connect mobile app (Tandem Diabetes Care), root cause recorded as software design change; initiated Mar 2024, posted 6 May 2024. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?id=206914
- Ch. 3 — Dijkstra E.W. The humble programmer (ACM Turing Lecture 1972; EWD340). Commun ACM 15, 859–866 (1972). doi:10.1145/355604.361591 (read as the EWD340 transcription, cs.utexas.edu/~EWD/transcriptions/EWD03xx/EWD340.html; journal details not checked)
- Ch. 3 — Eypasch E., Lefering R., Kum C.K., Troidl H. Probability of adverse events that have not yet occurred: a statistical reminder. BMJ 311, 619 (1995). doi:10.1136/bmj.311.7005.619 (abstract read)
- Ch. 3 — Butler R.W., Finelli G.B. The infeasibility of quantifying the reliability of life-critical real-time software. IEEE Trans Softw Eng 19, 3–12 (1993). doi:10.1109/32.210303 (read as the NASA preprint, ntrs.nasa.gov/api/citations/20040139817/downloads/20040139817.pdf)
- Ch. 3 — NASA Science. Universe overview (age of the universe, about 13.8 billion years). science.nasa.gov/universe/overview/
- Ch. 4 — IEC. IEC 62304:2006, Medical device software – Software life cycle processes, ed. 1.0 (9 May 2006). IEC Webstore page, webstore.iec.ch/en/publication/6792
- Ch. 4 — IEC. IEC 62304:2006+AMD1:2015 CSV, Medical device software – Software life cycle processes, consolidated ed. 1.1 (26 Jun 2015; stability date 2028). IEC Webstore page, webstore.iec.ch/en/publication/22794; preview pages, including the introduction to Amendment 1, read via webstore.ansi.org
- Ch. 4 — US Food and Drug Administration. Recognized Consensus Standards database: IEC 62304 ed. 1.1 2015-06 consolidated (and ANSI/AAMI IEC 62304:2006/A1:2016), recognition no. 13-79, complete standard; entry 14 Jan 2019. accessdata.fda.gov/scripts/cdrh/cfdocs/cfStandards/detail.cfm?standard__identification_no=38829
- Ch. 4 — IEC. Project records for IEC 62304 ED2, Health software – Software life cycle processes: iec:proj:122433 (stage 30.20, CD study/ballot initiated, 7 Aug 2026; next stage due 27 Nov 2026) and the earlier project iec:proj:23605 (deleted 7 Mar 2023); with iec:proj:11630 and iec:proj:21252 for editions 1.0 and 1.1. iss.rs/en/project/show/iec:proj:122433 (read via iss.rs, the Institute for Standardization of Serbia’s mirror of IEC project data; the iec.ch page was blocked)
- Ch. 4 — VDE. IEC 62304 Edition 2: Änderungen (VDE Health Fachinformation, 22 Jul 2026). vde.com/iec-62304-edition-2-aenderungen (secondary)
- Ch. 4 — Johner Institute. IEC 62304 2nd edition: all areas of application and changes (8 Jul 2026). blog.johner-institute.com/iec-62304-medical-software/iec-62304-2nd-edition-all-areas-of-application-and-changes/ (secondary)
- Ch. 4 — ISO. ISO 14971:2019, Medical devices – Application of risk management to medical devices, 3rd ed. (Dec 2019; confirmed 7 Mar 2025), catalogue page. iso.org/standard/72704.html
- Ch. 4 — International Medical Device Regulators Forum, SaMD Working Group. Characterization Considerations for Medical Device Software and Software-Specific Risk. IMDRF/SaMD WG/N81 FINAL:2025. 27 Jan 2025. imdrf.org/sites/default/files/2025-01/IMDRF_SaMD%20WG_Software-Specific%20Risk_N81%20Final_0.pdf
- Ch. 4 — Wikipedia. Software safety classification (web page quoting IEC 62304:2006+A1:2015, accessed 8 Oct 2026). en.wikipedia.org/wiki/Software_safety_classification (secondary)
- Ch. 4 — LDRA Ltd. Ease the Heartache of Medical Device Software Certification: Achieving Cost-effective Compliance with IEC 62304 – Amendment 1:2015. Technical briefing v2 (Apr 2018). ldra.com/wp-content/uploads/ldra/IEC_62304_Technical_Briefing_v2_04_18.pdf (secondary)
- Ch. 4 — Johner Institute. Software safety classes according to IEC 62304 (rewritten 15 Oct 2025). blog.johner-institute.com/iec-62304-medical-software/safety-class-iec-62304/ (secondary)
- Ch. 4 — US Food and Drug Administration. Content of Premarket Submissions for Device Software Functions. Final guidance, docket FDA-2021-D-0775. 14 Jun 2023 (supersedes the guidance of 11 May 2005). fda.gov/media/153781/download
- Ch. 4 — Matrix Requirements. IEC 62304:2015 impact on class A software (blog). matrixone.health/blog/iec-62304-2015-impact-on-class-a-software (secondary; checked in the subject review)
- Ch. 5 — US Food and Drug Administration. Infusion Pumps Total Product Life Cycle: Guidance for Industry and FDA Staff. Final guidance. 2 Dec 2014. fda.gov/media/78369/download
- Ch. 5 — US Food and Drug Administration. 510(k) K203629, IDx-DR (Digital Diagnostics, Inc.): database record and 510(k) summary; received 11 Dec 2020, decision 10 Jun 2021. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K203629; accessdata.fda.gov/cdrh_docs/pdf20/K203629.pdf
- Ch. 5 — US Food and Drug Administration. 510(k) K213037, IDx-DR v2.3 (Digital Diagnostics, Inc.): database record and 510(k) summary; received 21 Sep 2021, decision 17 Jun 2022. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K213037; accessdata.fda.gov/cdrh_docs/pdf21/K213037.pdf
- Ch. 6 — Microsoft. Windows XP, product lifecycle (Microsoft Learn, accessed 8 Oct 2026). learn.microsoft.com/en-us/lifecycle/products/windows-xp (manufacturer)
- Ch. 6 — National Audit Office (UK). Investigation: WannaCry cyber attack and the NHS. HC 414. 27 Oct 2017. nao.org.uk/reports/investigation-wannacry-cyber-attack-and-the-nhs/ (summary page read)
- Ch. 6 — Dworetzky T. ‘Ransomware’ attack hit U.S. medical devices, too. DOTmed News (18 May 2017). dotmed.com/legal/print/story.html?nid=37359 (secondary, reporting Bayer and Siemens Healthineers)
- Ch. 6 — Palo Alto Networks Unit 42. 2020 Unit 42 IoT Threat Report (10 Mar 2020). unit42.paloaltonetworks.com/iot-threat-report-2020/ (manufacturer)
- Ch. 6 — US Food and Drug Administration. Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions. Guidance for Industry and FDA Staff, docket FDA-2021-D-1158. 3 Feb 2026 (supersedes the version of 27 Jun 2025). fda.gov/media/119933/download
- Ch. 6 — Johner Institute. Off-the-shelf software (OTS) versus SOUP (web page, undated; quoting IEC 62304:2006+A1:2015). blog.johner-institute.com/iec-62304-medical-software/off-the-shelf-software-ots-versus-soup/ (secondary)
- Ch. 6 — US Food and Drug Administration. Off-The-Shelf Software Use in Medical Devices. Final guidance. 11 Aug 2023 (supersedes the edition of 27 Sep 2019; first issued 9 Sep 1999). fda.gov/media/71794/download
- Ch. 6 — Black Duck Software. 2026 Open Source Security and Risk Analysis (OSSRA) Report (Mar 2026). blackduck.com/resources/analyst-reports/open-source-security-risk-analysis.html (manufacturer)
- Ch. 6 — OWASP CycloneDX. CycloneDX v1.5 released (26 Jun 2023). cyclonedx.org/news/cyclonedx-v1.5-released/
- Ch. 6 — National Telecommunications and Information Administration (US Department of Commerce). The Minimum Elements For a Software Bill of Materials (SBOM). 12 Jul 2021. ntia.gov/files/ntia/publications/sbom_minimum_elements_report.pdf
- Ch. 6 — Medcrypt. Blog post on CISA’s 2026 Minimum Elements for a Software Bill of Materials (29 Jul 2026). medcrypt.com/blog/cisa-sbom-minimum-elements-2026-update (secondary; written by a device-security vendor; CISA’s own 2026 PDF was blocked and not opened)
- Ch. 6 — devops.com. CISA’s 2026 SBOM guidance adds hash requirements and AI coverage (31 Jul 2026). devops.com/cisas-2026-sbom-guidance-adds-hash-requirements-and-ai-coverage/ (secondary)
- Ch. 6 — HSToday. CISA updates software bill of materials guidance to strengthen supply chain security (2026). hstoday.us/subject-matter-areas/cybersecurity/cisa-updates-software-bill-of-materials-guidance-to-strengthen-supply-chain-security/ (secondary)
- Ch. 6 — ISO/IEC. ISO/IEC 5962:2021, Information technology — SPDX® Specification V2.2.1, ed. 1 (Aug 2021; to be revised), catalogue page. iso.org/standard/81870.html
- Ch. 6 — Ecma International. ECMA-424, CycloneDX Bill of materials specification, 1st ed. (Jun 2024) and 2nd ed. (Dec 2025). ecma-international.org/publications-and-standards/standards/ecma-424/
- Ch. 6 — National Telecommunications and Information Administration. Vulnerability-Exploitability eXchange (VEX) – An Overview (27 Sep 2021). ntia.gov/files/ntia/publications/vex_one-page_summary.pdf
- Ch. 6 — United States Code. 21 U.S.C. 360n-2, Ensuring cybersecurity of devices (FD&C Act section 524B), added by Pub. L. 117-328, div. FF, title III, sec. 3305, 29 Dec 2022; effective 29 Mar 2023. law.cornell.edu/uscode/text/21/360n-2 (read via Cornell LII)
- Ch. 6 — US Food and Drug Administration. Cybersecurity in Medical Devices Frequently Asked Questions (FAQs) (content current as of 26 Jun 2025). fda.gov/medical-devices/digital-health-center-excellence/cybersecurity-medical-devices-frequently-asked-questions-faqs
- Ch. 7 — US Food and Drug Administration. Medical Device Recalls: Class 2 recall Z-1233-2023, Monaco RTP System (Elekta), builds 5.11.00–5.11.03, root cause recorded as software design; initiated 28 Feb 2023, posted 8 Mar 2023. accessdata.fda.gov/scripts/cdrh/cfdocs/cfRes/res.cfm?id=198827
- Ch. 8 — AAMI. AAMI TIR45:2012 and AAMI TIR45:2023, Guidance on the use of AGILE practices in the development of medical device software (not opened; editions as listed in the FDA Recognized Consensus Standards database)
- Ch. 8 — US Food and Drug Administration. Recognized Consensus Standards database: AAMI TIR45:2012, recognition no. 13-36 (entry 15 Jan 2013), and AAMI TIR45:2023, recognition no. 13-143 (entry 26 May 2025; declarations of conformity to 13-36 accepted until 2 Jul 2028). accessdata.fda.gov/scripts/cdrh/cfdocs/cfStandards/detail.cfm?standard__identification_no=46298; accessdata.fda.gov/scripts/cdrh/cfdocs/cfstandards/detail.cfm?standard__identification_no=46295
- Ch. 8 — Johner Institute. TIR 45: Agile software development (2016, updated for the 2023 edition). blog.johner-institute.com/iec-62304-medical-software/tir-45-agile-software-development/ (secondary)
- Ch. 8 — AAMI. Key Updates: AAMI TIR45:2023 – Guidance on Agile Practices (training page, undated). aami.org/training/training-suites/software-cybersecurity/key-updates-aami-tir45-2023-guidance-on-agile-practices (publisher’s training page)
- Ch. 8 — ISPE. ISPE GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition) (Jul 2022), product page. ispe.org/publications/guidance-documents/gamp-5-guide-2nd-edition
- Ch. 8 — Wyn S., Clark C. What you need to know about GAMP 5 Guide, 2nd Edition. Pharmaceutical Engineering (Jan/Feb 2023). ispe.org/pharmaceutical-engineering/january-february-2023/what-you-need-know-about-gampr-5-guide-2nd-edition
- Ch. 8 — OpenRegulatory. IEC 62304:2006 mapping of requirements to documents (9 Jun 2022). openregulatory.com/document_templates/iec-623042006-mapping-of-requirements-to-documents (secondary; clause titles of IEC 62304, standard not opened)
- Ch. 8 — OpenRegulatory. Doing a software release in compliance with IEC 62304 (24 May 2023, updated 1 Oct 2024). openregulatory.com/articles/software-release-iec-62304 (secondary)
- Ch. 9 — Google Cloud DORA. Accelerate State of DevOps Report 2024 (22 Oct 2024). dora.dev/research/2024/dora-report/2024-dora-accelerate-state-of-devops-report.pdf (manufacturer)
- Ch. 9 — US Food and Drug Administration. Quality Management System Regulation (QMSR) web page; and Medical Devices; Quality System Regulation Amendments, final rule, FR Doc. 2024-01709 (2 Feb 2024; effective 2 Feb 2026). fda.gov/medical-devices/postmarket-requirements-devices/quality-management-system-regulation-qmsr
- Ch. 9 — ISO. ISO/TR 80002-2:2017, Medical device software — Part 2: Validation of software for medical device quality systems, ed. 1 (13 Jun 2017), catalogue page. iso.org/standard/60044.html
- Ch. 9 — US Food and Drug Administration. Computer Software Assurance for Production and Quality System Software. Final guidance, docket FDA-2022-D-0795. 24 Sep 2025. fda.gov/regulatory-information/search-fda-guidance-documents/computer-software-assurance-production-and-quality-system-software
- Ch. 9 — US Food and Drug Administration. Computer Software Assurance for Production and Quality Management System Software. Final guidance, docket FDA-2022-D-0795. 3 Feb 2026 (supersedes the guidance of 24 Sep 2025 and Section 6 of the 2002 software validation guidance). fda.gov/media/188844/download
- Ch. 9 — Johner Institute. IT security for legacy devices (web page, undated; quoting the IEC 62304 Amendment 1 definition of legacy software). blog.johner-institute.com/iec-62304-medical-software/it-security-for-legacy-devices/ (secondary)
- Ch. 9 — Spyro-soft. How to implement a legacy software gap analysis as required by IEC 62304:2015 (blog, undated). spyro-soft.com/blog/how-to-implement-a-legacy-software-gap-analysis-as-required-by-iec-62304-2015 (secondary)
- Ch. 10 — Gulshan V., Peng L., Coram M., et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 316, 2402–2410 (2016). doi:10.1001/jama.2016.17216 (authors include company staff)
- Ch. 10 — PyTorch. torchvision.models.resnet50, documentation (stable release, accessed 8 Oct 2026). docs.pytorch.org/vision/stable/models/generated/torchvision.models.resnet50.html
- Ch. 10 — US Food and Drug Administration. De Novo DEN190040, Caption Guidance (Bay Labs, Inc.): database record and decision summary; received 27 Aug 2019, granted 7 Feb 2020. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN190040; accessdata.fda.gov/cdrh_docs/reviews/DEN190040.pdf
- Ch. 10 — US Food and Drug Administration. 510(k) K201992, Caption Guidance (Caption Health, Inc.): 510(k) summary; decision 18 Sep 2020. accessdata.fda.gov/cdrh_docs/pdf20/K201992.pdf
- Ch. 10 — PIC/S and EMA GMP/GDP Inspectors Working Group. Draft Annex 22: Artificial Intelligence. Consultation draft (Jul 2025). picscheme.org/docview/9715
- Ch. 10 — US Food and Drug Administration. Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) – Discussion Paper and Request for Feedback. 2 Apr 2019. fda.gov/media/122535/download
- Ch. 11 — Zech J.R., Badgeley M.A., Liu M., Costa A.B., Titano J.J., Oermann E.K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLOS Med 15, e1002683 (2018). doi:10.1371/journal.pmed.1002683 (authors include company staff)
- Ch. 11 — International Medical Device Regulators Forum, AI/ML-enabled Working Group. Good machine learning practice for medical device development: Guiding principles. IMDRF/AIML WG/N88 FINAL:2025. 27 Jan 2025. imdrf.org/sites/default/files/2025-02/IMDRF_AIML%20WG_GMLP_N88%20Final.pdf
- Ch. 12 — Krause J., Gulshan V., Rahimy E., Karth P., Widner K., Corrado G.S., et al. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology 125, 1264–1272 (2018). doi:10.1016/j.ophtha.2018.01.034 (read as arXiv:1710.01711v3; journal version not opened) (authors include company staff)
- Ch. 12 — Medicines and Healthcare products Regulatory Agency, US Food and Drug Administration, Health Canada. Good machine learning practice for medical device development: guiding principles (27 Oct 2021). gov.uk/government/publications/good-machine-learning-practice-for-medical-device-development-guiding-principles
- Ch. 12 — US Food and Drug Administration. De Novo DEN170073, ContaCT (Viz.ai, Inc.): database record and decision summary; received 29 Sep 2017, granted 13 Feb 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN170073; accessdata.fda.gov/cdrh_docs/reviews/DEN170073.pdf
- Ch. 12 — Daneshjou R., et al. Diverse Dermatology Images (DDI) dataset, project page (accessed 8 Oct 2026). ddi-dataset.github.io/
- Ch. 13 — NIST/SEMATECH. e-Handbook of Statistical Methods, section 7.2.4.1, Confidence intervals (web page, undated). itl.nist.gov/div898/handbook/prc/section2/prc241.htm
- Ch. 13 — MedCalc Software. Sensitivity and specificity (MedCalc manual, web page, undated). medcalc.org/en/book/sensitivity-specificity.php (secondary)
- Ch. 13 — Wong A., Otles E., Donnelly J.P., Krumm A., McCullough J., DeTroyer-Cooley O., et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med 181, 1065–1070 (2021). doi:10.1001/jamainternmed.2021.2626
- Ch. 14 — Wolf R.M., Channa R., Liu T.Y.A., et al. Autonomous artificial intelligence increases screening and follow-up for diabetic retinopathy in youth: the ACCESS randomized control trial. Nat Commun 15, 421 (2024). doi:10.1038/s41467-023-44676-z (authors include the company’s founder; youth accuracy of the earlier SEE study as reported in this paper)
- Ch. 14 — Heaven W.D. Google’s medical AI was super accurate in a lab. Real life was a different story. MIT Technology Review (27 Apr 2020). technologyreview.com/2020/04/27/1000658/google-medical-ai-accurate-lab-real-life-clinic-covid-diabetes-retina-disease/ (secondary)
- Ch. 14 — Beede E. Healthcare AI systems that put people at the center. Google Keyword blog (25 Apr 2020). blog.google/technology/health/healthcare-ai-systems-put-people-center/ (manufacturer)
- Ch. 14 — Beede E., Baylor E., Hersch F., Iurchenko A., Wilcox L., Ruamviboonsuk P., et al. A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (2020). research.google/pubs/a-human-centered-evaluation-of-a-deep-learning-system-deployed-in-clinics-for-the-detection-of-diabetic-retinopathy/ (abstract read) (authors include company staff)
- Ch. 14 — Ruamviboonsuk P., Tiwari R., Sayres R., Nganthavee V., Hemarat K., Kongprayoon A., et al. Real-time diabetic retinopathy screening by deep learning in a multisite national screening programme: a prospective interventional cohort study. Lancet Digit Health 4, e235–e244 (2022). doi:10.1016/S2589-7500(22)00017-6 (abstract read, via DOAJ) (authors include company staff; funded by Google and Rajavithi Hospital)
- Ch. 15 — Daneshjou R., et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. arXiv:2203.08807 v1 (Mar 2022). arxiv.org/pdf/2203.08807 (preprint)
- Ch. 15 — Daneshjou R., Vodrahalli K., Novoa R.A., et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv 8, eabq6147 (2022). doi:10.1126/sciadv.abq6147 (abstract read)
- Ch. 15 — Seyyed-Kalantari L., Zhang H., McDermott M.B.A., Chen I.Y., Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med 27, 2176–2182 (2021). doi:10.1038/s41591-021-01595-0 (text read; tables not accessible)
- Ch. 15 — US Food and Drug Administration. Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests. Final guidance. 13 Mar 2007. fda.gov/media/71147/download
- Ch. 15 — Emergo by UL. India CDSCO finalizes guidance on medical device software (30 Jul 2026). emergobyul.com/news/india-cdsco-finalizes-guidance-medical-device-software (secondary)
- Ch. 15 — Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act). OJ L (12 Jul 2024). artificialintelligenceact.eu/article/3/ (Articles 3, 4, 6, 8–15, 26, 43 and 113, as amended, read via artificialintelligenceact.eu, a transcription of the OJ text)
- Ch. 15 — Emergo by UL. Decoding MDCG 2025-6: interplay between the MDR/IVDR and the AI Act (24 Jun 2025). emergobyul.com/news/decoding-mdcg-2025-6-interplay-between-mdrivdr-and-aia (secondary)
- Ch. 16 — Dratsch T., Chen X., Rezazade Mehrizi M., Kloeckner R., Mähringer-Kunz A., Püsken M., et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology 307, e222176 (2023). doi:10.1148/radiol.222176 (abstract read)
- Ch. 16 — US Food and Drug Administration. Clinical Decision Support Software. Final guidance, docket FDA-2017-D-6569. 6 Jan 2026, corrected 29 Jan 2026 (supersedes the September 2022 version); with the guidance web page. fda.gov/media/109618/download; fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software
- Ch. 16 — US Food and Drug Administration. Clinical Performance Assessment: Considerations for Computer-Assisted Detection Devices Applied to Radiology Images and Radiology Device Data in Premarket Notification (510(k)) Submissions. Final guidance, docket FDA-2009-D-0503. 28 Sep 2022. fda.gov/media/77642/download
- Ch. 16 — US Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles (13 Jun 2024). fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles
- Ch. 17 — Digital Diagnostics. LumineticsCore product page (last modified 5 Aug 2026; accessed 8 Oct 2026). digitaldiagnostics.com/idx-dr/ (manufacturer)
- Ch. 18 — NEMA, DICOM Standards Committee. DICOM PS3.6 2026d, Data Dictionary, chapter 6: Registry of DICOM Data Elements. dicom.nema.org/medical/dicom/current/output/chtml/part06/chapter_6.html
- Ch. 18 — HL7 International. HL7 Version 2 Product Suite, product brief (web page, undated). hl7.org/implement/standards/product_brief.cfm?product_id=185
- Ch. 18 — HL7 International. FHIR Publication (Version) History (web page, accessed 8 Oct 2026). hl7.org/fhir/history.html
- Ch. 18 — HL7 International. FHIR R5 (v5.0.0), Resource (26 Mar 2023). hl7.org/fhir/R5/resource.html
- Ch. 18 — HL7 International. FHIR v6.0.0-ballot5, CodeSystem FHIR-version (page generated 17 Jul 2026). hl7.org/fhir/6.0.0-ballot5/codesystem-FHIR-version.html
- Ch. 18 — IHE Radiology Technical Committee. IHE Radiology Technical Framework Supplement: AI Results (AIR), Rev. 1.3, Trial Implementation. 8 Aug 2025. ihe.net/uploadedFiles/Documents/Radiology/IHE_RAD_Suppl_AIR.pdf
- Ch. 18 — IHE Radiology Technical Committee. IHE Radiology Technical Framework Supplement: AI Workflow for Imaging (AIW-I), Rev. 1.1, Trial Implementation. 6 Aug 2020. ihe.net/uploadedFiles/Documents/Radiology/IHE_RAD_Suppl_AIW-I.pdf
- Ch. 18 — Martinez-Gutierrez J.C., Kim Y., Salazar-Marioni S., et al. Automated large vessel occlusion detection software and thrombectomy treatment times: a cluster randomized clinical trial. JAMA Neurol 80, 1182–1190 (2023). doi:10.1001/jamaneurol.2023.3206 (abstract read, via scholars.aku.edu)
- Ch. 18 — UTHealth Houston, McWilliams School of Biomedical Informatics. News release on the study of artificial intelligence software and endovascular thrombectomy treatment times (2023). sbmi.uth.edu/news/story/uthealth-houston-study-artificial-intelligence-software-improves-endovascular-thrombectomy-treatment-times-for-stroke-patients (secondary)
- Ch. 18 — Medical Device Coordination Group. MDCG 2019-16 rev.1, Guidance on Cybersecurity for medical devices. Dec 2019, rev.1 Jul 2020 (quoting Regulation (EU) 2017/745, Annex I, sections 17.2 and 17.4). health.ec.europa.eu/system/files/2022-01/md_cybersecurity_en.pdf
- Ch. 18 — IEC. IEC 80001-1:2021, Application of risk management for IT-networks incorporating medical devices – Part 1: Safety, effectiveness and security in the implementation and use of connected medical devices or connected health software, ed. 2.0 (21 Sep 2021). IEC Webstore page, webstore.iec.ch/en/publication/34263; ISO catalogue page (stage 90.92, to be revised, 4 Feb 2026), iso.org/standard/72026.html
- Ch. 18 — ISO. ISO/TR 80001-2-6:2014, Application of risk management for IT-networks incorporating medical devices – Part 2-6: Application guidance – Guidance for responsibility agreements (catalogue entry read via implementer.digitalhealth.gov.au, Australian Digital Health Agency)
- Ch. 19 — Brady A.P., Allen B., Chong J., Kotter E., Kottler N., Mongan J., et al. Developing, purchasing, implementing and monitoring AI tools in radiology: practical considerations. A multi-society statement from the ACR, CAR, ESR, RANZCR & RSNA. Can Assoc Radiol J 75, 226–244 (2024). doi:10.1177/08465371231222229
- Ch. 19 — Kwong J.C.C., Erdman L., Khondker A., Skreta M., Goldenberg A., McCradden M.D., et al. The silent trial: the bridge between bench-to-bedside clinical AI applications. Front Digit Health 4, 929508 (2022). doi:10.3389/fdgth.2022.929508
- Ch. 19 — IntuitionLabs. GAMP 5 categories explained (16 Oct 2025, updated 8 Aug 2026). intuitionlabs.ai/articles/gamp-5-categories-explained (secondary)
- Ch. 19 — US Code of Federal Regulations (eCFR). 21 CFR Part 11, Electronic Records; Electronic Signatures (62 FR 13464, 20 Mar 1997). ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- Ch. 19 — Blumenthal R., Erdmann N., Heitmann M., Lemettinen A.-L., Stockton B.M. Machine learning risk and control framework. Pharmaceutical Engineering (Jan/Feb 2024). ispe.org/pharmaceutical-engineering/january-february-2024/machine-learning-risk-and-control-framework
- Ch. 19 — ISPE. ISPE GAMP Guide: Artificial Intelligence (Jul 2025), product page. ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence
- Ch. 19 — Stockton B., Staib E., Heitmann M. New GAMP Guide addresses challenges posed by AI-enabled computerized systems. Pharmaceutical Engineering (Sep/Oct 2025). ispe.org/pharmaceutical-engineering/september-october-2025/new-gampr-guide-addresses-challenges-posed-ai
- Ch. 19 — European Commission, DG SANTE. Stakeholders’ Consultation on EudraLex Volume 4 – Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22 (7 Jul to 7 Oct 2025). health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en
- Ch. 19 — European Commission. EudraLex – Volume 4, Good Manufacturing Practice guidelines (web page, accessed 8 Oct 2026). health.ec.europa.eu/medicinal-products/eudralex/eudralex-volume-4_en
- Ch. 20 — Michigan Medicine Health Lab. Study of 24 U.S. hospitals shows onset of COVID-19 led to spike in sepsis alerts (6 Dec 2021), reporting Wong A., et al. Quantification of sepsis model alerts in 24 US hospitals before and during the COVID-19 pandemic. JAMA Netw Open 4, e2135286 (2021). doi:10.1001/jamanetworkopen.2021.35286. michiganmedicine.org/health-lab/study-24-us-hospitals-shows-onset-covid-19-led-spike-sepsis-alerts (secondary; the paper itself was blocked and not opened)
- Ch. 20 — Davis S.E., Lasko T.A., Chen G., Siew E.D., Matheny M.E. Calibration drift in regression and machine learning models for acute kidney injury. J Am Med Inform Assoc 24, 1052–1061 (2017). doi:10.1093/jamia/ocx030 (abstract read)
- Ch. 20 — US Food and Drug Administration, CDRH Digital Health Center of Excellence. Request For Public Comment: Measuring and Evaluating Artificial Intelligence-enabled Medical Device Performance in the Real-World. Docket FDA-2025-N-4203. 30 Sep 2025. fda.gov/medical-devices/digital-health-center-excellence/request-public-comment-measuring-and-evaluating-artificial-intelligence-enabled-medical-device
- Ch. 21 — US Food and Drug Administration. Artificial Intelligence in Software as a Medical Device (web page with milestones, content current as of 25 Mar 2025). fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-software-medical-device
- Ch. 21 — US Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Predetermined Change Control Plans for Machine Learning-Enabled Medical Devices: Guiding Principles (Oct 2023; web page republished 18 Aug 2025). fda.gov/medical-devices/software-medical-device-samd/predetermined-change-control-plans-machine-learning-enabled-medical-devices-guiding-principles
- Ch. 21 — US Food and Drug Administration. Predetermined Change Control Plans for Medical Devices. Draft guidance, docket FDA-2024-D-2338. Aug 2024. fda.gov/regulatory-information/search-fda-guidance-documents/predetermined-change-control-plans-medical-devices
- Ch. 21 — Dayma K., Patel P., Hildreth K., Jamaspishvili T. Predetermined change control plan adoption and documentation transparency in U.S. Food and Drug Administration–cleared radiology artificial intelligence/machine learning devices. Radiol Artif Intell 8(5) (online 29 Jul 2026). doi:10.1148/ryai.260385 (abstract read)
- Ch. 21 — US Food and Drug Administration. 510(k) K253281, UpDoc (Updoc, Inc.): database record and 510(k) summary; received 29 Sep 2025, decision 23 Dec 2025. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K253281; accessdata.fda.gov/cdrh_docs/pdf25/K253281.pdf
- Ch. 22 — Abbott. Important Cybersecurity Advisory: pacemaker firmware update, letter to physicians (US), 28 Aug 2017. cardiovascular.abbott/content/dam/cv/cardiovascular/pdf/reports/Pacemaker-Firmware-Update-Doctor-Letter-Aug2017-US.pdf (manufacturer)
- Ch. 22 — US Food and Drug Administration. Medical Device Recalls: recall event 78093 (St. Jude Medical pacemakers, Merlin PCS programmer software and Merlin@home software; records Z-0029 to Z-0038-2018, Class 2), initiated 28 Aug 2017, posted 12 Jun 2018. accessdata.fda.gov/scripts/cdrh/cfdocs/cfres/res.cfm?start_search=1&event_id=78093
- Ch. 22 — American College of Cardiology. FDA approves firmware addressing cybersecurity vulnerabilities in Abbott implantable pacemakers (31 Aug 2017). acc.org/latest-in-cardiology/articles/2017/08/31/12/13/fda-approves-firmware-addressing-cybersecurity-vulnerabilities-in-abbott-implantable-pacemakers (secondary)
- Ch. 22 — CSO Online. 465,000 Abbott pacemakers vulnerable to hacking, need a firmware fix (4 Sep 2017). csoonline.com/article/3222068/465000-abbott-pacemakers-vulnerable-to-hacking-need-a-firmware-fix.html (secondary)
- Ch. 22 — US Food and Drug Administration. FDA warns patients and health care providers about potential cybersecurity concerns with certain Medtronic insulin pumps. Press release, 27 Jun 2019. biospace.com/fda-warns-patients-and-health-care-providers-about-potential-cybersecurity-concerns-with-certain-medtronic-insulin-pumps (read via the BioSpace reprint; the FDA safety communication returned 404)
- Ch. 22 — US Food and Drug Administration. Medical Device Recalls: Class 2 recall Z-1581-2020, MiniMed insulin pump MMT-508 (Medtronic), initiated 27 Jun 2019, posted 26 Mar 2020, status open; and recall event 83433 (16 records, MiniMed 508 and Paradigm models). accessdata.fda.gov/scripts/cdrh/cfdocs/cfRES/res.cfm?id=175194; accessdata.fda.gov/scripts/cdrh/cfdocs/cfRes/res.cfm?start_search=1&event_id=83433
- Ch. 22 — Cybersecurity and Infrastructure Security Agency (NCCIC). ICS Medical Advisory ICSMA-19-178-01: Medtronic MiniMed 508 and Paradigm Series Insulin Pumps (27 Jun 2019, date inferred from the advisory number). cisa.gov/news-events/ics-medical-advisories/icsma-19-178-01
- Ch. 22 — Cybersecurity and Infrastructure Security Agency (NCCIC). ICS Medical Advisory ICSMA-18-219-02: Medtronic MiniMed 508 and Paradigm Series insulin pumps, remote controllers (7 Aug 2018; updated). cisa.gov/news-events/ics-medical-advisories/icsma-18-219-02
- Ch. 22 — US Food and Drug Administration. Cybersecurity (Digital Health Center of Excellence web page, content current as of 6 Jul 2026), listing the Playbook for Threat Modeling Medical Devices (30 Nov 2021). fda.gov/medical-devices/digital-health-center-excellence/cybersecurity
- Ch. 22 — US Food and Drug Administration. Cybersecurity Vulnerabilities with Certain Patient Monitors from Contec and Epsimed: FDA Safety Communication. 30 Jan 2025, updated 2 Jul 2025. fda.gov/medical-devices/safety-communications/cybersecurity-vulnerabilities-certain-patient-monitors-contec-and-epsimed-fda-safety-communication
- Ch. 22 — MITRE and Medical Device Innovation Consortium. Playbook for Threat Modeling Medical Devices (Nov 2021; FDA-funded). Release: mitre.org/news-insights/news-release/mitre-and-medical-device-innovation-consortium-create-playbook-threat; mdic.org/resource/playbook-for-threat-modeling-medical-devices/
- Ch. 23 — Regenscheid A. Platform Firmware Resiliency Guidelines. NIST SP 800-193. May 2018. csrc.nist.gov/pubs/sp/800/193/final (abstract page read)
- Ch. 23 — IEC. IEC 81001-5-1:2021, Health software and health IT systems safety, effectiveness and security – Part 5-1: Security – Activities in the product life cycle, ed. 1.0 (Dec 2021). standards.iteh.ai/catalog/standards/iso/227a3148-4206-42e3-ad21-05f80dda507a/iec-81001-5-1-2021 (standard’s front matter read via a reseller preview)
- Ch. 23 — ISO. IEC 81001-5-1:2021, catalogue page (published 21 Dec 2021; stage 90.92, to be revised, since 19 Sep 2025). iso.org/standard/76097.html
- Ch. 23 — Blue Goat Cyber. Did AAMI SW96 replace TIR57? (blog post, 21 Jul 2026). bluegoatcyber.com/blog/did-aami-sw96-replace-tir57-fda-2026 (secondary)
- Ch. 23 — AAMI. ANSI/AAMI SW96:2023, Standard for medical device security – Security risk management for device manufacturers (not opened; year and alignment with ISO 14971 as reported by AHA News and MedTech Intelligence)
- Ch. 23 — AHA News. FDA recognizes ANSI/AAMI medical device standard to enhance cybersecurity (8 Nov 2023). aha.org/news/headline/2023-11-08-fda-recognizes-ansiaami-medical-device-standard-enhance-cybersecurity (secondary)
- Ch. 23 — MedTech Intelligence. FDA recognizes AAMI SW96 cybersecurity guidance document (15 Nov 2023). medtechintelligence.com/news_article/fda-recognizes-aami-sw96-cybersecurity-guidance-document/ (secondary)
- Ch. 24 — Armis. URGENT/11 (research page; updated 1 Oct 2019 and 15 Dec 2020). armis.com/research/urgent-11/ (manufacturer)
- Ch. 24 — US Food and Drug Administration. FDA informs patients, providers and manufacturers about potential cybersecurity vulnerabilities for connected medical devices and health care networks that use certain communication software. Press release, 1 Oct 2019. fda.gov/news-events/press-announcements/fda-informs-patients-providers-and-manufacturers-about-potential-cybersecurity-vulnerabilities
- Ch. 24 — Whooley S. B. Braun, Baxter, Carestream, Green Hills affected by Ripple20 cyber vulnerabilities. MassDevice (29 Jun 2020). massdevice.com/b-braun-baxter-carestream-green-hills-affected-by-ripple20-cyber-vulnerabilities/ (secondary)
- Ch. 24 — FIRST. FIRST has officially published the latest version of CVSS (v4.0). Press release, 1 Nov 2023. first.org/newsroom/releases/20231101
- Ch. 24 — FIRST. Common Vulnerability Scoring System v4.0: Specification Document (document version 1.2). first.org/cvss/v4-0/specification-document
- Ch. 24 — Chase M., Christey Coley S. Rubric for Applying CVSS to Medical Devices. MITRE (web page dated 21 Oct 2020). mitre.org/md-cvss-rubric
- Ch. 24 — US Food and Drug Administration. New Medical Device Development Tool (MDDT) Qualification for Cybersecurity. Bulletin, 20 Oct 2020. content.govdelivery.com/accounts/USFDA/bulletins/2a6cf79
- Ch. 24 — US Food and Drug Administration. Postmarket Management of Cybersecurity in Medical Devices. Final guidance. 28 Dec 2016. fda.gov/files/medical%20devices/published/Postmarket-Management-of-Cybersecurity-in-Medical-Devices---Guidance-for-Industry-and-Food-and-Drug-Administration-Staff.pdf
- Ch. 24 — ISO/IEC. ISO/IEC 29147:2018, Information technology — Security techniques — Vulnerability disclosure, ed. 2 (23 Oct 2018; to be revised since 26 Sep 2025), catalogue page. iso.org/standard/72311.html
- Ch. 24 — ISO/IEC. ISO/IEC 30111:2019, Information technology — Security techniques — Vulnerability handling processes, ed. 2 (1 Oct 2019; to be revised since 26 Sep 2025), catalogue page. iso.org/standard/69725.html
- Ch. 24 — European Commission. Guidance – MDCG endorsed documents and other guidance (web page, accessed 8 Oct 2026). health.ec.europa.eu/medical-devices-sector/new-regulations/guidance-mdcg-endorsed-documents-and-other-guidance_en
- Ch. 25 — US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (list; page updated 4 Sep 2026). fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
- Ch. 25 — IntuitionLabs. FDA-approved AI medical devices list: complete 2026 guide (19 Jul 2026), citing counts by TheImagingWire and Innolitics. intuitionlabs.ai/articles/fda-approved-ai-medical-devices-list (secondary)
- Ch. 25 — US Food and Drug Administration. Digital Health Advisory Committee meeting of 6 Nov 2025, Generative Artificial Intelligence-Enabled Digital Mental Health Medical Devices: agenda, discussion questions and brief summary. fda.gov/media/189389/download; fda.gov/media/189392/download; fda.gov/media/190450/download
- Ch. 25 — GE HealthCare. GE HealthCare drives growth with investment in AI-enabled medical devices and tops FDA’s list of AI authorizations for 4th year with 100. Press release, 23 Jul 2025. gehealthcare.com/en/about/newsroom/press-releases/ge-healthcare-drives-growth-with-investment-in-ai-enabled-medical-devices-and-tops-fdas-list-of-ai-authorizations-for-4th-year-with-100 (manufacturer)
- Ch. 25 — US Food and Drug Administration. 510(k) K200921, qER (Qure.ai, Mumbai): database record; received 6 Apr 2020, decision 17 Jun 2020. accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm?ID=K200921
- Ch. 25 — DAIC. Qure.ai’s chest X-ray reporting tool receives additional FDA clearances (26 Feb 2026). dicardiology.com/content/qureais-chest-x-ray-reporting-tool-receives-additional-fda-clearances (secondary, reporting the manufacturer)
- Ch. 25 — Qure.ai. Regulatory and privacy (web page, accessed 8 Oct 2026). qure.ai/regulatory-and-privacy (manufacturer)
- Ch. 25 — HLTH. Aidoc wins FDA nod for comprehensive foundation model AI (26 Jan 2026). hlth.com/insights/news/aidoc-wins-fda-nod-for-comprehensive-foundation-model-ai-2026-01-26 (secondary, reporting the manufacturer)
- Ch. 25 — Tijori Alerts. Lords Mark receives India’s first Biomescan SaMD manufacturing licence (10 Aug 2026). tijorialerts.com/company-updates/lords-mark-receives-indias-first-biomescan-samd-manufacturing-licence-139868/ (secondary, reporting the manufacturer)
- Ch. 25 — US Food and Drug Administration. November 20–21, 2024: Digital Health Advisory Committee Meeting Announcement (Total Product Lifecycle Considerations for Generative Artificial Intelligence-Enabled Medical Devices). fda.gov/advisory-committees/advisory-committee-calendar/november-20-21-2024-digital-health-advisory-committee-meeting-announcement-11202024
- Ch. 25 — McGuireWoods. A pathway for clinical AI developers opens: FDA clears first software as a medical device with patient-facing LLM. Legal alert, 6 Jul 2026. mcguirewoods.com/client-resources/alerts/2026/7/a-pathway-for-clinical-ai-developers-opens-fda-clears-first-software-as-a-medical-device-with-patient-facing-llm/ (secondary, reporting the manufacturer’s claim)
- Ch. 25 — Aguilar M. A ‘historic’ FDA clearance raises the question: is the LLM an interface or the decision-maker? STAT (2 Jul 2026). statnews.com/2026/07/02/fda-clearance-raises-questions-updoc-use-generative-ai-diabetes-treatment/ (secondary; read in part, paywalled)
- Ch. 26 — Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI). OJ L (24 Jul 2026); in force 27 Jul 2026. eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32026R1744 (read in two parts, with gaps)
- Ch. 26 — Taylor Wessing. Update: AI-enabled medical devices and IVDs confirmed as high risk (Life Sciences Legal Lens, vol. 2; 2026). taylorwessing.com/en/international-life-sciences-newsletter/life-sciences-legal-lens-vol-2/update-ai-enabled-medical-devices-and-ivds-confirmed-as-high-risk (secondary)
- Ch. 26 — Covington & Burling. 5 key takeaways from FDA’s revised clinical decision support (CDS) software guidance (8 Jan 2026). cov.com/en/news-and-insights/insights/2026/01/5-key-takeaways-from-fdas-revised-clinical-decision-support-cds-software-guidance (secondary)
- Ch. 26 — OpenRegulatory. MDCG 2019-11 explained (web page dated 12 Aug 2026). openregulatory.com/mdcg/mdcg-2019-11 (secondary)
- Ch. 26 — GMP Insiders. MDCG 2019-11 Rev.1: qualification and classification of software under the MDR and IVDR (web page). gmpinsiders.com/mdcg-2019-11-rev1-mdr-ivdr/ (secondary)
- Ch. 26 — Sobande S. European revision of primary software guidance MDCG 2019-11, revision 1. Emergo by UL (20 Jun 2025). emergobyul.com/news/european-revision-primary-software-guidance-mdcg-2019-11-revision-1-small-changes-meaningful (secondary)
- Ch. 26 — Rao A. CDSCO draft guidance on medical device software. India Briefing (12 Nov 2025). india-briefing.com/news/cdsco-draft-guidance-medical-software-40691.html (secondary)
- Ch. 26 — Nishith Desai Associates. From code to compliance: the CDSCO guidance on medical device software (3 Aug 2026) (secondary)
- Ch. 26 — US Food and Drug Administration. How to Determine if Your Product is a Medical Device (content current 29 Sep 2022). fda.gov/medical-devices/classify-your-medical-device/how-determine-if-your-product-medical-device
- Ch. 26 — US Food and Drug Administration. Classify Your Medical Device (content current 26 Aug 2026). fda.gov/medical-devices/overview-device-regulation/classify-your-medical-device
- Ch. 27 — US Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft guidance, docket FDA-2024-D-4488. Cover date 7 Jan 2025 (web page published 6 Jan 2025). fda.gov/media/184856/download; fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketing (PDF read to partway through Performance Validation)
- Ch. 27 — Medical Device Coordination Group. MDCG 2020-1, Guidance on Clinical Evaluation (MDR) / Performance Evaluation (IVDR) of Medical Device Software. Mar 2020. ec.europa.eu/docsroom/documents/40323
- Ch. 27 — Commission Implementing Decision (EU) 2021/1182 on harmonised standards for medical devices, consolidated text to 17 Jun 2026. eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02021D1182-20260617
- Ch. 27 — Critical Software. IEC 62304 Edition 2 changes (web article, undated). asd.criticalsoftware.com/en/newsroom/iec-62304-edition-2-changes-august-2026 (secondary)
- Ch. 28 — US Food and Drug Administration. Deciding When to Submit a 510(k) for a Software Change to an Existing Device. Final guidance, docket FDA-2016-D-2021. 25 Oct 2017. fda.gov/media/99785/download
- Ch. 28 — Medical Device Coordination Group. MDCG 2020-3, Guidance on significant changes regarding the transitional provision under Article 120 of the MDR. Mar 2020 (rev.1 of Sep 2023 not opened). ec.europa.eu/docsroom/documents/40301
- Ch. 28 — Central Drugs Standard Control Organisation, India. IVD Medical Devices FAQ, Doc No. CDSCO/IVD/FAQ/04/2022, with addendum of 28 Mar 2025. cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/FAaddendum.pdf (read by the G-06 researcher)
- Ch. 28 — Artixio. CDSCO medical device post approval change application (19 Aug 2026). artixio.com/post/cdsco-medical-device-post-approval-change-application (secondary)
- Ch. 28 — Central Drugs Standard Control Organisation, India. Addendum No. 02 to Doc No. CDSCO/IVD/FAQ/04/2022 (13 Mar 2026). cdsco.gov.in/opencms/export/sites/CDSCO_WEB/Pdf-documents/Addendum-Doc-No-CDSCOIVDFAQ-042022dated-13032026.pdf (read by the G-06 researcher)
- Ch. 28 — Regulation (EU) 2017/745 on medical devices, Annex X section 5 (changes to the approved type). OJ L 117 (5 May 2017). Text read as reproduced at advisera.com/13485academy/mdr/conformity-assessment-based-on-type-examination
- Ch. 29 — DLA Piper. FDA issues revised cybersecurity premarket submission guidance (27 Feb 2026). dlapiper.com/en/insights/publications/2026/02/fda-issues-revised-cybersecurity-premarket-submission-guidance (secondary)
- Ch. 29 — NSF. FDA updates cybersecurity guidance: shift toward QMSR alignment rather than new requirements (3 Feb 2026). nsf.org/life-science-regulatory-news/fda-updates-cybersecurity-guidance-shift-toward-qmsr-alignment-rather-than-new-requirements (secondary)
- Ch. 29 — AAMI. Food and Drug Administration adds AAMI cybersecurity guidance to Recognized Consensus Standards Database (AAMI CR515:2025, Cybersecurity considerations unique to machine learning-enabled medical devices). Press release via Newsfile, 24 Mar 2026. newsfilecorp.com/release/289622/Food-and-Drug-Administration-Adds-AAMI-Cybersecurity-Guidance-to-Recognized-Consensus-Standards-Database (standards developer’s press release)
- Ch. 29 — Regulation (EU) 2024/2847 of the European Parliament and of the Council of 23 October 2024 (Cyber Resilience Act). OJ L (20 Nov 2024); in force 10 Dec 2024. eur-lex.europa.eu/eli/reg/2024/2847/oj (recital 25 read on EUR-Lex; Articles 2 and 71 read via springlex.eu)
- Ch. 29 — Pure Global. India CDSCO medical device software guidance 2026 (9 Aug 2026). pureglobal.com/news/india-cdsco-medical-device-software-guidance-2026 (secondary)
- Ch. 29 — BioSpectrum India. CDSCO releases guidance document on medical device software (27 Jul 2026). biospectrumindia.com/news/93/28210/cdsco-releases-guidance-document-on-medical-device-software.html (secondary)
- Ch. 29 — LexCounsel. CDSCO publishes guidance on medical device software under MDR 2017. Mondaq (5 Aug 2026). mondaq.com/cdsco-publishes-guidance-on-medical-device-software-under-mdr-2017/1826286 (secondary)
- Ch. 29 — Press Information Bureau, Government of India. Digital Personal Data Protection Rules, 2025: explainer (17 Nov 2025). static.pib.gov.in/WriteReadData/specificdocs/documents/2025/nov/doc20251117695301.pdf
- Ch. 29 — S.S. Rana & Co. MeitY notifies final Digital Personal Data Protection Rules, 2025 (14 Nov 2025). ssrana.in/articles/meity-notifies-final-digital-personal-data-protection-rules-2025/ (secondary)
- Ch. 29 — Hogan Lovells. India’s Digital Personal Data Protection Act 2023 brought into force (Nov 2025). ca.hoganlovells.com/en/publications/indias-digital-personal-data-protection-act-2023-brought-into-force- (secondary)
Take the PDF with you.
Reading online needs no sign-up. For the PDF edition, laid out for print and offline reading, tell us where to send it and we will email you.
Check your inbox.
The download link is on its way to your email. If it has not arrived in a few minutes, check your spam folder or write to info@unplex.tech.
Preview only — nothing was sent
Let's Talk