Healthcare AI data breaches: Why companies need to protect more than patient records

Healthcare and life science companies are not only protecting traditional patient records – things like names, diagnoses, treatment notes, lab results, and insurance information. They may also be protecting clinical trial data, research data, vendor systems, proprietary models, and other information used to support patients and drug development. When these companies and systems are involved in data breaches, the legal risks can become broader than a standard data breach. 

A useful example is the recent Novo Nordisk incident. Novo Nordisk is the pharmaceutical company that makes drugs such as Ozempic and Wegovy and recently disclosed a cyber incident involving patient data from clinical trials. It has also been reported that a hacking group claimed to have stolen over a terabyte of data and attempted to extort the company, though Novo Nordisk has not confirmed the entire scope of those claims. 

The Novo Nordisk incident shows that healthcare breaches can involve more than names and medical records. Companies also need to consider whether a cyber incident involves clinical trial information, healthcare provider data, research files, AI model information, or other confidential business materials. When this type of data is involved, companies may need to take precautions to protect the data, document their response, and comply with applicable privacy and cybersecurity obligations.

Recent health-related breaches also highlight how these incidents create major legal and business consequences. The Change Healthcare cyberattack that occurred in 2024 shows the scale of healthcare cybersecurity risks with reports that 197.2 million individuals were impacted by this breach. For scale, 197.2 million people would be more than half of Americans. The 23andMe breach also highlights the litigation risks tied to sensitive genetic information, with a bankruptcy judge approving a $46.75 million settlement for victims in the breach. While the facts of these examples differ, they all highlight why companies handling health-related data need to treat cybersecurity, AI governance and breach response as connected compliance and a priority in business practices. 

Why are AI-related healthcare breaches unique?

A standard healthcare breach analysis often focuses on whether names, medical records, insurance information, or other identifiable patient information was exposed. These concerns are highly important and sensitive. But as artificial intelligence becomes more integrated into healthcare, it creates additional risks because AI tools rely on large volumes of sensitive data that may also be tied to valuable research and business information.

For example, an AI model used in drug development may involve clinical trial data, molecular research, imaging data, prompts, outputs and other model training information. If these systems are accessed by an unauthorized actor, the incident may raise privacy concerns and may also create IP and competitive risks. 

This means healthcare AI security should not only focus on protecting patient records. Companies should also consider whether their AI models, training datasets, research systems, and vendor environments are adequately secured.

Why does de-identification matter?

De-identification can be key when healthcare data is used with AI. In simple terms, de-identification means removing information that could identify a person up to a certain legal standard. This matters because companies may want to use patient data or clinical data to train, test or improve their models. 

If health data is properly de-identified before being used with AI, the company may reduce privacy risks. However, if the data is not properly de-identified, several questions remain. The company may need to ask whether its privacy notices clearly explained that health data could be used for AI training or analysis. State privacy laws may also require clear disclosures about how personal information is collected and used. For example, California’s CCPA requires covered businesses to provide notice about the categories of personal information collected and the purposes for which that information is collected or used. This is especially important when health-related data is later used in a way that patients or consumers may not have expected. 

The company may also need to consider whether protected health information was shared with an outside AI vendor, whether a business associate agreement was required, and whether patient authorization was needed.

What happens if PHI may have been compromised? 

Breach response is also an important issue. If a healthcare company experiences a ransomware attack or an unauthorized access incident, they may need to assess whether protected health information (PHI) was compromised, This may include reviewing what information was involved, whether it could identify an individual, who accessed it, whether it was viewed or acquired and if the risk was mitigated or not.

The Novo Nordisk incident highlights why this analysis matters. Novo Nordisk stated that the clinical trial information involved was not directly linked to patient names or other direct identifiers. However, the incident still raised questions about patient-related data, de-identification, and whether any protected health information may have been compromised. This is why companies should be able to document how health-related data was stored, protected, and separated from information that could identify individuals.

What should companies take away? 

For healthcare and life science-related companies, the key takeaway is that AI governance should include privacy and cybersecurity from the start. Before using health-related data with AI, companies should ask:

  • What data goes into the AI system?
  • Is the data identifiable or properly de-identified?
  • Who can access the data, prompts, outputs, or model?
  • Are outside vendors involved?
  • Can the vendor use the data to train or improve its own models?
  • What security controls apply if the system is breached?

Companies should also treat AI models, training datasets, prompts, outputs, and research tools as sensitive assets as well since they can carry PHI and may lead to inadvertent disclosures. Vendor contracts should address confidentiality, data retention, security controls and model training. 

AI may help healthcare and life science companies innovate and operate faster, but innovation cannot replace privacy and regulatory accountability. Companies using AI with health-related data need to confirm de-identification practices, strengthen vendor contracts, and properly prepare for responses before incidents occur.