The notion that personal health information shared with a physician remains strictly confidential, confined to the patient-doctor relationship and perhaps the insurer, is becoming increasingly antiquated. While the Health Insurance Portability and Accountability Act (HIPAA) is widely perceived as a robust shield for medical data, its actual scope is far more limited than its reputation suggests. The federal law governs the handling of health information by specific entities—hospitals, physicians, insurers, and their business associates—but it conspicuously excludes a vast and growing ecosystem of health-related data generated outside these traditional channels. This includes sensitive information captured by period-tracking applications on smartphones, data from internet searches about medical conditions, genetic information submitted to genealogy companies, and the myriad metrics collected by wearable health devices that monitor everything from heart rate to sleep patterns.
Even the health records that fall under HIPAA’s purview are subject to disclosure, sale, and transfer to government entities in ways that often surprise individuals, revealing significant gaps in the supposed privacy protections. This erosion of data safeguards is particularly concerning given the current trajectory of U.S. government policy. There is a concerted push to amass domestic health data and, increasingly, to acquire it from abroad. This initiative proceeds even as a growing body of scientific research demonstrates that the primary safeguard relied upon for these data-gathering efforts—the anonymization of data by removing identifying information—is far less effective than official pronouncements suggest.
As a professor of law specializing in health information privacy and medical data regulation, I have extensively studied the intricate pathways through which sensitive health information moves among healthcare providers, government agencies, and law enforcement. My research, including work on a federally funded study concerning opioid prescribing, underscores the immense value of health data for scientific advancement. However, it also highlights the inherent dangers of collecting such information without robust, meaningful safeguards.
The Evolving Landscape of Medical Privacy
HIPAA grants individuals several fundamental rights concerning their health records. Patients can access their medical information, request corrections to inaccuracies, and expect that covered healthcare providers will not disclose their data without authorization. However, the law contains numerous provisions that permit the release of certain health information without explicit patient consent. Covered entities, such as hospitals, are authorized to release specific types of records under approximately a dozen categories without requiring patient authorization or notification.
These permissible disclosures extend to information necessary for treatment, payment, and routine healthcare operations. Critically, HIPAA also allows for the disclosure of health information for public health reporting, law enforcement purposes, judicial and administrative proceedings, health plan oversight, research, and for what are broadly termed "essential government functions." These exceptions, coupled with additional statutory allowances, create numerous avenues for health information to be shared beyond the immediate patient-provider context.
Once health data is transferred outside the purview of HIPAA-regulated entities, the protections afforded by the law effectively cease to apply. This is particularly relevant in the context of state-operated prescription drug monitoring programs (PDMPs). These programs compile detailed logs of controlled substance prescriptions, including who filled them and when. Federal law enforcement agencies can often access these logs through administrative subpoenas, which do not require judicial oversight or approval. These programs have expanded significantly, encompassing more than just opioid prescriptions, and now function as a broad dragnet that shares health data across state lines. This can inadvertently expose individuals seeking sensitive reproductive or gender-affirming healthcare to surveillance, even when they are in states with different legal frameworks. The destination of health records and the rules governing their flow can vary significantly, and the perceived breach of privacy can depend heavily on who is making the disclosure decisions and who ultimately gains access to the information.
Government Initiatives for Health Data Acquisition
A notable development in the increasing governmental pursuit of health data has been the initiative led by Health and Human Services Secretary Robert F. Kennedy Jr. Since the spring of 2025, Secretary Kennedy has been advocating for expanded federal access to Americans’ medical records. The stated objective of this push is to investigate the potential link between vaccines and autism, a hypothesis that has been extensively studied and conclusively refuted by decades of scientific research, which has demonstrated no such causal relationship.
According to reports from KFF Health News, the Department of Health and Human Services (HHS) has been actively engaging with state health information exchanges (HIEs). These are often less-understood systems that facilitate the secure exchange of detailed, identifiable patient records among hospitals and clinics. HHS has been inquiring about the potential use of these records for vaccine research. One proposal put forth by state organizations envisions providing HHS with access to approximately 90% of Americans’ medical records by 2028. In Nebraska, substantial federal grant funding has been directed to a statewide health information exchange non-profit organization that has cooperated with this data-gathering effort.
The value of large health datasets for public health initiatives is undeniable. Aggregated records can be instrumental in identifying drug side effects, tracking disease outbreaks, and revealing disparities in healthcare access and quality that might be missed by smaller-scale studies. Historically, public health has relied on a degree of individual privacy surrender for the collective good. The concern, however, is not that the government should never collect health data, but rather that the meaningful safeguards designed to protect this data have not kept pace with the expanding scale of data collection and the sophisticated capabilities of modern data analytics.
When HHS seeks access to medical records for a vaccine and autism study—a question already definitively answered by science—the department has been reticent to disclose crucial details. These include the number of states involved, the specific types of data being collected, the individuals or entities with access to this information, and the precise measures being taken for its protection. This approach appears to invert the typical logic of research, where a well-defined hypothesis guides the data collection process, rather than the collection of extensive data paving the way for a research question.
Moreover, the creation of a comprehensive repository containing identifiable records for tens of millions of individuals presents a significant target for data breaches. It also opens the door to secondary uses of the data for purposes to which individuals have not consented, and potential abuses by current or future administrations with differing policy priorities.
The Unreliability of "Anonymized" Data
Officials have offered assurances that the collected data will be aggregated and stripped of direct identifiers, thereby preventing the singling out of individuals. However, decades of computer science research cast significant doubt on the efficacy of such assurances. A study published in Nature in June 2026 further underscored this vulnerability, demonstrating that in the era of artificial intelligence, the removal of identifiers from patient records does not uniformly protect all individuals.
The researchers evaluated AI diagnostic models that had been trained on clinical data, including chest X-rays, electrocardiograms, and electronic health records. Their objective was to determine if an external party could ascertain whether a specific individual’s data had been used in the model’s development. For instance, confirming that an individual’s record contributed to the training of a cancer prediction tool could reveal that the individual has cancer. This type of exploit is known as a "membership inference attack."
The research team found that while the average risk of re-identification from data stripped of identifiers often appeared acceptably low, certain patients faced a near-certainty of being re-identified. This burden fell disproportionately on underrepresented groups, categorized by race, insurance status, or diagnosis. Those most exposed were frequently individuals already vulnerable to discrimination. Researchers have long established that the removal of identifiers from comprehensive datasets does not reliably protect the individuals within them, and that re-identification becomes progressively easier as the volume and richness of the data increase. Current AI technology enables these attacks to be conducted remotely and with unprecedented speed.
Global Implications of Data Extraction
The U.S. government’s extensive appetite for health data extends beyond its borders. Reports from June 2026 by ProPublica revealed that the State Department has been conditioning lifesaving aid to African nations on gaining access to their citizens’ health data.
Under the Trump administration’s "America First" global health strategy, Uganda agreed to grant the United States real-time access to nine of its health data systems for a period of seven years. This included access to the nation’s central health information repository and the system managing individual electronic medical records. In exchange, Uganda was slated to receive up to $1.7 billion over five years, a sum that has reportedly decreased annually and falls below previous levels of U.S. support. Kenya entered into a comparable agreement, while Zambia, Zimbabwe, and Ghana reportedly walked away from the initial terms of similar deals.
The U.S. government has pledged that the data acquired through these agreements will be aggregated and anonymized. However, privacy experts have voiced concerns that these agreements are vague and lack standard limitations on the volume of data collected and the permissible uses of that data. A Ugandan digital rights lawyer characterized the choice facing his country as the essence of "digital colonialism," where accepting the deal entails the risk of exploitation, while refusing it could mean a lack of essential aid, potentially leading to preventable deaths.
The Common Thread: Faith in Anonymization
Both the domestic collection of health records and the international data-for-aid agreements appear to be predicated on a shared assumption: that anonymization effectively neutralizes the risks associated with pooling sensitive health data. The available evidence strongly suggests otherwise. This does not imply that health data should never be collected or studied. However, it necessitates a critical examination of official reassurances, rigorous scrutiny of existing safeguards, and a recognition that individuals whose bodies generate this data deserve a meaningful voice in its governance.
For governments seeking access to sensitive medical records, a more transparent approach would involve demonstrating the necessity of such access and providing concrete evidence of the effectiveness of their proposed safeguards. The evolution of privacy law has largely been shaped by an era when data was primarily stored in physical filing cabinets. Today, governments, whether operating in Kalamazoo or Kampala, are navigating a digital landscape where even an anonymized digital record can potentially be traced back to an individual. This fundamental shift demands a recalibration of privacy protections to meet the challenges of the 21st century.







