Trained on You
The scribe records the visit to draft the note. But the conversation it captures — your patient's words, your reasoning spoken aloud — has a second life the pitch never mentions: once it is de-identified, it leaves the privacy law entirely, and can train a model the vendor owns.
I. The third party in the room
Consider what the microphone hears. Not a transaction — a confession. Over the course of a visit a patient tells a stranger in a white coat the things they have told no one else: the symptom they were too frightened to name, the medication they stopped taking and lied about, the drinking, the fear, the home they are afraid to go back to. And the physician answers in kind, thinking aloud, reasoning through the differential in real time — the working mind of a clinician, narrated. It is one of the last conversations in modern life that both parties assume is private, and the assumption is not naïve. It is the whole premise of the encounter. People do not tell the truth to doctors because they are brave. They tell the truth because they believe the room is sealed.
For most of the profession’s history the room was sealed, more or less, by the simple physics of speech: words said aloud in an exam room vanished as they were spoken, surviving only as much of them as the physician later chose to write down. The ambient scribe — the AI that listens to the visit and drafts the note, the subject of the first essay in this series — changes that physics completely. Now the conversation is captured, in full, at the source. Every word, on both sides, streamed to a server the moment it is spoken. It is important not to overstate how far this has spread: the ambient scribe is not yet in every room. It remains a minority of American practice — on the order of a third of physician groups by 2025, by one market analysis, with no national count of the share of actual encounters it touches. But adoption is climbing on a one-way curve, and in the rooms where the microphone is on, the capture is total.
The first essay asked where the value of that recording goes — and found it booked by the institution as coding revenue while the physician kept the liability. This essay asks a different and quieter question about the same recording: where does it go? Not the note it produces — the raw material. The conversation itself. And the answer, which almost no physician using these tools has been told and which is hiding in plain sight in documents the vendors publish, is that the conversation does not stay in the sealed room. Under a provision written into the privacy law a quarter-century ago, there is a single step that lifts it out of that law entirely — after which it can be kept, pooled, and used to train a commercial model, and no one whose voice is in the recording has to agree.
II. The switch
The step is called de-identification, and to see why it matters you have to see that HIPAA is narrower than almost everyone believes. The privacy rule does not protect health information. It protects individually identifiable health information — “protected health information,” PHI — and it protects nothing else. Strip the identifiers, and the rule, by its own terms, stops applying.
This is not an interpretation. It is the text. Under 45 CFR § 164.502(d), health information that has been de-identified “is considered not to be individually identifiable health information,” and then the operative sentence: “The requirements of this subpart do not apply to information that has been de-identified.” The subpart is the privacy rule. Once data clears the de-identification bar, the rule releases it — no restriction on use, no restriction on disclosure, no authorization required, because there is no longer a protected thing to require authorization for.
The bar itself is not high. Section 164.514 offers two ways over it. One is “Expert Determination”: a qualified statistician certifies that the risk of re-identification is “very small.” The other, the one most used because it is mechanical, is “Safe Harbor”: remove eighteen enumerated identifiers — name, geography smaller than a state, dates more specific than a year, medical-record and device numbers, and so on — and, absent actual knowledge that the remainder could still identify someone, the data is de-identified as a matter of law. The words a patient spoke, and the clinical reasoning spoken back, are not on the list of eighteen. Scrub the identifiers and what remains — the substance of the conversation — walks out of the privacy rule carrying none of its protections.
There is a sharper way to feel the size of this. HIPAA does contain a bar on selling patient data: under § 164.502(a)(5)(ii), a covered entity or business associate “may not sell protected health information” without the patient’s authorization. It reads like a firm floor. But it is a floor under protected health information — and de-identified data, by the definition three paragraphs down in the same regulation, is not that. The prohibition on selling a patient’s medical information and the freedom to sell the same information de-identified are not in tension in the rule. They are the rule, working exactly as written. The one act HIPAA most plainly forbids becomes permitted the moment the identifiers come off.
None of this was a loophole someone discovered. De-identification has been in the regulation since it took effect in 2000, designed for a legitimate and valuable purpose — to let health data feed research and public health without exposing the people in it. What is new is not the switch. What is new is the machine that has been wired to it: a recorder in a fast-growing share of exam rooms, generating de-identifiable clinical conversation at a scale no one contemplated when the rule was drafted.

III. Said out loud, in the terms
Here the discipline of this journal is to describe what the people involved have already disclosed, rather than to allege something hidden — and the ambient-scribe industry has disclosed this, in its own published documents, for anyone who reads them.
Begin with the market leader, because it is the most candid. Microsoft’s Dragon Copilot is used across many of the country’s largest systems, and its privacy white paper, updated this past spring, sets out the pipeline without euphemism. During a visit, it says, the tool “securely streams Customer Data in the form of audio data (including voice of the patient, the patient’s care team, and other encounter participants) to the cloud,” and generates a transcript. Select customer data — “audio recordings and transcriptions, clinical documentation” — is then transmitted to “a designated Dragon Copilot research environment where it is processed and anonymized within 90 days.” The anonymization, the document states, is “specifically designed to meet” the standard “for de-identification under the United States’ Health Insurance Portability and Accountability Act.” And then the sentence itself, stated plainly:
Dragon Copilot’s AI/ML models are trained solely on anonymized data.
Read it beside the regulation and the architecture is complete. The conversation is captured; it is anonymized to the HIPAA de-identification standard, which — per § 164.502(d) — removes it from the privacy rule; and the resulting data trains the model. Each step is disclosed, each step is lawful, and the sum is that the patient’s words become model-training material by operation of a rule most physicians assume protects them.
One more line from the same document closes the door the patient might think they still have. A patient can revoke consent; the vendor will purge data that links back to them. But: “Resulting anonymized data created by Microsoft prior to the patient’s revocation of consent will not be purged.” What was learned from the conversation is kept. Consent, withdrawn, reaches only the identified copy — never the anonymized data already absorbed into the training corpus. The recording can be deleted. The training is permanent.
Microsoft is not an outlier for doing this; it is an outlier for documenting it so clearly. The pattern recurs across the industry’s terms. Suki’s terms of service reserve the right to use customer data for “system tuning” and “training of acoustic models and other models,” and to “de-identify and/or anonymize” that data, after which the terms treat it as the company’s own aggregate data. Abridge’s business associate agreement grants the vendor the right to “use PHI for purposes of de-identification of the PHI” — the contractual on-ramp to the same exit — and the company has publicly described training its newest model on de-identified data. Vendors differ in how long they retain the raw audio, and several delete it within weeks; that is worth crediting, and this essay will return to why it is less protective than it sounds. But retention of the audio is a separate question from what the model keeps, and on the substance the disclosures converge: the clinical conversation is a training input, and de-identification is the mechanism that makes it a lawful one.
IV. The consent that was given, and the one that wasn’t
Here the argument has to be careful, because the easy version of it is wrong. It is tempting to say that none of this is consented to, and that would be false. In most settings that use an ambient scribe, the patient is told and agrees; the disclosure is commonly folded into the consent-to-treat and privacy paperwork a patient already signs, and a practice that records visits will, as routine and often as policy, say so and ask. Physicians are rarely compelled either — at most institutions the scribe is opt-in, taken up by the clinicians who find it helps and left alone by those who don’t. Consent, in the ordinary sense the word carries at the front desk and in the exam room, is usually present on both sides. An argument that pretends otherwise deserves to lose.
But look at what that consent covers, and what it does not. The patient agrees to be recorded so the visit can be documented — a treatment purpose, and a reasonable one. The physician agrees to use a tool that drafts the note. Neither is asked the question that actually matters: whether the de-identified residue of the encounter — the patient’s words, the physician’s reasoning — may be used to train a commercial model the vendor will own. And neither is given a way to refuse it, for the reason already established. The recording reaches the vendor, a business associate, on nothing more than the contractual “satisfactory assurance” HIPAA requires; treatment and operations need no patient authorization to begin with, which HHS states plainly — the rule “permits, but does not require” consent for them; and once the data is de-identified there is no protected thing left to authorize. The consent that is obtained is genuine. It is also beside the point. It governs the recording. It does not govern, and offers no way to decline, the training.
That is the precise gap, and it is worth putting in the plain form a patient would use: I agreed to be recorded. I did not agree to become training data, and no one gave me a way to say no. The physician whose diagnostic voice sits in the same file can say the identical sentence. A caveat, because the record only supports so much: HIPAA nowhere says “artificial intelligence” or “model training” — the rule predates the use. What the documents establish is narrower and sufficient: de-identified data falls outside the rule, and the vendors’ own terms say they train on it. That no one’s authorization is required for that training is an inference from those two facts — but not a strained one. It is the reading the vendors themselves have acted on.
V. Do they really have a choice?
Return to the physician’s “choice,” because it is more constrained than the word admits. The clinician who opts into a scribe does so from inside a system that was never opted into at all.
The electronic record was not so much adopted by physicians as required of them, and enforced with money. The 2009 HITECH Act paid clinicians to install certified electronic records and then, beginning in 2015, reduced the Medicare payments of those who had not — a mandate carried by a penalty, not an invitation. Office-based adoption climbed from 42 percent in 2008 to 95 percent in the years that followed, which is the shape an adoption curve takes when a payer insists. With the record came the rest of the apparatus — the coding rules, the billing edits, the quality measures, the documentation load that the first essay’s two hours of desk work for every hour of care describes. Very little of it was designed by the people who would spend their evenings inside it.
The ambient scribe enters that history not as a free choice but as relief from an imposed one — a way to endure a documentation burden the imposed system created. A physician who takes it up to get the evening back is exercising a real preference, but inside constraints they had no hand in setting. That is a thinner sort of consent than the word usually implies, and the profession’s collective say was thin too. The body treated as medicine’s representative voice, the American Medical Association, counted roughly three-quarters of American doctors as members in the 1950s and now speaks for fewer than a quarter of active physicians — a minority that nonetheless, through its coding panels, helps author the billing machinery all the rest must work inside. When the systems that shape a physician’s day were built, the physician was mostly not in the room.
Which is the pattern the series keeps turning up, one layer deeper each time. The value of the note goes to whoever books the coding; the substance of the conversation goes to whoever holds the corpus; and at every step the people who generate the thing of value are handed a tool, told it will help, and never asked about the terms on which what they make is kept. Everyone wants to collect on the effort. Almost no one thinks to ask the person making it.
This has, throughout, been an argument about American law — and the switch it depends on, de-identify and the privacy rule lets go, is a specifically HIPAA feature. Other wealthy democracies draw the line elsewhere: the European Union’s GDPR also exempts truly anonymized data, but sets a demanding bar for reaching it and surrounds secondary use with separate consent norms; the United Kingdom, Canada, Japan, and other OECD systems each balance secondary use against patient control in their own way. Whether the training corpus that HIPAA makes frictionless to build is as easily built abroad is a real and separate question — one a later essay will take up on its own.
VI. What is kept is not the recording
The industry’s most reassuring fact — many vendors delete the raw audio within days or weeks — turns out to be the one most likely to mislead, and it is worth being precise about why.
Deletion of the audio answers the wrong question. The audio was never the durable asset. The durable asset is what the model learned from it — a pattern of clinical language, reasoning, and phrasing distilled from millions of encounters and fixed in the weights of a system the vendor owns. That distillation outlives the recording by design. The company that told you it deleted the audio in thirty days can be, at the same moment and with no contradiction, a company that will hold what it learned from your visit for as long as the model exists — the permanence the market leader disclosed when it said anonymized data survives a patient’s revoked consent. The recording is transient. The training is not. Asking how long they keep the audio is a little like asking a lender how long they keep the pen you signed with.
It is the shape of the first essay again. There, the value the scribe unlocked — richer coding of the same visit — was booked by the institution while the physician carried the risk. Here, the raw material of a different kind of value — the clinical conversation, generated jointly by a patient who bared something and a physician who reasoned aloud — is captured and refined into an asset kept by a third party that was in the room only as a microphone. In both cases the people who produced the thing of value do not own the thing of value. The compensation series traced that pattern through the pricing of physician labor; the intelligence, it turns out, is priced the same way. The work is the physician’s and the patient’s. The asset built from it is somebody else’s.
The last thread to handle honestly is re-identification, because de-identification is not the perfect seal the word implies. Safe Harbor removes eighteen identifiers; Expert Determination certifies only that the residual risk is “very small,” which is a deliberately chosen phrase and is not zero. A rich enough record — a rare diagnosis, an unusual phrasing, a constellation of dates and details — can in principle be linked back, and the literature on re-identification of “anonymized” datasets is a running caution against treating the label as a guarantee. This essay does not rest on that risk, and I flag it as a live but secondary concern rather than a proven harm here. The primary point needs no re-identification at all. Even if the de-identification holds perfectly and no patient is ever unmasked, the transfer has already happened: the substance of the conversation has become a corporate training asset, lawfully, without consent, and permanently.
VII. The honest case, and the law that has not caught up
The strongest objection to everything above deserves its full weight, because it is real and I think it is partly right.
De-identification is not a trick. It is one of the most useful provisions in health-information law, and the reason is that pooled clinical data, stripped of identity, is how medicine learns. Registries, quality measurement, public-health surveillance, and — yes — better clinical AI all depend on it. A scribe trained on a large, de-identified corpus of real clinical language is genuinely better at drafting a usable note than one trained on less, and the physician down the line, and that physician’s patient, benefit from the improvement. Training on de-identified data is not self-evidently wrong; in much of medicine it is how progress is made. Grant all of that — I do — and the argument survives it, because the argument was never that the training is illegitimate. It is that the people who generate the raw material have been given no say in the matter and no stake in the result, and that a privacy law most of them trust is the very instrument that arranges it. The question is not whether the corpus should exist. It is who owns it, who profits from it, and whether the patient and the physician who made it were ever asked.
That question currently has no regulator tending it. The Federal Trade Commission has warned, in a 2024 statement, that “there is no AI exemption from the laws on the books” and that a company may not quietly “use customer data for secret purposes, such as to train or update their models” — but that is a general admonition, not a rule aimed at this practice, and it governs deception, not disclosed and lawful use. On the health side, the gap is starker: as of this writing there is no HHS guidance addressing ambient AI scribes or the use of de-identified clinical conversation to train commercial models at all. The controlling text is a de-identification standard written more than two decades ago, for a world in which “health data” meant a billing record in a filing cabinet, not the recorded voice of every patient in every room. The rule is not being broken. It is being used, precisely as written, for something it was never written to govern.
VIII. Whose intelligence
The scribe was the visible transfer — the value of the note, moving to whoever books the coding. This is the invisible one: the conversation itself, moving to whoever holds the corpus. It is quieter because nothing appears to leave. The patient goes home, the audio is deleted on schedule, the privacy notice was signed, every box is checked. And a durable asset built out of the most private conversation in medicine now sits on a company’s servers, lawfully, with no one’s permission and no one’s stake but the company’s.
There is a test any physician can run this week, and it costs nothing. Open the terms of service or the business associate agreement for the scribe your institution uses — the document is almost always public — and search it for two words: “de-identify” and “train.” Then ask your vendor, or your informatics office, the two questions the documents are written to answer: Do you retain our audio or transcripts, and for how long? And do you use de-identified data from our encounters to train or improve your models? You are entitled to the answers, and the answers will tell you which side of the switch your patients’ conversations are on.
The modest, unwelcome question to carry out of the room is no longer only the first essay’s — the value that time unlocked, who is keeping it? It is also this one: the conversation you and your patient believed was sealed — who, now, has been trained on it? The next essay follows the asset one step further, to the thing the corpus is used to build, and asks what it would take for the people who generated the intelligence to own a share of it.
Satyanarayan Hegde, MD, is a pediatric pulmonologist and the founder of Access Pediatric.