Where Does Ambient AI Scribe Audio Go After a Visit? Retention, Deletion, and Vendor Contract Terms Explained
A practical guide for IT admins on what happens to patient audio after ambient AI scribe transcription, including retention windows, model training risk, and the exact contract language to demand.
TL;DR
- Ambient AI scribe audio does not disappear the moment a note is generated. Most vendors retain raw audio somewhere in the pipeline, at least briefly, for quality assurance, error correction, or model improvement.
- Retention periods range from near-immediate deletion (minutes to hours) to 30, 60, or 90+ days, and sometimes indefinitely unless a customer negotiates a specific carve-out.
- Whether audio trains third-party models depends on contract tier, not brand reputation. General-purpose AI infrastructure often permits training use by default unless a covered entity signs an enterprise agreement that explicitly excludes it.
- A defensible AI scribe contract needs a stated retention window in days, a no-training clause covering every subprocessor, a current subprocessor list, and a deletion certification process the customer can audit.
- Voice is treated as biometric data under some state laws (Illinois' BIPA, for example), which adds a compliance layer beyond standard HIPAA that most ambient scribe contracts do not address.
- Thyra's Longitudinal AI Scribe minimizes raw audio persistence and puts retention terms in front of IT admins during procurement, not after an incident.
Ambient scribes are spreading fast across endocrinology and primary care because they cut documentation time in ways dictation and templates never could. HIMSS has repeatedly flagged third-party vendor risk as a top security concern in its annual cybersecurity surveys, and ambient AI scribes are a new, fast-growing entry in that vendor list. The convenience of "just talk and the note writes itself" hides an infrastructure question most clinicians never ask: where does the audio actually go, and for how long?
This matters more in specialty practices than people assume. The ADA's Standards of Care in Diabetes (2024 edition) explicitly calls for psychosocial screening as part of routine diabetes management, meaning endocrinology visits regularly include discussion of diabetes distress, disordered eating, mental health history, and fertility planning. That content, captured verbatim in raw audio, is a different liability surface than a structured note field, and most BAAs were written before ambient audio capture was common enough to address it directly.
How Long Do AI Scribe Vendors Retain Patient Audio Recordings?
There is no single industry-standard retention period. In practice, ambient scribe retention tends to fall into three operational patterns.
- Real-time processing with near-immediate deletion. Audio streams to a transcription engine, converts to text, and the raw file is discarded within minutes to hours. Safest from a privacy standpoint, but it limits a clinician's ability to replay a recording if a note looks wrong, and it limits the vendor's ability to retrain on your specific patient population.
- Short-term retention for quality assurance. Audio is held for a defined window, commonly 7 to 30 days, so human reviewers or automated QA pipelines can check transcription accuracy or support a clinician disputing what was actually said.
- Extended retention tied to model improvement. Some vendors keep audio for 90 days or longer, sometimes without a hard expiration, because it feeds ongoing training or fine-tuning pipelines. This is the pattern IT admins should scrutinize hardest, because "model improvement" language can be broad enough to include using your patients' voices to improve a product sold to other health systems.
Microsoft's published enterprise data commitments for Copilot products state that customer data submitted through commercial channels is not used to train the underlying foundation models without an explicit opt-in. That commitment matters here because Dragon Copilot and several competing ambient scribe tools route language model calls through enterprise cloud AI infrastructure rather than consumer-facing APIs, and the enterprise commitment only applies if your specific contract routes through that commercial channel rather than a lower, developer-tier agreement. This is exactly the kind of clause number a sales deck will not show you. The honest answer for any specific vendor is: read the current data processing addendum, not the marketing page.
Decision rule: if a vendor's contract does not state a specific day count for raw audio retention, assume the worst case (indefinite retention) and price that risk into the procurement decision. Do not accept "as long as reasonably necessary" as a functional substitute for a number.
Does Ambient AI Scribe Audio Get Used to Train Third-Party Models?
Sometimes, and the deciding factor is almost always the contract tier and subprocessor chain, not the vendor's public reputation.
Many ambient scribe products are not single-vendor systems. A clinical documentation company may license a speech-to-text engine from one provider and a large language model from another, layering clinical logic on top. Several vendors in this category, including Abridge, Nabla, and Suki, publish trust center pages stating that patient data is not used to train their own models without customer consent. That statement is meaningful, but it typically covers the primary vendor's own model, not necessarily every subprocessor in the chain.
A concrete failure mode: a primary BAA excludes training on patient audio, but the speech-to-text subprocessor's own standard terms permit using audio for "service improvement" unless a separate no-training addendum is signed with that subprocessor directly. The primary contract's promise doesn't reach that layer. This is not hypothetical; it is structural to how ambient documentation tools are built when speech recognition, medical NLP, and generative summarization are stitched together from multiple providers.
What IT admins should actually verify:
- Whether the no-training clause covers all subprocessors in the pipeline, named individually, not just the primary vendor.
- Whether "de-identified data" used for model improvement meets the HIPAA Safe Harbor or Expert Determination standard (per HHS Office for Civil Rights guidance), or whether it is audio with the patient name stripped from the transcript while voice biometrics and clinical detail remain intact.
- Whether opting out of training use requires an enterprise-tier contract that costs more, which is common, and whether that tier should be the default your organization negotiates rather than an add-on.
Illinois' Biometric Information Privacy Act (BIPA) treats voiceprints as biometric identifiers requiring specific consent and retention schedules, separate from HIPAA. Washington's My Health My Data Act (2023) extends consumer consent requirements to health-adjacent data collected outside traditional HIPAA covered entities. Neither law was written with ambient scribes in mind, but both create exposure for any organization operating across state lines whose contract only references HIPAA.
Get early access to Thyra
Built by a practicing endocrinologist. JJ personally reviews every application.
What Should Be in an AI Scribe Vendor Contract About Data Deletion?
A defensible contract states a retention period in days, names every subprocessor that touches raw audio, and requires a deletion certification the covered entity can audit on request.
Retention period, stated as a number. A specific day count, ideally under 30 days for raw audio, with note text retained per your organization's normal medical record schedule, which is a separate and usually much longer requirement.
Explicit training exclusion, named by subprocessor. A clause stating patient audio, transcripts, and derived data will not train, fine-tune, or improve any model, including third-party foundation models, without a separate signed opt-in. The absence of this clause should be treated as a yes, the vendor can use it, not a no.
Full subprocessor disclosure. A current list of every subprocessor in the audio-to-note pipeline, updated when infrastructure changes, with advance notice before a new subprocessor is added. HL7's FHIR Provenance and Consent resources are increasingly referenced in health IT RFPs as a way to structure this kind of audit trail; asking whether a vendor can map its pipeline to those resources is a reasonable procurement question.
Deletion certification and audit rights. Even an annual attestation letter confirming audio matching the stated retention period has been deleted, plus audit rights during a security review or breach investigation.
Breach notification specific to audio data. Standard HIPAA breach rules apply, but voice is biometric data under some state frameworks, so the contract should address audio-specific incident response, not just generic PHI language.
Data residency and encryption at rest and in transit. Where the audio physically lives, and whether encryption meets standards your security team has already vetted, per HHS Office for Civil Rights guidance on the HIPAA Security Rule.
| Retention Pattern | Typical Window | Primary Risk to Watch |
|---|---|---|
| Immediate deletion after transcription | Minutes to hours | Limited ability to re-verify a note against original audio if a clinician disputes accuracy |
| Short-term QA retention | 7 to 30 days | Vendor staff or contractors may have access to raw audio during the review window |
| Extended retention for model improvement | 90 days or longer, sometimes undefined | Highest risk of audio informing third-party model training without explicit consent |
| Customer-controlled retention (enterprise tier) | Configurable, often 0 to 30 days | Requires negotiation and a higher contract tier, but gives IT admins direct control |
Who Actually Processes the Audio Behind the Scenes?
A typical pipeline looks like this: a microphone captures audio in the exam room, a streaming API sends it to a speech-to-text engine, the transcript passes to a large language model for clinical summarization, and structured output returns to the EHR as a draft note. Dragon Copilot, Abridge, Nabla, Suki, and Ambience Healthcare represent the current competitive set, and each generally runs on cloud infrastructure from major providers, routing language model calls through enterprise AI services rather than consumer-facing APIs.
The enterprise versions of those services typically carry different default data use terms than consumer or developer-tier access to the same underlying models, which is why the specific contract tier your organization signs matters more than the brand name on the product. Two health systems using what looks like the same scribe product can have meaningfully different audio retention and training exposure depending on which agreement their legal and IT teams actually negotiated.
For IT admins, the practical takeaway is to ask for the full subprocessor map during procurement, not just the top-line vendor's privacy policy. If the vendor cannot produce that map or treats it as proprietary, that is itself useful information about how much visibility you will have during an incident.
What Questions Should IT Admins Ask Before Signing an AI Scribe Contract?
- What is the exact retention period for raw audio, in days, and is that number in the contract itself or only in a policy document the vendor can change unilaterally?
- Is patient audio used, in any form, to train or fine-tune models, including third-party foundation models accessed through the vendor's pipeline?
- Can we obtain a current subprocessor list, and will we be notified before a new one is added?
- What happens to audio and transcripts if we terminate the contract? Is there a defined export and deletion timeline post-termination?
- Does the BAA specifically reference audio recordings as a PHI category, or does it use generic language that could exclude voice data from certain protections?
- Who has access to raw audio during any QA window, and is that access logged and auditable?
- If a patient requests deletion under state privacy law, does that extend to voice already ingested into a training pipeline, and is deletion from a trained model even technically possible?
That last question is worth sitting with. If a model has already been fine-tuned on audio including a specific patient's voice, deleting the original recording does not necessarily remove that patient's influence from the trained model. This is a genuinely unresolved question across the AI industry, not something specific to healthcare. If a vendor cannot explain how deletion requests interact with model training, treat that as a gap, not a resolved risk.
How Does Thyra Handle Ambient Scribe Audio and Retention?
Thyra's approach is to minimize how much raw audio needs to exist in the first place, and to keep whatever retention does occur inside a contract IT admins can review before deployment.
Thyra's Longitudinal AI Scribe is built into the same clinical data environment as the rest of the platform, the Smart Inbox, CGM viewer, orders, and protocols, so audio processing happens within infrastructure already governed by the practice's BAA rather than being routed to a separate consumer-facing AI product bolted onto the EHR. The practical security benefit of that single clinical brain design goes beyond workflow convenience: fewer external hops for sensitive audio means fewer subprocessors to vet and fewer contract gaps to close during procurement.
We do not claim zero retention, because that claim is not honest for any ambient scribe product that needs a window for QA and error correction. What we commit to in writing is a stated retention period for raw audio, a contractual exclusion from third-party model training without explicit customer opt-in, and a subprocessor list IT admins can request during procurement. For a specialty like end