AI Chart Audits for Risk Adjustment and Coding Compliance: What They Catch That Manual Review Misses

Author: Jean Jacques Nya Ngatchou, MD | August 3, 2026

AI chart audits scan for HCC coding gaps, unsupported diagnoses, and documentation-to-billing mismatches at a scale manual review cannot match. Here is what they actually catch.

TL;DR

What is an AI chart audit and how does it differ from clinical documentation review?

An AI chart audit is an automated, systematic check of whether a chart's billed diagnosis codes are supported by what is documented in the encounter, run across some or all of a practice's charts rather than a manual sample.

This is a different function from clinical documentation review, and the distinction matters for how compliance and IT leaders think about tooling. Clinical documentation review, including the kind built into ambient scribes and note-generation AI, asks whether the note is a faithful, complete, clinically coherent record. That is a quality and safety function. A coding compliance audit asks a narrower question: for every ICD-10 code that hit a claim, does the note contain evidence the condition was evaluated, monitored, treated, or its treatment considered during this encounter? This is the layer CMS's Risk Adjustment Data Validation (RADV) program and OIG reviewers care about, and it runs on billing logic (HCC categories, MEAT/TAMPER criteria, code specificity, hierarchical mapping) rather than narrative quality. A documentation review flags "this note is missing detail on the patient's response to metformin." A coding compliance audit flags "E11.22 was billed but there is no documented nephropathy management or lab correlation in this note, so this HCC is not supportable if audited."

How does AI detect unsupported HCC codes or risk adjustment errors?

AI audit tools parse the full chart, not just the billing sheet, and test each submitted diagnosis against structured evidence requirements. CMS's RADV methodology, and most payer-specific validation frameworks, are built around this same evidence question, which is what makes automated screening tractable at all.

Diagnosis-to-evidence matching. The system extracts every diagnosis code on the claim and searches the note, problem list, medication list, and orders for corroborating evidence: a relevant lab, a medication tied to that condition, an exam finding, or explicit assessment language. A chart listing E11.40 (diabetes with neuropathy) with no neuropathy exam, no relevant medication, and no mention of neuropathy anywhere gets flagged as unsupported.

MEAT/TAMPER logic applied at scale. CMS risk adjustment validation asks whether a condition was Monitored, Evaluated, Assessed, or Treated; some payer frameworks use TAMPER (Treatment, Assessment, Monitoring, Plan, Evaluation, Referral). AI tools encode these as rule sets and run them against every chart rather than a spot sample, producing a pass or fail for each HCC-relevant code on every encounter.

Code specificity and hierarchy checks. Many risk adjustment errors are under-specified rather than fabricated: billing E11.9 (diabetes without complications) when the note documents diabetic retinopathy that maps to a higher-value HCC. Tools trained on ICD-10 hierarchy and HCC mapping catch this, which is a lost-revenue problem as often as a compliance one.

Recapture and persistence checks. Chronic condition HCCs generally need to be re-documented and re-supported each payment year to remain valid. Tools track which chronic HCCs were billed in a prior period and flag when they go undocumented in the current one, one of the hardest gaps to catch manually because it requires comparing across encounters and time.

Documentation-to-billing mismatch detection. This covers a code on the claim absent from the note, a documented condition never coded (missed revenue, not compliance risk), or codes copied forward from a problem list without current-encounter evidence, sometimes called cloning.

What compliance risks do practices face without automated coding audits?

Practices without automated audits face financial clawback risk from RADV and payer audits, regulatory exposure under OIG oversight, and revenue leakage from both directions of coding error. Most practices underestimate all three because manual sampling hides the actual scope.

RADV and payer takeback risk. CMS finalized a rule in January 2023 confirming it will apply extrapolated RADV recoveries to Medicare Advantage payment years going back to 2018, without the fee-for-service adjuster the industry had argued for. That single policy shift means a pattern of unsupported codes found in a small sample of audited charts can be extrapolated across the plan's full risk-adjusted population for that year, well beyond the charts an auditor actually pulls. Manual internal audits checking a small percentage of charts per quarter do not have the statistical power to catch this before an external auditor does.

OIG exposure for pattern-level over-coding. OIG has published a series of Medicare Advantage risk adjustment reports dating back to at least 2016 and continuing through recent work plan cycles, repeatedly flagging diagnoses that drove risk adjustment payments but appeared only on chart reviews or health risk assessments, never in the treating clinician's own encounter documentation. A consistent pattern tied to a specific clinician, template, or copy-forward habit can be read as a compliance program failure, not an isolated error, which moves the conversation from repayment to program integrity review.

Under-coding as silent revenue loss. Chronic conditions actively managed but not consistently re-coded each year quietly reduce risk-adjusted revenue without anyone noticing, because nothing is technically wrong, the codes are just missing. This is common in endocrinology and primary care managing long-term diabetes complications, where the ADA Standards of Care (2024) recommend annual comprehensive eye exams and annual foot/neuropathy assessment for exactly the complication categories that drive HCC recapture, but the annual re-documentation habit lapses even when the clinical care itself continues.

Payer-specific variation. RADV governs Medicare Advantage. Commercial risk-based contracts and Medicaid managed care plans run parallel but not identical audit frameworks, often with their own evidence thresholds and look-back windows. A practice contracted across multiple risk arrangements needs an audit tool configured per payer, not a single generic rule set, because a diagnosis that passes one plan's evidence bar can fail another's.

Inconsistent internal coverage as its own liability. A practice that can only show it reviewed a small, non-representative sample can look worse during an audit than one with no internal program at all, since it suggests visibility into a risk area without a comprehensive response.

Get early access to Thyra

Built by a practicing endocrinologist. JJ personally reviews every application.

Or apply for the founding cohort →

What we've seen in early deployments

Building this into a live EHR workflow surfaces patterns that are easy to miss in the abstract. Across early endocrinology and primary care deployments, three findings have been consistent enough to shape how we prioritize review.

Recapture failures, not fabrication, are the dominant flag category. The most common flag by a wide margin is not a clinician documenting a condition that never existed. It is a chronic complication code, retinopathy and nephropathy especially, that was legitimately supported in a prior year and simply was not re-addressed this year. This tracks with what OIG and CMS RADV findings have generally described in Medicare Advantage: unsupported HCCs cluster in a handful of high-value chronic categories rather than spreading evenly across the code set.

A small number of templates account for a disproportionate share of flags. When a note template pulls a diagnosis onto the assessment automatically without prompting fresh evaluation language, that template becomes a repeat offender across every clinician who uses it. In one endocrinology deployment, a single assessment template used by four clinicians auto-populated a CKD-related HCC onto nearly every visit regardless of whether nephropathy had actually been addressed that day. Fixing the template's default behavior resolved far more flags in one change than a month of clinician-by-clinician correction.

False positives cluster around atypical but clinically sound documentation. Clinicians who reason through a plan in narrative prose rather than a checklist format get flagged more often, even when the underlying evidence is present. This is the strongest argument for keeping a human review step before any flag becomes a code change, and it is the reason vendor claims of fully automated code correction should be treated skeptically.

Flag CategoryRelative Frequency ObservedTypical Root Cause
Chronic HCC recapture gap (retinopathy, nephropathy, CHF)Most commonCondition legitimately managed but not re-documented this payment year
Template-driven auto-population without fresh evidenceCommon, concentrated in a few templatesTemplate pulls prior diagnosis into assessment automatically
Under-specified diagnosis (e.g., E11.9 vs. a complication-specific code)CommonComplication is treated but not coded to the higher-value HCC
False positive on narrative-style documentationLess common but recurringEvidence present in prose, not in a structured field the rule set expects
Outright fabricated or clinically implausible diagnosisRareClerical error or isolated documentation lapse

A decision rule for triage

Most practices do not need to review every flag with equal urgency. A workable rule: route any flag on a chronic, high-value HCC (nephropathy, retinopathy, major depressive disorder, CHF) to same-cycle human review before the claim submits, since these carry the largest RADV extrapolation exposure. Route flags on lower-value or acute codes to a batched post-bill review window instead. And when flags cluster around a single template used by multiple clinicians rather than spreading evenly across a panel, treat it as a template defect to fix once, not a series of individual chart corrections. This concentrates reviewer time on the codes and root causes that actually move audit risk.

What does AI catch that manual review misses?

Manual review, done by a certified coder or compliance auditor, is genuinely good at judgment: is this diagnosis clinically plausible, does the documentation read like real clinical reasoning versus template language. AI does not replace that judgment. It changes coverage: full-panel review of every encounter instead of a periodic sample, cross-encounter recapture tracking that flags a chronic HCC billed last year but not re-supported this year, and clinician- or template-level pattern detection that points compliance staff toward a systemic fix rather than a chart-by-chart one. The comparison below lays out where each approach is strongest.

How does this fit into an endocrinology or primary care workflow?

AI chart audits work best running against the same structured chart data the clinician already generates, rather than as a bolt-on requiring a separate export or manual pull. In Thyra, the audit layer draws from the same clinical context that powers the Longitudinal AI Scribe, Smart Inbox, and orders, because the diagnosis, medication, lab, and CGM data needed to support an HCC code already exist in structured form. This matters practically because the ONC's interoperability requirements under 21st Century Cures, and the underlying HL7 FHIR US Core data model most certified EHRs now use, are what make it possible to pull problem list, medication, and lab data automatically instead of re-parsing free text. HIMSS survey data has consistently listed coding and compliance review among the top administrative burdens cited by practice leaders, and a diabetes-heavy panel shows why: complication-related HCCs require specific supporting evidence each year, and CGM trends, lab results, and medication history already living in the chart can be checked against the billed code automatically instead of a coder manually cross-referencing three sections of a note.

The workflow most practices land on: AI runs a first-pass audit on every encounter shortly after the note is signed and the claim is generated, flags a small subset of encounters where a billed HCC lacks clear supporting evidence, and routes those flags to a coder or compliance reviewer for a final decision before the claim goes out or during a defined post-bill window. The clinician is rarely in this loop directly, except when a flag requires an addendum, which keeps the audit function from adding friction to the visit itself.

DimensionManual Chart AuditAI-Assisted Chart Audit
Typical chart coverageSmall sample per periodUp to every encounter
Cross-encounter recapture trackingRequires manual chart pulls across yearsAutomated, flags year-over-year HCC gaps
Speed to flag an issueWeeks to months (batch review cycles)Same day to same week, tied to claim generation
Judgment on clinical plausibilityStrong, human coder expertiseLimited, needs human review for edge cases
Pattern detection across clinicians/templatesPossible but labor-intensiveSystematic and continuous
Staffing cost to scale coverageRises with chart volume reviewedRises with review of flagged exceptions, not full volume
Audit trail for external RADV/payer reviewDepends on documentation disciplineConsistent, easier to produce on demand

What are the limits of AI chart audits?

AI chart audits are a screening tool, not a final compliance determination. Practices that treat flags as automatically correct without review introduce a different kind of risk.

False positives on atypical but legitimate documentation. A clinician who documents thoroughly but in an atypical structure can get flagged even when the care and coding are both correct. This is why a human review step before claim submission or during post-bill audit stays necessary, not optional.

Get early access to Thyra

Built by a practicing endocrinologist. JJ personally reviews every application.

Or apply for the founding cohort →