NHS Scribes Dropped 'Null.' Patients Caught It, Not GPs.

An NHS AI scribe dropped the word 'null' and a negative test became a demyelination diagnosis. Healthwatch England found 27 scribe products in use and patients, not clinicians, catching the errors — because fluent prose is exactly what defeats a human review step.

By Rajesh Beri·September 1, 2026·16 min read
Share:
A desktop microphone standing on a GP's consulting-room desk beside a printed clinical letter, with one line struck through in red pen and the pen resting on the page.

Illustration generated using AI

A patient in England was told they had demyelination — the nerve damage underlying conditions including multiple sclerosis. The result had actually read "null demyelination." An AI scribe dropped the word that reversed the meaning, no clinician caught it, and the patient — querying the MRI result — did. That is the entire problem in one sentence: the error survived because the prose read perfectly well.

If your organisation has approved an AI summarisation rollout on the basis that a human checks the output before it is filed, you do not have a control. You have a reading exercise. And the evidence now says reading is the wrong task to give the reviewer.


What the Watchdog Actually Found

Healthwatch England, the statutory patient champion for the NHS, published findings on 31 August 2026 documenting AI scribes getting drug names and diagnoses wrong inside live consultations, with 27 different scribe products in use across the health service in England. An ambient scribe is software that listens to a clinician-patient conversation and generates the note, letter or coding suggestion that goes into the record.

Four cases, and they are worth reading as a taxonomy rather than as anecdotes:

  • A dropped negation. "Null demyelination" became demyelination. One token, and the diagnosis inverted.
  • A similar-sounding substitution. The scribe confused the prescribed drug with a different one of a similar name. The patient, not the doctor, spotted it.
  • An omitted instruction. A summary letter left out the consultant's direction to obtain a repeat prescription for migraine medication — which could have left the patient unable to get it.
  • A fabricated medication. A London GP, Dr Shier Ziser Dawood, described in the British Journal of General Practice a note recording that she had told a patient to continue Prozac, a drug she had neither prescribed nor discussed.

Note the common property. Every one of these produces a grammatical, plausible, clinically coherent sentence. None of them looks like a bug. Healthwatch's own framing is the tell: it has heard "multiple stories from patients who have noticed these errors when a health professional hasn't."

This is not confined to one health system. In Australia, a patient's post-operative letter after kidney stone surgery recorded that she micro-dosed psychedelic mushrooms — something she had never done — while she was on a workers' compensation claim. A fabricated line in a clinical record became an insurance problem. The Royal Australian College of GPs estimates around 40% of Australian GPs use scribes regularly.


The Regulator Just Made Your Review Step the Entire Control

On 29 July 2026 the MHRA clarified that ambient voice technology used solely for transcription, summarising clinical conversations, drafting letters or suggesting clinical codes for review is not regulated as a medical device. Products intended to support diagnosis, treatment or prevention — or that take automated actions such as placing orders without a clinician reviewing first — remain devices.

That guidance superseded NHS England's earlier position, set out in an April 2025 policy and a June 2025 priority notification, that any product performing clinical summarisation had to be registered as at least a Class I device. Be precise about what moved: the MHRA states that the guidance "does not change the law" and only sets out how existing device regulation applies. Nothing was deregulated. But the bar a deploying trust actually faces went down, not up, five weeks before the watchdog published its error cases.

Read the MHRA's own sentence carefully, because it is now doing all the load-bearing work: "Clinicians remain responsible for reviewing and verifying AI generated transcripts, summaries and other outputs before they are used in patient care." MHRA chief executive Lawrence Tallon framed the guidance as clarity — "where a product supports diagnosis or treatment, regulatory protections will still apply."

The practical consequence for anyone running one of these systems: there is no pre-market safety and effectiveness test standing behind a summarisation-only product. NHS England's guidance, version 3 of 27 April 2025 and last updated on the same day as the MHRA guidance, still requires a deploying organisation to produce DCB0160 documentation — a safety case, a hazard log and a monitoring framework — plus the supplier's DCB0129 clinical safety case report, a DTAC assessment, DSPT compliance, a DPIA, and a medical device determination. Your DCB0160 is the safety case. And its central hazard control is a human reading prose.

The same structural gap shows up in device-cleared AI, where 1,357 FDA-cleared AI devices had three tested on patient outcomes, and in the finding that fewer than half of health systems have a sandbox to test vendor AI at all. The regulatory question and the local-evidence question keep coming apart.


Negation Is 30% of Hallucinations — and the One a Reader Cannot Catch

The "null demyelination" case is not an outlier. It is the second-largest hallucination class in the best clinician-annotated dataset published on this problem.

Asgari and colleagues, in npj Digital Medicine, had clinicians label LLM-generated clinical notes across 18 experimental configurations covering 49,590 transcript sentences and 12,999 note sentences. They measured a 1.47% hallucination rate — 191 sentences — of which 44% were rated major, meaning they could change the patient's diagnosis or management if left uncorrected. The omission rate was 3.45% (1,712 instances), with 16.7% major.

Two caveats belong on the same line as those numbers, and neither is in the coverage. The notes were generated by GPT-4-32k-0613 over PriMock, a public corpus of mock primary-care consultations — not by any of the 27 deployed products, and not on real patients. And every author was an employee of Tortus AI, an ambient scribe vendor selling into the NHS (the paper states the affiliation and then declares no competing interests). That does not make the taxonomy wrong — it is still the best clinician-annotated map of these failure classes that exists, and a vendor measuring a generic GPT-4 baseline has more reason to understate the problem than to inflate it. It does mean you should read what follows as the shape of the risk, not as a measured rate for the product on your desk.

The breakdown of those 191 hallucinations is the part to put in front of your clinical safety officer:

Hallucination type Count Share
Fabricated 82 43%
Negation 56 30%
Contextual 33 17%
Causality 20 10%

A negation error is one where the model reverses the polarity of a clinically relevant fact — "no chest pain" becomes "chest pain", "null demyelination" becomes "demyelination". The authors found hallucinations concentrated in the plan section, contradicting what was actually said in the consultation, with 21% of major hallucinations landing there. Their explanation of why this class is uniquely dangerous is the single most important sentence in the literature for anyone designing a review step: without the full context of the consultation, readers may struggle to discern which negation is true and which is false.

That is not a statement about careless reviewers. It is a statement that the information needed to catch the error is not present in the artifact being reviewed. You cannot proofread your way out of it.


Reading the Summary Is Not a Control. Here Is the Evidence.

Every measurement of human review over AI-generated clinical text points the same way: it catches some errors, misses the consequential ones, and degrades with familiarity.

Errors survive the pipeline. A Mayo Clinic Proceedings: Digital Health study ran 14 simulated ambulatory encounters through five ambient scribe platforms. Transcripts from four of them averaged 13.9 errors per case, 19.5% of which were transmitted into the clinical note. Mean error across key clinical elements in the notes was 26.3%. An average of three errors per case had potential to cause moderate-to-severe harm on the AHRQ scale. Omissions were roughly three-quarters of all errors — and an omission is the one thing a reviewer reading only the output can never see. The authors' own verdict on the safeguard is blunt: clinician proofreading remains the de jure safeguard against these errors, but early simulation studies suggest it may not be an effective one.

Clinicians know it and still ship it. A survey of 1,003 UK GPs in August 2025, published in BMJ Health & Care Informatics, found 14% currently using scribes and 39% intending to. Among the 141 who use them, 44% find errors in 10–30% of AI-generated documents, 32% encounter errors often or always, and 14% have seen errors with significant to critical implications for care. It is a self-selected Doctors.net.uk panel and its authors call the study exploratory, so that is clinician self-report on a small base, not a measured rate. Reported error rates spike exactly where enterprise deployments live: multi-party conversations (38%), complex histories (35%), languages other than English (31%), speech impairment (21%) and noisy environments (20%).

The time budget for review is tiny. A study of 1,800 clinicians across five academic medical centres found scribes saved about 16 minutes per eight-hour shift, and 13 fewer minutes in the EHR during working hours. That is the entire budget. Instruct clinicians to carefully re-read every generated note against the conversation and you have spent the benefit and then some — which is precisely why they will not do it.

Reviewers may get worse over time. In the first real-world clinical evidence of deskilling in medicine, 1,443 non-AI-assisted colonoscopies across four Polish centres showed adenoma detection falling from 28.4% before routine AI introduction to 22.4% after — a 20% relative decline, published in The Lancet Gastroenterology & Hepatology in August 2025. Take this one as suggestive rather than settled. It is observational and before-and-after; its authors say plainly that factors other than the AI rollout may have influenced the finding; and it drew several critical letters in the same journal. It tells you the direction of a risk, not an effect size you can plan against.

None of this is specific to medicine. We have seen the same shape in software: the vaguer an AI explanation was, the more novice reviewers trusted it, ChatGPT made requirements inspectors measurably worse at finding defects, and one vendor writing and reviewing 208,145 pull requests is not two controls. Fluent output plus a tired reviewer is a failure mode, not an oversight regime.


Steel-Manning the Rollout

The case for scribes is real and should not be dismissed by anyone reading this. Documentation burden is a genuine driver of clinician attrition, burnout improvements from ambient scribes have been larger than the raw time savings, and the same npj study that measured a 1.47% hallucination rate also reported that iterating on prompts and workflow drove major errors below documented human note-taking performance — a claim to weigh alongside the fact that its authors sell scribes — human notes average at least one error and four omissions each.

The honest reading is not "AI scribes are unsafe." It is that the comparator changed. Human note-taking errors are mostly omissions of things nobody wrote down. LLM errors include confident assertions of things that were never said, in fluent prose, in the plan section, at scale, across 27 products with no common monitoring. A control designed for the first failure mode does not cover the second.


Stop Reviewing Prose. Start Reconciling Fields.

The fix is not asking reviewers to read more carefully. It is changing what you hand them: a short list of extracted values to confirm, not a paragraph to proofread.

Concretely, the pipeline should emit two artifacts, not one. The prose summary is the deliverable. Alongside it, a structured object containing every fact that can be wrong in a way that matters:

  1. Every clinical or commercial entity as a discrete field — drug, dose, route, frequency; or in a claims context, policy number, date of loss, coverage decision, amount. Similar-sounding substitutions are trivially catchable as a value mismatch and nearly invisible as prose.
  2. Every negation as an explicit polarity flag on a named target. demyelination: absent. A reviewer confirming a checkbox catches a flipped polarity. A reviewer reading a sentence does not.
  3. Every instruction as an owner, an action and a deadline. The migraine repeat prescription was not a wrong fact; it was a missing row. Omissions only become visible when the schema has a slot for them.
  4. A source span for every extracted value. Any field the model cannot point back to a location in the transcript is a candidate fabrication and should be surfaced as one. This is the same discipline as grounding a RAG answer in retrieved passages — see the trade-offs in buying the index and building the eval set.
  5. Surface only the deltas. The reviewer's screen should show six fields to confirm, not 400 words to scan. Humans are good at confirming discrete values against a source. They are bad at detecting a missing word in fluent text.

Two things make this possible and both are procurement items. You must retain the source transcript and audio for at least as long as the record they generated — NHS England's guidance asks organisations to establish retention for interaction audio, transcripts, user interactions and outputs, but does not mandate a period, so it is a question you have to ask and write into the contract. Without the source there is no diff, no audit, and no way to adjudicate a patient's dispute. And you must be able to get structured output with spans, not just formatted prose, which several products do not offer.


What to Ask a Vendor Before You Sign

Ask for per-failure-class numbers, not an accuracy percentage. A single accuracy figure averages across the classes that do not hurt you and the classes that do — the same reason an average agent score of 73 hid a 30% usable rate.

  • What is your intended-purpose statement, and what determination did it produce? Under the 29 July 2026 guidance, if the answer is "summarisation only", the product is not a device and no regulator has tested it. That is a fact to plan around, not an objection.
  • Give us measured rates for: dropped negation, entity substitution, omitted instruction, speaker misattribution, and hypothetical-stated-as-fact. Five numbers, on your data, with the annotation protocol.
  • Provide the DCB0129 clinical safety case report and the hazard log it is built on.
  • Does the API return structured fields with transcript spans, or only prose?
  • What are the retention controls for audio, transcript and intermediate output, and can we set them?
  • What happens on the hard cases — three speakers, an interpreter, a strong regional accent, a noisy room? Those are the documented hotspots, so they belong in the pilot, not in production.

Build the eval set from the failure taxonomy rather than from happy-path recordings; the largest catalogue of enterprise AI failures found hallucination is not the top category, and your classes will be specific to your domain. Score pass/fail per class and gate releases on it — the mechanics are the same as any LLM regression gate you would build with Promptfoo, Langfuse or Braintrust, and the tooling (Langfuse, Braintrust) is not the hard part. Defining the classes is.


This Is Not a Healthcare Story

Any workflow where a model turns a conversation or a document into a record that someone later relies on has this exact failure surface. Meeting-notes products that write to CRM — the category Granola turned into enterprise infrastructure — commit the same negation errors, and "the customer did not agree to the renewal terms" is a sentence with real money in it. Claims adjudication summaries feed coverage decisions. Contract-review summaries feed signature. Contact-centre wrap-ups feed dispute records. Incident post-mortems feed regulators.

In every one of those, the standing control is a human who reads the output. In every one, the reviewer is reading fluent text without the source in front of them. And in regulated manufacturing, the discipline already exists and is called reconciliation — validating an instrument integration against its change control means diffing values against a source of truth, not reading a report and nodding. Even a model that reports 99% confidence needs to be sampled repeatedly before you trust the number.


What to Do Now

This Week:

  1. Inventory every place in your organisation where a model turns speech or a document into a record a human later acts on. Include the ones bought on a corporate card.
  2. For each, write down in one sentence what the actual control is. If the sentence is "a human reviews the output", flag it.
  3. Confirm whether you still hold the source transcript for those systems. If you do not, you cannot reconcile anything and that is the first fix.

This Month:

  1. Build a 50-item eval set from the five failure classes above, drawn from your own recordings, including at least ten hard cases — multi-party, accented, noisy, non-native-language.
  2. Run it against your incumbent product and report pass/fail by class. Expect the negation and omission numbers to be the ugly ones.
  3. Change one review screen from "read the summary" to "confirm these eight extracted fields against the highlighted transcript spans" and measure detection against your seeded errors.

Before Renewal:

  1. Put per-failure-class reporting and structured-output-with-spans into the contract as a deliverable, not an aspiration.
  2. Set audio and transcript retention explicitly and align it with the retention of the record produced.
  3. Establish a correction path for the subject of the record — the patient, the customer, the counterparty. Healthwatch's central ask is that patients can report and get errors corrected; the same right is your cheapest detection channel, because the subject of a record is the one reader who knows what actually happened.

The Bottom Line

The public has already priced this. In a YouGov poll of 4,039 UK adults for Healthwatch between 16 and 27 April 2026, nearly 90% of people who had an appointment in the past 12 months did not know AI scribes were in use, 81% wanted to be told and asked, and 69% said they would feel more comfortable if professionals committed to verifying accuracy. Strong opposition ran at 21% against 11% strong support. That commitment to verify is the one being made on your behalf, in every deployment, by a reviewer who has 16 minutes a shift and a paragraph that reads fine.

We built a control on the assumption that a wrong answer looks wrong. The whole achievement of these models is that it does not.

Stop asking people to read the summary. Give them the fields, and the transcript, and the diff.

Continue Reading

FDA Cleared 1,357 AI Devices. Three Were Tested on Patients. Hospitals Test Vendor AI. Fewer Than Half Have a Sandbox. The Vaguer the AI Explanation, the More Novices Trusted It ChatGPT Made Spec Reviewers Worse. Save It for Round 2. Same Vendor Wrote and Reviewed 208,145 PRs. Split Them. 10,000 AI Failures Exposed. Hallucination Isn't #1. Braintrust vs Langfuse vs Promptfoo: Don't Pay for the Gate 4 Instruments, 8 Hours. The Weeks Were Your Change Control.

Share:

Frequently Asked Questions

Why did clinicians miss the AI scribe errors that patients caught?

Because the errors produced fluent, plausible prose. A dropped negation, a similar-sounding drug substitution or an omitted instruction all read as a grammatical, clinically coherent sentence. Researchers in npj Digital Medicine found hallucinations concentrated in the plan section and noted that without the full context of the consultation, readers struggle to tell which negation is true and which is false. The information needed to catch the error is not in the artifact being reviewed.

Are AI scribes regulated as medical devices in the UK?

Not if they only transcribe and summarise. MHRA guidance published on 29 July 2026 confirmed that ambient voice technology used solely for transcription, summarising clinical conversations, drafting letters or suggesting clinical codes for clinician review is not a medical device. Products that support diagnosis or treatment, or that take automated actions without clinician review, remain regulated. The guidance overrode NHS England's earlier position that summarisation required at least Class I registration.

How often do AI clinical scribes make errors?

A clinician-annotated study of GPT-4-generated notes from mock primary-care consultations measured a 1.47% hallucination rate across 12,999 note sentences, with 44% rated major, and a 3.45% omission rate. A separate test of five ambient scribe platforms across 14 simulated encounters found a 26.3% mean error rate on key clinical elements and about three errors per case with potential for moderate-to-severe harm. In a survey of 1,003 UK GPs, among the 141 who currently use scribes, 44% reported finding errors in 10 to 30% of AI-generated documents.

What should replace reading the AI summary as a review step?

Field-level reconciliation. Have the pipeline emit a structured object alongside the prose containing every entity and value, every negation as an explicit polarity flag on a named target, and every instruction as an owner, action and deadline — each with a source span in the transcript. Surface only the deltas, so the reviewer confirms a handful of discrete values against highlighted source text rather than proofreading a paragraph.

Does this risk apply outside healthcare?

Yes. Any workflow where a model turns a conversation or document into a record someone later relies on has the same failure surface: meeting notes written to CRM, claims adjudication summaries, contract-review summaries, contact-centre wrap-ups and incident post-mortems. A dropped negation in 'the customer did not agree to the renewal terms' carries the same shape of harm, and the standing control is usually the same human reading fluent text without the source in front of them.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Related Articles

Latest Articles

View All →