The Joy of Being Confident and Wrong
There is a certain Zen-like beauty in the way a Large Language Model (LLM) evaluates a clinical note. It looks at a wall of text, sees that the patient has two arms and a heartbeat, and gives the whole thing a gold star because it found exactly what it was looking for. The fact that the doctor forgot to mention the patient is currently experiencing an active heart attack is irrelevant. If it isn't on the page, it isn't in the reality of the judge.
This is the "Hallucination of Completeness," a fancy term for what most of us call being spectacularly unobservant. In a recent study regarding AI-generated clinical summaries, automated judges were found to be excellent at verifying that the information present was accurate. However, they were statistically abysmal at noticing when critical, life-saving data was missing entirely. It turns out that when you train a machine to find needles in haystacks, it gets so excited about the needles that it forgets the haystack is supposed to contain a tractor, too.
We are currently building a feedback loop where AI models grade other AI models based on a shared delusion. The student turns in a half-finished exam, and the teacher—who is also a robot with a short attention span—gives it an A because the three sentences that were written had excellent grammar. It is a closed system of automated incompetence that feels very efficient until someone ends up in the wrong operating room.
The Professional Art of Looking Busy
In high-stakes environments like medicine or law, we have historically relied on humans to notice the silence. A seasoned nurse notices the medication that wasn't ordered; a cynical lawyer notices the clause that wasn't included in the contract. But humans are expensive and they require things like "sleep" and "dignity." AI, on the other hand, will happily scan 10,000 documents a second and tell you they look great because it has no concept of what a void looks like.
- The AI judge confirms the presence of a 'Treatment Plan.'
- It ignores that the 'Treatment Plan' is just the word 'Aspirin' repeated four times.
- It gives a high 'faithfulness' score because 'Aspirin' was indeed mentioned in the source text.
- The hospital administrator buys a new yacht with the money saved on auditing.

Photo by https://kaboompics.com/ on Pexels
This creates a dangerous incentive structure. If the automated auditor only checks for the presence of specific keywords, the generative AI will learn to pepper those keywords throughout its output like digital seasoning. It doesn't matter if the narrative makes sense or if the patient's history is a fractured mess of contradictions. As long as the 'Judge' sees the word 'History' and 'Physical,' it marks the task as complete. We are essentially teaching computers to be the ultimate middle managers: obsessed with metrics, oblivious to results.
Why Reality Is a Low-Priority Update
If you ask an LLM to judge a summary, it performs a pattern match. It is looking for overlap. If the summary says 'The cat sat on the mat' and the source says 'The cat sat on the mat,' the judge is thrilled. If the source says 'The cat sat on the mat and then the house blew up,' and the summary leaves out the explosion, the judge still sees the cat and the mat and thinks everything is fine. To the AI, the explosion is just 'extra noise' it wasn't specifically told to look for.
This is particularly fun when you realize that we are using these judges to 'fine-tune' the next generation of models. We are literally breeding AI to be more omit-happy. We are selecting for the versions of the software that are the best at hiding their own ignorance. It’s like a corporate culture where nobody is allowed to report bad news, so the quarterly reports always look fantastic right up until the company vanishes in a puff of logic.
By the time we hit 2025, we could have clinical systems that produce beautiful, poetic, and entirely useless documentation. The notes will be formatted perfectly. The billing codes will be extracted with surgical precision. The only minor issue will be that the actual medical reality of the patient has been edited out to make room for a high 'Completeness Score' from an algorithm that wouldn't know a missing limb from a missing comma.
What This Actually Means
We are currently mistaking "consistency" for "accuracy." Just because two AI models agree that a document is good doesn't mean the document is actually good; it just means they are both hallucinating in the same direction. When we remove the human from the loop because the AI 'judges' say the AI 'writer' is doing a great job, we aren't automating quality control. We are automating a cover-up.
In professional settings, the most important information is often the stuff that isn't there—the negative test result, the skipped dosage, the lack of a pulse. If our auditing tools are biologically incapable of seeing a vacuum, they are useless for safety. We are essentially installing smoke detectors that only go off if they see a fire extinguisher. It’s a very clean, very quiet way to let a building burn down.
Ultimately, this is a call for a return to skepticism. If a tool tells you a complex medical record is '100% verified,' it probably just means the tool found the font agreeable. We need to stop treating AI 'judgment' as a benchmark and start treating it as a very fast, very confident intern who has never actually seen a patient but is really good at filling out forms.
Quick Answers
Does this mean AI clinical notes are always wrong?
No, it means they are selectively right, which is actually more dangerous because it lures you into a false sense of security before failing on the one detail that matters.
Can't we just tell the AI to look for missing things?
We can try, but since LLMs operate on probability and presence, asking them to find a 'nothing' is like asking a flashlight to find a shadow; the moment it looks, the shadow disappears into a hallucination of something else.
Is there a fix for omission blindness?
Not a total one, but it involves using humans to audit the auditors, which everyone hates because it reveals that the 'efficiency' we bought was actually just a pile of unread mistakes.



