An AI answer that shows its sources feels checked. That feeling is the problem. A source list tells you the model consulted something. It does not tell you that the sentence sitting above it is actually in there.
In most categories that gap is survivable. In medical equipment it usually is not, because the reader is making a compatibility or suitability decision and the sentence they act on is a specification. A wrong display size is embarrassing. A wrong transducer compatibility list is a return, a delayed procedure, and a conversation with an account you spent years earning.
This article is about the specific way sourced answers go wrong, why the obvious defence does not work, and what GearDock AI does instead. Some of it is unflattering about approaches we tried and threw away, which is the only way to write this honestly.
The Failure That Looks Like Proof
The failure worth naming is the citation-shaped one. A model writes a fluent paragraph, attaches a marker to it, and the marker points at a real page published by the real manufacturer. Everything about the output signals diligence. The claim itself is simply not on that page.
It is worse than an uncited invention, and the reason is behavioural rather than technical. An unsourced number gets checked by somebody. A sourced number gets copied into a catalog. The citation does not just fail to prevent the error, it actively lowers the guard that would have caught it.
In device content it usually takes one of three shapes.
- The source is about the family, not the model. A manufacturer's ultrasound page covers several systems in one range. A figure lifted from it may belong to a different one, and the page gives no signal about which.
- The source supports part of the sentence. Half a claim carrying a citation reads exactly like a whole claim carrying a citation, and nothing in the layout distinguishes them.
- The source is real and the sentence is a summary of it whose meaning drifted. Paraphrase is where clinical overreach enters a catalog, because a sales brochure paraphrasing an instructions-for-use document is already one step away from what the document permits.
Why Matching Words Cannot Catch It
The obvious defence is to require the words of a sentence to appear in the source it cites. It is deterministic, it costs nothing to run, and it is measurably wrong. Two statements, both checked against the same real manufacturer page during development:
| Statement | What it actually is | Word overlap |
|---|---|---|
| Push handles at both ends make the unit easy to move between rooms | A faithful paraphrase of the source | 31% |
| The Vivid T9 is manufactured in a solar-powered facility in Norway | Entirely invented | 36% |
There is no threshold that separates those two, because the thing that differs between them is meaning and the thing being measured is vocabulary. No amount of tuning fixes that. Only a reader can tell them apart.
The consequence showed up in the drafts long before anyone traced it back to the rule. The sentences that survived were the ones copied off the page, so a video library blurb outranked every sentence a writer had composed. Copying was the only reliable way to score, and the output read exactly like that.
A Different Byte Is Not a Different Claim
What Runs Before an Answer Reaches You
The writing step is the only part of GearDock that reaches a language model at all. Everything that decides what is true is ordinary deterministic code, and it runs on both sides of the model.
Resolve What Is Being Asked About
Nothing is written until the request maps to a real product. GearDock recognizes 77 manufacturers and 460 product families, and it reads the identifier shapes this industry actually uses, which are not consistent between vendors. When the request is genuinely ambiguous it asks one focused question rather than picking the likeliest reading and proceeding.
Find Sources in a Fixed Order
Documents in your own workspace are read first, because a file somebody on your team attached and reviewed outranks anything found on the open web. Beyond that, the manufacturer's own sitemap is tried before any search, because it needs no credential, no index, and no model, and it answers the question directly.
Write With the Markers Still Attached
The model is asked to end every factual sentence with the marker of the passage supporting it. Language is free at this stage. Wording like designed for and delivers is the writer's business. Numbers, identifiers, manufacturers, and regulatory, clinical, safety, compatibility and performance statements are not.
Judge Each Factual Atom on Its Own
Statements are pulled out of the prose one at a time and judged against the evidence cited for them, sorted into 9 classes because a description is not adjudicated the way a regulatory status is. The output is a ledger rather than a verdict: this statement, this class, this source, supported or not, and why.
Verify the Prose That Carries No Atom
Ordinary descriptive sentences carry no figure or identifier for the deterministic layer to check, and those are exactly the sentences the word-overlap rule got wrong. They are read against the same evidence the writer saw, by a reader that can tell paraphrase from invention.
Score the Draft and Send It Back If It Falls Short
Before anyone sees it, the answer is scored on 7 dimensions: grounding, structure, clarity, specificity, readability, intent match and format match. A draft that falls short is revised. What arrives is the revision, not the first attempt.

The Check Runs on Sentences, Not on the Answer
Admission is not a pass or fail decision about the reply. Each statement is judged separately, and the ones that fail are withheld while the rest of the answer stands. That is why a GearDock answer is sometimes shorter than you expected rather than hedged. Softening an unsupported claim keeps it in the catalog. Removing it does not.
The citation markers are load-bearing right up to that moment and then they are gone. The marker is how a sentence proves itself to the ledger, and it is also the last thing anybody wants sitting in the middle of a product page they are about to paste into a storefront. So the markers go in, the proof runs, the markers come out, and the link between each admitted statement and the sources that carried it travels on separately. That link is what the source panel reads, what answers where did you get that, and what the challenge feature re-checks.
Stripping the Markers Too Early Makes the Whole Thing a No-Op
The Part Most Tools Do Not Show You
Open the panel under a draft and there are three lists, not one. The sources that carried a statement. The sources that were reviewed and set aside. And underneath both, the list that matters most: the statements the draft wanted to make and could not.

The withheld sentences in that capture are worth reading slowly, because none of them looks like a hallucination. The system is suited for clinical settings that require ultrasonic imaging and pulsed Doppler assessment. Offers quantitative flow analysis to support cardiovascular evaluations. Provides real-time visual data for clinical decision making.
Those are ordinary sentences from an ordinary product page. They are plausible, they are probably true of the device, and every one of them was removed because nothing retrieved actually said it. That is the whole behaviour in one screen. Plausible is not a source.
It is also why that draft is short: sixteen words against a target of two hundred to five hundred, four of eight buyer questions answered, and a remedy stated in plain terms, which is to attach the manufacturer's datasheet or manual. A tool optimising for a full-looking page would have written the other nine sentences, and every one of them would have read exactly as authoritative as the two it kept.
What Happens When the Checker Cannot Run
Any system that adds a verification step has to answer what it does when that step is unavailable, and most answers are worse than the question suggests. If verification failing open means unverified copy ships looking verified, the check has made things worse, because it created confidence it can stop supplying without saying so.
An Outage Makes GearDock Quieter, Never Looser
Two Confidence Numbers, Not One
A single confidence score hides the failure that matters most here. GearDock reports two, and keeping them apart is deliberate.
- Grounded coverage
- What share of the factual statements in an answer resolve to a real admitted source at all. This is the number most tools mean when they say grounded.
- Provenance accuracy
- Whether the label on that source is the right one. A manufacturer's page tagged as a regulatory record is fully covered and completely mislabelled, and coverage alone cannot see it.
An answer can score perfectly on the first and badly on the second. Reported as one number, that answer looks excellent. This is why every statement carries a plain label saying what kind of thing stands behind it: manufacturer verified, regulator verified, industry source, GearDock knowledge, your workspace, you provided this, or wording only when a sentence carries no fact at all.
How to Test Any Tool for This
None of this has to be taken on trust, from us or from anyone else. These five tests take about twenty minutes and need no access beyond a trial account.
- Ask about a product you know well, and include one specification that is genuinely not published anywhere. See whether you get a number. A tool that returns one has told you what it does with every gap you cannot check.
- Take any sourced sentence, open the source, and use the browser's find function on the key term in that sentence. Do this five times. The failures cluster in specifications and compatibility, not in prose.
- Ask about a product whose console, software tiers, and accessories carry different model numbers. Watch whether the answer keeps them apart or produces one confident blend of all three.
- Give it a document that contradicts something on the manufacturer's public page, and see which one wins and whether it tells you there was a conflict at all.
- Ask what it could not confirm. A system that tracks admission can answer this. A system that only retrieves will describe its sources again.
What This Does Not Promise
None of this makes output correct, and any vendor implying otherwise is overselling. A manufacturer can publish an error. A datasheet can be superseded and stay online. An extraction can take the wrong figure out of a dense table. Every one of those produces a wrong statement that is correctly sourced, correctly labelled, and correctly admitted.
What the machinery buys is different and more modest. When something is wrong you can find out where it came from, and you can correct it across every product it touched instead of one complaint at a time. That is an achievable promise, and it is worth keeping separate from the unachievable one.
The other limit is deliberate. GearDock AI writes. A person decides what the catalog says, and generated content sits as a draft until somebody with the right role releases it. Drafting, approval, release and export are separate steps by design, and no amount of confidence in the checking collapses them into one.
