Skip to content
GearDock

AI Content Workflows

Why an AI Answer With Sources Can Still Be Wrong

A source list tells you a model consulted something. It does not tell you the sentence above it is in there. That gap has a shape, and it can be closed.

The GearDock Team11 min read

An AI answer that shows its sources feels checked. That feeling is the problem. A source list tells you the model consulted something. It does not tell you that the sentence sitting above it is actually in there.

In most categories that gap is survivable. In medical equipment it usually is not, because the reader is making a compatibility or suitability decision and the sentence they act on is a specification. A wrong display size is embarrassing. A wrong transducer compatibility list is a return, a delayed procedure, and a conversation with an account you spent years earning.

This article is about the specific way sourced answers go wrong, why the obvious defence does not work, and what GearDock AI does instead. Some of it is unflattering about approaches we tried and threw away, which is the only way to write this honestly.

The Failure That Looks Like Proof

The failure worth naming is the citation-shaped one. A model writes a fluent paragraph, attaches a marker to it, and the marker points at a real page published by the real manufacturer. Everything about the output signals diligence. The claim itself is simply not on that page.

It is worse than an uncited invention, and the reason is behavioural rather than technical. An unsourced number gets checked by somebody. A sourced number gets copied into a catalog. The citation does not just fail to prevent the error, it actively lowers the guard that would have caught it.

In device content it usually takes one of three shapes.

  • The source is about the family, not the model. A manufacturer's ultrasound page covers several systems in one range. A figure lifted from it may belong to a different one, and the page gives no signal about which.
  • The source supports part of the sentence. Half a claim carrying a citation reads exactly like a whole claim carrying a citation, and nothing in the layout distinguishes them.
  • The source is real and the sentence is a summary of it whose meaning drifted. Paraphrase is where clinical overreach enters a catalog, because a sales brochure paraphrasing an instructions-for-use document is already one step away from what the document permits.

Why Matching Words Cannot Catch It

The obvious defence is to require the words of a sentence to appear in the source it cites. It is deterministic, it costs nothing to run, and it is measurably wrong. Two statements, both checked against the same real manufacturer page during development:

Word overlap against the same manufacturer page. The invention scores higher than the truth.
StatementWhat it actually isWord overlap
Push handles at both ends make the unit easy to move between roomsA faithful paraphrase of the source31%
The Vivid T9 is manufactured in a solar-powered facility in NorwayEntirely invented36%

There is no threshold that separates those two, because the thing that differs between them is meaning and the thing being measured is vocabulary. No amount of tuning fixes that. Only a reader can tell them apart.

The consequence showed up in the drafts long before anyone traced it back to the rule. The sentences that survived were the ones copied off the page, so a video library blurb outranked every sentence a writer had composed. Copying was the only reliable way to score, and the output read exactly like that.

A Different Byte Is Not a Different Claim

The same lesson turns up in miniature. A statement reading 21.5-inch high-definition display was admitted, and the identical statement written with a non-breaking hyphen was refused as unconfirmed against evidence that stated it in as many words. Models trained on typeset text produce that character constantly, and medical specifications are hyphenated almost without exception: screen sizes, voltage ranges, gauge sizes, lead counts. Folding those characters together before comparing is not a nicety. Without it, most of what a specification page is made of gets refused for the wrong reason.

What Runs Before an Answer Reaches You

The writing step is the only part of GearDock that reaches a language model at all. Everything that decides what is true is ordinary deterministic code, and it runs on both sides of the model.

  1. Resolve What Is Being Asked About

    Nothing is written until the request maps to a real product. GearDock recognizes 77 manufacturers and 460 product families, and it reads the identifier shapes this industry actually uses, which are not consistent between vendors. When the request is genuinely ambiguous it asks one focused question rather than picking the likeliest reading and proceeding.

  2. Find Sources in a Fixed Order

    Documents in your own workspace are read first, because a file somebody on your team attached and reviewed outranks anything found on the open web. Beyond that, the manufacturer's own sitemap is tried before any search, because it needs no credential, no index, and no model, and it answers the question directly.

  3. Write With the Markers Still Attached

    The model is asked to end every factual sentence with the marker of the passage supporting it. Language is free at this stage. Wording like designed for and delivers is the writer's business. Numbers, identifiers, manufacturers, and regulatory, clinical, safety, compatibility and performance statements are not.

  4. Judge Each Factual Atom on Its Own

    Statements are pulled out of the prose one at a time and judged against the evidence cited for them, sorted into 9 classes because a description is not adjudicated the way a regulatory status is. The output is a ledger rather than a verdict: this statement, this class, this source, supported or not, and why.

  5. Verify the Prose That Carries No Atom

    Ordinary descriptive sentences carry no figure or identifier for the deterministic layer to check, and those are exactly the sentences the word-overlap rule got wrong. They are read against the same evidence the writer saw, by a reader that can tell paraphrase from invention.

  6. Score the Draft and Send It Back If It Falls Short

    Before anyone sees it, the answer is scored on 7 dimensions: grounding, structure, clarity, specificity, readability, intent match and format match. A draft that falls short is revised. What arrives is the revision, not the first attempt.

The GearDock AI workspace showing a short drafted product description for a GE HealthCare Vivid T9 ultrasound system, with a key features list and a two row specification table. A line underneath records that it was written from one official record with nine details left out.
A real draft, as it arrives. Two admitted specifications, and a line underneath saying it was written from one official record with nine details left out.

The Check Runs on Sentences, Not on the Answer

Admission is not a pass or fail decision about the reply. Each statement is judged separately, and the ones that fail are withheld while the rest of the answer stands. That is why a GearDock answer is sometimes shorter than you expected rather than hedged. Softening an unsupported claim keeps it in the catalog. Removing it does not.

The citation markers are load-bearing right up to that moment and then they are gone. The marker is how a sentence proves itself to the ledger, and it is also the last thing anybody wants sitting in the middle of a product page they are about to paste into a storefront. So the markers go in, the proof runs, the markers come out, and the link between each admitted statement and the sources that carried it travels on separately. That link is what the source panel reads, what answers where did you get that, and what the challenge feature re-checks.

Stripping the Markers Too Early Makes the Whole Thing a No-Op

This ordering is easy to get wrong and expensive to get wrong quietly. If citation markers are cleaned out of the draft before adjudication rather than after it, every statement arrives at the check with no citation, every check passes vacuously, and the firewall reports success while doing nothing at all. The output looks identical. That is the trap: a governance layer that has been switched off produces cleaner-looking copy than one that is working.

The Part Most Tools Do Not Show You

Open the panel under a draft and there are three lists, not one. The sources that carried a statement. The sources that were reviewed and set aside. And underneath both, the list that matters most: the statements the draft wanted to make and could not.

An expanded Sources and details panel. It names the manufacturer, product line and model that were established, lists one source used and three reviewed but not used, and then lists nine withheld statements under a heading stating that the retrieved sources do not confirm them.
One official record used, three reviewed and set aside, and nine statements withheld. Each withheld sentence is quoted back in full rather than summarised as a count.

The withheld sentences in that capture are worth reading slowly, because none of them looks like a hallucination. The system is suited for clinical settings that require ultrasonic imaging and pulsed Doppler assessment. Offers quantitative flow analysis to support cardiovascular evaluations. Provides real-time visual data for clinical decision making.

Those are ordinary sentences from an ordinary product page. They are plausible, they are probably true of the device, and every one of them was removed because nothing retrieved actually said it. That is the whole behaviour in one screen. Plausible is not a source.

It is also why that draft is short: sixteen words against a target of two hundred to five hundred, four of eight buyer questions answered, and a remedy stated in plain terms, which is to attach the manufacturer's datasheet or manual. A tool optimising for a full-looking page would have written the other nine sentences, and every one of them would have read exactly as authoritative as the two it kept.

What Happens When the Checker Cannot Run

Any system that adds a verification step has to answer what it does when that step is unavailable, and most answers are worse than the question suggests. If verification failing open means unverified copy ships looking verified, the check has made things worse, because it created confidence it can stop supplying without saying so.

An Outage Makes GearDock Quieter, Never Looser

A verifier that is unconfigured, times out, errors, or returns something unreadable produces no verdicts at all. Every statement keeps the refusal it already had from the deterministic layer. The visible result of an outage is a shorter answer, never a longer one that skipped its checks.

Two Confidence Numbers, Not One

A single confidence score hides the failure that matters most here. GearDock reports two, and keeping them apart is deliberate.

Grounded coverage
What share of the factual statements in an answer resolve to a real admitted source at all. This is the number most tools mean when they say grounded.
Provenance accuracy
Whether the label on that source is the right one. A manufacturer's page tagged as a regulatory record is fully covered and completely mislabelled, and coverage alone cannot see it.

An answer can score perfectly on the first and badly on the second. Reported as one number, that answer looks excellent. This is why every statement carries a plain label saying what kind of thing stands behind it: manufacturer verified, regulator verified, industry source, GearDock knowledge, your workspace, you provided this, or wording only when a sentence carries no fact at all.

How to Test Any Tool for This

None of this has to be taken on trust, from us or from anyone else. These five tests take about twenty minutes and need no access beyond a trial account.

  1. Ask about a product you know well, and include one specification that is genuinely not published anywhere. See whether you get a number. A tool that returns one has told you what it does with every gap you cannot check.
  2. Take any sourced sentence, open the source, and use the browser's find function on the key term in that sentence. Do this five times. The failures cluster in specifications and compatibility, not in prose.
  3. Ask about a product whose console, software tiers, and accessories carry different model numbers. Watch whether the answer keeps them apart or produces one confident blend of all three.
  4. Give it a document that contradicts something on the manufacturer's public page, and see which one wins and whether it tells you there was a conflict at all.
  5. Ask what it could not confirm. A system that tracks admission can answer this. A system that only retrieves will describe its sources again.

What This Does Not Promise

None of this makes output correct, and any vendor implying otherwise is overselling. A manufacturer can publish an error. A datasheet can be superseded and stay online. An extraction can take the wrong figure out of a dense table. Every one of those produces a wrong statement that is correctly sourced, correctly labelled, and correctly admitted.

What the machinery buys is different and more modest. When something is wrong you can find out where it came from, and you can correct it across every product it touched instead of one complaint at a time. That is an achievable promise, and it is worth keeping separate from the unachievable one.

The other limit is deliberate. GearDock AI writes. A person decides what the catalog says, and generated content sits as a draft until somebody with the right role releases it. Drafting, approval, release and export are separate steps by design, and no amount of confidence in the checking collapses them into one.

The Guided Content workflow in GearDock, showing four numbered stages: choose product, confirm sources, create draft, review. A line at the foot of the page states that every draft goes to Reviews first and that this page cannot approve, finalize, release, or export content.
The separation, stated on the screen that does the drafting. Confirm sources comes before create draft, and the page that writes cannot be the page that releases.

Private Beta

Try This on Your Own Catalog

Reading about a workflow only goes so far. Bring a slice of your products and the documents you hold for them, and we will run it with you.