Check your company’s visibility in AI
Free, with no credit card required.
AI recommends your company but adds that a product feature is available in every plan. The link leads to a real page. When you open it, however, the feature is available only in a particular package. Your brand earned a mention and a citation, but the customer received an incorrect promise.
This requires assessment of a specific claim rather than a source counter. This article examines whether answers agree with the materials they reference. It is neither a general visibility guide nor a study of model error rates. The examples are hypothetical. Source documentation was checked on 5 October 2026.
Separate factual accuracy from citation correctness
Ask two questions: “Is the information true under the conditions being assessed?” and “Does the referenced material actually support it?” A statement can be true while its citation is misplaced. A source may also support old information that no longer describes the current offering.
In Enabling Large Language Models to Generate Text with Citations, the researchers distinguish answer correctness from citation quality. We use that distinction without extrapolating their 2023 experimental results to current products.
Record the link’s role too. A reference beside a sentence, a combined source list and a “related” resource may have different meanings. The Gemini Apps documentation explains that the app may show sources or related content and that not every answer includes such links. Do not automatically assign the entire list to every sentence.
Choose claims that matter to customers
Start with information customers may use to decide: feature scope, plan conditions, compatibility, market availability or implementation requirements. A general judgement such as “a convenient tool” needs different criteria from a verifiable fact about an integration.
Split sentences containing several promises into separate claims. “The product supports integration and guarantees a support reply within 24 hours” contains two things to check. Otherwise, a correct part may conceal an incorrect one.
Define the audit’s boundaries: questions, AI products, modes, languages, market and period. Preserve the original question, full answer and date. Asking for sources afterwards changes the conversation; save that result as a subsequent step rather than as the original answer’s sources. Label API results as API results without equating them with the user-facing app.
Build a claim and evidence record
The assessment unit is one claim together with its assigned sources. If several links jointly support the statement, assess their combined content and retain all URLs. More references do not mean more assessed claims.
- Context: question, AI product, mode, language, market and answer date.
- Claim: exact wording and the conditions to which it applies.
- Attribution: assigned URLs and how they are connected to the sentence.
- Evidence: a short source extract, its location on the page and retrieval date.
- Current state: official material describing current conditions and its owner.
- Assessment: source support, factual accuracy in context and reasoning.
- Next steps: priority, responsible person, correction and recheck method.
Read the content, not just the page title or search snippet. Check the product name, version, market, plan and qualifications. If you cannot reliably connect a combined source list to a sentence, record uncertain attribution instead of inventing evidence.
Assess sources using explicit rules
These labels are a proposal for your own audit. They are neither Semly metrics nor an official scale used by AI providers. Assess the complete claim and its material conditions.
| Citation assessment | When to use it | What to record |
|---|---|---|
| Supported | The assigned materials support the whole claim in the assessed context. | Passages supporting the fact and its conditions. |
| Partially supported | They support part of the claim but not its full scope or an important condition. | What is supported and what is missing. |
| Contradicted | The cited source provides information inconsistent with the claim. | The exact discrepancy, version and date. |
| Not supported by the source | The materials are readable but provide no evidence for the claim. | The missing basis; do not automatically treat it as falsehood. |
| No source provided | The answer provides no material that can be assigned to the claim. | The absence of a citation; assess factual accuracy separately. |
| Unverifiable | Lack of access, uncertain attribution or a missing necessary version prevents a conclusion. | The reason and missing evidence. |
Have a second person review disputed cases. If you change the classification rules, record the new version and reassess the materials being compared. This prevents a change in labels alone from creating a change in results.
Work through an example involving service conditions
Imagine fictional software whose current documentation makes an integration available only in the Pro plan. Its support policy describes a usual first-response time of up to 48 business hours and no guaranteed SLA. The general product page does not describe file exports. This is a training example, not a description of Semly or another real company.
| AI claim | Assigned source content | Audit finding |
|---|---|---|
| The integration works in every plan. | The integration is available only in the Pro plan. | Contradicted; “every plan” changes the scope of the offering. |
| A support reply is guaranteed within 24 hours. | Usually within 48 business hours; no guaranteed SLA. | Contradicted; the time changed and a guarantee was added. |
| The product supports CSV export. | The general product description does not mention exports. | Not supported by this source; check the feature documentation separately. |
Do not mark the final claim as false simply because the linked page does not discuss it. Another resource may confirm the export feature. The fact would then be true, while the citation remains misplaced.
Today’s page content does not prove what was on that page when the answer was generated. Without a preserved version, assess agreement with today’s material and state the limitation of the historical assessment.
Correct the cause and preserve the decision trail
If your documentation is incorrect or ambiguous, improve it where the product owner maintains it. Put the condition beside the promise rather than hiding an exception in a distant FAQ. Where several resources describe the same offering, establish a shared source of current facts; the division of roles between a blog and a knowledge base can help.
If the issue lies in external material, prepare a specific correction with evidence for its author. If your source is correct but the answer distorts its conditions, retain the example and consider reporting it through the product’s available channel. Do not assume a report or page edit will immediately change answers.
Prioritise by the consequence for the customer. Promising a nonexistent integration requires a different response from using an imprecise adjective. Assign an owner and a recheck date to the correction. Keep the new answer alongside the previous one rather than overwriting the history.
Report citation quality separately from visibility
In Semly, you can inspect stored answers and their sources. These are inputs to the assessment. Brand presence, positive sentiment or position in an answer does not establish whether every sentence agrees with the documentation. Verify evidence and conditions separately.
If you need a custom metric, define it as: fully supported cited claims / assessed cited claims. Include the first four table categories in the denominator. Separately report claims without a source, unverifiable claims and reasons for exclusion. If the denominator is zero, report no result rather than 0%.
Show counts and question scope beside the percentage. This metric describes the assessed set, not the overall “truthfulness of AI”. Do not combine different modes or markets without explanation. Repetitions do not guarantee independent observations.
A quality audit and a test of a correction’s effect are different tasks. We explain a comparison plan in the before-and-after article update test; do not infer causation from a single improved answer.
Preserve useful content and follow SEO policies
An audit may call for clarifying one condition rather than publishing more similar pages. Google’s spam policies describe scaled content abuse involving pages primarily created to manipulate rankings without user value.
Keep information clear and current, and make structured data agree with visible text. The Google Search AI features documentation states that AI Overviews and AI Mode require neither a special Schema.org type nor separate AI files. This guidance applies to those Google features, not to all models collectively.
- I preserved the question, full answer, AI product, mode and date.
- I split sentences containing several verifiable promises.
- I assigned sources to specific claims or marked uncertainty.
- I read the conditions, product version and scope of the cited materials.
- I assessed factual accuracy and citation support separately.
- I recorded the retrieval date and historical assessment limitations.
- I assigned a correction owner and a recheck method.
- I separated missing citations, inaccessible sources and confirmed errors in the report.
FAQ: auditing AI citation correctness
Does a link to an official page validate the entire answer?
No. Check the specific passage and assigned source. Official material may support only part of the answer or different product conditions.
Does a missing source mean AI provided false information?
No. It means that answer provides no explicit evidence. You can assess the fact separately using the relevant documentation.
Can you assess an old answer using today’s page?
You can compare it with the current state, but that does not establish what the source contained in the past. Record both dates and flag the missing historical version.
Does Semly automatically decide whether every citation is correct?
In this process, Semly supplies stored answers and sources for analysis. Determining whether a resource supports a specific fact and its conditions is a separate task for the auditor.
Does correcting documentation guarantee accurate AI answers?
No. It makes clear information easier to access but does not guarantee source selection or correct interpretation. Recheck and preserve comparable answers.
Share: