Check your company’s visibility in AI
Free, with no credit card required.
More citations after improving an article are a promising sign, but they do not yet prove the update worked. During the same period, the system's responses, competitors' sources or your measurement process may have changed. To evaluate the update, you need a recorded baseline and clear comparison rules.
The protocol below is a working proposal for marketing teams. It covers observable responses in selected AI systems. It does not describe their ranking mechanisms or assume that monitoring alone establishes causation.
What should you measure when updating one article?
Choose one primary outcome tied to the purpose of the change. If you are improving a guide so that it becomes a source for customer questions more often, a sensible target is the percentage of responses citing that guide's specific URL. If you are assessing product recommendations, a link to an article alone may not answer your business question.
For this protocol, use the following definition: citation rate for the tested URL = number of successfully obtained responses with at least one visible citation of that URL / total number of successfully obtained responses in the tested set × 100%.
A response with three links to the tested resource counts once. Decide in advance how to identify redirects and remove tracking parameters. Record citations of other pages on the same domain separately. This prevents an increase in homepage citations from being counted as a successful guide update.
A valid response without a visible citation of the tested URL scores zero on this measure. A retrieval error, outage or empty result is missing data. If the interface shows a related link whose role is unclear, assign it to a separate category rather than automatically counting it as a citation.
Measure the visible source link. Its absence does not prove that the model is unfamiliar with the content. This test also does not measure overall market share or the actual queries of all users.
We cover a wider range of metrics in our guide to measuring brand visibility in AI. Here, we focus on evaluating a single editorial change.
Write down the protocol before publishing the update
A short plan prepared before measurement reduces the risk of selecting questions and dates to fit the result you want. Decide what you will do if the outcome is inconclusive. Do not change your success criterion after the first favourable reading.
| Protocol field | What to record |
|---|---|
| Hypothesis and change | The effect you expect and the part of the article you are improving, such as adding a table of integration requirements. |
| Unit receiving the change | A specific URL or an entire cluster of related pages; a list of pages receiving and not receiving the update. |
| Questions | A fixed list of natural customer questions, with each question assigned to a resource and an intent. |
| Measurement conditions | System, product or interface, available model version, search mode, language, market, schedule and number of repetitions. |
| Primary outcome | One outcome definition, URL identification rules, and criteria for valid responses and missing data. |
| Comparison and timing | Comparison resource, baseline period, deployment date, transition period and end of observation. |
| Decision | The improvement that matters to the business, limitations that prevent a conclusion, and when to repeat the test. |
Questions should fit the content before the update too. Do not add only questions that the new paragraph answers word for word. A neutral example is “What requirements must be met to connect an online store to a warehouse management system?” The instruction “Answer using our new article” tests instruction following rather than the natural likelihood of selecting a source.
Separate questions containing the brand name from those without it. Keep their wording unchanged in both periods. For manual measurements, start fresh conversations and save the first response before requesting additional sources. Keep API results and consumer app results in separate series.
Establish a baseline and set the schedule
A single response before the change and one afterwards do not reveal natural fluctuations. Take several measurements over time with the same questions and similar frequency. Save full responses, links, measurement times and available information about the mode and model.
An illustrative operational schedule could include:
- two weeks of baseline measurement before the update;
- saving the old version and recording when the new content went live;
- a predefined one-week transition period, reported separately;
- two weeks of observation after the change, followed by assessment on a fixed date.
This is an example work plan, not a waiting period recommended by providers. Choose the periods according to measurement frequency and outcome variability. A transition week does not guarantee that systems have retrieved the new version. Where logs or indexing information are available, use them as supporting evidence of access, not proof that the content was used in a particular response.
Do not stop the test immediately after a favourable reading. If the product, search mode or response collection method changes, mark the break in comparability. Start a new series where necessary rather than combining different conditions into one result.
Report planned measurements, valid responses and missing results separately for each period. If responses to difficult questions are disproportionately missing after the update, the higher citation rate may reflect a change in the data mix. When missingness differs, also compare a common set of questions with similar coverage.
Separate the update from changes in the environment
Add a comparison resource: a similar guide or cluster that you do not update during the test. It should serve a similar intent, have comparable seasonality and be measured under the same conditions. Inspect the trajectory of both series before deployment, not just their averages.
The update is assigned to a page or cluster, so groups should be defined at that level. Two question sets pointing to the same updated page do not create an independent comparison group. A similar problem arises if a shared template change affects both resources or new internal links also affect the comparison page.
In a larger test, you can randomly select some comparable pages for updating and improve the remaining pages after observation ends. Treat a single pair of articles as a pilot. Repeating questions increases the observation count but does not replace having more independent pages.
Keep an event log: a PR campaign, new review, offer change, website rebuild, outage or change in measurement method. Simultaneous growth in both groups may indicate a common change. Differences between groups also need explanation, especially if an event affected only one of them.
Make content changes that can be evaluated
Choose a coherent scope, such as clarifying integration requirements and adding a verifiable example. Keep a copy of both versions. If you change the text, URL, title, linking and promotion together, you are evaluating the whole package. State that in your hypothesis rather than attributing the result to one paragraph.
The update should help the reader by removing outdated information, explaining conditions or providing material that can be checked. Google's guidance for generative features in Google Search emphasises useful, original content and SEO fundamentals. Its spam policies prohibit practices including scaled content creation primarily to manipulate rankings. Do not create separate, near-identical pages for every question variant or hidden instructions forcing a brand recommendation.
Google's rules apply to Google Search. Evaluate ChatGPT, Gemini Apps and Perplexity in their own environments. Our knowledge base covers the content refresh process itself in content recycling for AI and LLMs (in Polish).
Calculate the change: an example with a comparison group
The numbers below are an invented calculation example. They are not from a Semly study or a customer campaign. For simplicity, they cover one system, fixed conditions and equal question coverage. Each period contains 300 valid responses from repeated measurements in each group.
| Group | Before the update | After the update | Change |
|---|---|---|---|
| A: updated resource | 60 / 300 = 20% | 105 / 300 = 35% | +15 percentage points |
| B: comparison resource | 60 / 300 = 20% | 75 / 300 = 25% | +5 percentage points |
Chart: illustrative data only. The percentage of valid responses citing the tested resource's URL rose from 20% to 35% in group A and from 20% to 25% in group B. The difference in changes is 10 percentage points.
Calculate the difference in changes: (35% − 20%) − (25% − 20%) = 10 percentage points. Comparing group A alone would show a rise of 15 percentage points. Including group B reveals that some improvement also occurred for the resource that was not updated.
This comparison draws on the difference-in-differences method. A causal interpretation requires assumptions, including similar outcome trajectories in both groups had the update not happened. The World Bank explains parallel trends and the limits of checking them. In this simplified example, 10 percentage points describes the difference in changes; it is not automatically the effect of the update.
Do not call the result statistically significant based on the table alone. 300 repeated responses do not mean 300 independent users or pages. Measurements of the same questions and resources can be dependent. Assessing uncertainty requires accounting for the data structure and how updates were assigned.
Also examine how the result is distributed. Does the improvement cover many questions or one frequently repeated question? Does it persist on later dates? Report systems separately. A combined result can be misleading if the proportions of responses from different systems have changed.
How can you use Semly for this measurement?
Use Semly to monitor questions, brand presence and sources shown in responses. Response previews let you revisit a specific measurement and check whether a link points to the tested article, another page on your domain or an external resource. The available range of systems depends on your plan and monitoring configuration.
Keep a separate update log for the test: URL, content version, deployment date, assigned questions and comparison resource. Match it to responses and sources from the same periods. The citation-rate definition in this article is a proposal for your own analysis, not the name of a built-in metric or experiment feature in Semly.
Separate three levels in the report: citations of the tested resource, brand mentions or recommendations, and traffic and conversions. A change at one level does not automatically imply a change at the others. Check organic performance in Google Search Console alongside visitor behaviour in website analytics, accounting for the scope and limitations of those data.
How should you decide what to do after the test?
Assess measurement quality first, then the size of the change. Before extending the update to more pages, work through this checklist:
- Questions, schedule and conditions are comparable in both periods.
- Missing data are documented and have not materially changed the composition of the set.
- The result concerns the correct URL, not just any link to the domain.
- The comparison group was not affected by the update or an obvious spillover.
- Known concurrent events and environmental changes are recorded.
- The improvement appears in more than one reading and has been checked by question.
- The assessment uses the business criterion set in advance.
- The report separates observation from causal claims and identifies the next step.
If the result is favourable, persists and shows no major confounding factors, test a similar change on more comparable resources. If both groups grow similarly, look for a common cause. If conditions or data coverage have changed, the result may remain inconclusive. Failure to detect an improvement does not prove the change will never help.
An accurate report for the example data could read: “The citation rate for resource A rose by 15 percentage points and for resource B by 5. The difference in changes was 10 percentage points. The result applies to the tested question set and needs replication across more resources.”
FAQ: testing an article update in AI
Is checking a response the day after the update enough?
That reading can be an initial observation. Evaluation requires a baseline, repeated measurements and comparable conditions. There is no universal deadline by which every system will use a new page version.
What if I do not have a similar article for comparison?
You can observe outcomes before and after the change, but clearly disclose the absence of a comparison group. You are then showing a change over time and accompanying events. Your basis for attributing the result to the update is weaker.
Are 300 responses enough for a reliable test?
There is no universal number. Outcome variability, the number of independent pages, dependencies between measurements and the size of the change you want to detect all matter. The number 300 in the example is used only to demonstrate the calculation.
Do more citations mean more sales?
A citation is an observed source in an AI response. Sales require separate measurement of customer behaviour and conversions. Report the outcomes separately rather than deriving revenue from the citation rate alone.
Is this protocol an official Google instruction?
It is a proposed measurement plan prepared for Semly readers. References to Google concern content quality and search policies. Google does not endorse the schedule, sample size or update evaluation method presented here.
Share:
