First, define the title's entry action
If "seeing a citation", "returning to a brand later" and "completing a subscription" are all treated as a visit to the original website, the discussion loses its denominator at the outset. The title's entry action is an identifiable source-link or traditional-result opening in the same results interface after an AI answer or summary has appeared. It is an entry event, not an overall judgement of a website's value and not a synonym for a commercial outcome.
In the reproducible Pew definition, a next URL after the results page counted as a match when it matched a first-page traditional-result link or one of the first three source URLs in the AI summary. When the same URL appeared in both places, it was assigned to the summary source. [1] This definition is enough to answer whether an identifiable, result-adjacent opening occurred. It cannot cover a user typing an address, a cross-device visit, a return after a long delay, or a later page with no source marker. Those cases should remain unknown or non-matching; they cannot be inferred as an opening from the AI answer merely because they occurred in roughly the same session.
To turn the title's question into a repeatable measure, a publisher should create a unique exposure_id for each exposure and record the entry type, target URL and timestamp. answer_exposed, source_link_opened, traditional_result_opened, answer_session_end, verification_anchor_viewed and task_completed should be independent events. A 60-second post-exposure window can be registered in advance as the primary window, with five minutes as a sensitivity window; the report can also state whether an opening occurred before the end of the same search session. These windows and event names are this article's measurement design, not a standard set by Pew, and changing a time window cannot create a causal interpretation.
The report should also leave a clear place for "unknown". An opening without a source marker is not zero clicks, and a delayed opening is not automatically a direct click; they are simply cases the current attribution rule cannot determine. If a dashboard silently puts unknowns in the not-opened bucket, the team will understate the effect of cross-device and later decisions. If it puts them all in the opened bucket, it will count brand habits and existing relationships as answer-generated entries. Acknowledging the measurement gap first is what makes it possible to decide which log or question is needed in the next round.
This separation may look finicky, but it determines whether the answer can genuinely address the title. A source being displayed is not a source being opened; a source being opened is not the original text being verified; and completing an on-site action does not mean the user arrived because of the answer. Direct brand visits, email, notifications, bookmarks, returns from an existing subscription and later browsing within the site may have commercial value, but without an entry marker linked to this exposure they cannot be rewritten as the "visit to the original website" in the title. Registration, payment and long-term retention are separate outcome variables and cannot fill a missing click attribution.
1. What can currently be confirmed is a behavioural boundary in one Google sample
The closest field observation to the title's question comes from Pew. The study tracked browser activity for 900 US adults in the KnowledgePanel Digital panel from 1 to 31 March 2025; the research team captured results from 7 to 17 April 2025 to reconstruct the search pages. Of 68,879 independent Google searches, 12,593 included an AI summary. Under the next-URL rule, the match rate for traditional-result links was 8% when a summary appeared and 15% when it did not; the match rate for the first three summary sources was 1%; and the session-end rate was 26% with a summary versus 16% without one. [1] [1] [1]
These figures provide the minimum answer needed by the title: in this time period, country, browser sample and counting rule, searches with a summary still recorded matched openings; the original-site entry did not fall to zero. At the same time, the observed proportions are not the answer to "how much did AI make users click less?" Summary and no-summary searches were not randomly assigned treatments. The query, the user's task, the device and the page presentation could all affect both whether a summary appeared and what the user did next. Pew's data also did not measure query risk, task intent, the reason for an opening, reading or verification after opening, transactions or long-term retention. Thus, 8% versus 15% can serve as an associated behavioural boundary, but it cannot be written as a general causal effect.
Nor can a stable group of clickers be extracted from the overall proportions. 8% is not the click rate for "high-risk users", 15% is not the rate for "low-risk users", the 1% match to summary sources is not the share of people who "trust citations only", and the 26% versus 16% session-end figures are not a portrait of "people satisfied by the answer". All of them depend on the same observation rule and sample boundaries. Recasting those numbers as user motives would add conclusions about variables that were never measured.
That is also why the importance of the original website cannot be judged from total search traffic alone. Total volume stacks together different entry routes, tasks and time horizons. It cannot tell an editor whether a particular answer exposure produced a direct source opening, or whether an opening led to verification, an action or a relationship outcome. If an overall traffic decline is treated as one crisis, the team will change headlines, summaries, internal links and commercial components at the same time without knowing where the gap lies, and still will not know which change affected the user's task.
Pew also observed that the presence of a summary and the session-end rate changed together, but these are parallel results in the same observational sample. A session ending may mean that the answer was sufficient; it may also mean that the user paused, switched channels or never intended to continue the task. Without a post-exposure task chain, "ended" cannot be interpreted as "satisfied", and it cannot serve as negative evidence for a particular kind of clicker. The most reliable wording stops here: under the next-URL rule, a specific sample recorded some entry events, while "who will stay" has not been identified.
2. Abstract-level studies can flag the question, but cannot name the clicker
Beyond that behavioural boundary, the available material is closer to a research map than a user list. In the abstract record for CHI Extended Abstracts 2025, Kaiser et al. compared real search tasks by 1,526 US participants using generative AI and Google; the abstract reports that generative-AI users reached correct answers faster and relied less on primary web sources. [2] This can suggest that an answer interface changes reliance on web pages, but it does not provide sequential source-opening events, an opening rate, task strata, risk variables, post-opening verification or an effect size. Turning "relied less on primary web sources" into "users did not click", or into a claim that a particular task caused a click difference, goes beyond what the abstract supports.
The title and OpenAlex abstract of the paper by Degachi, Niforatos and Kortuem list agent persona, source-clicking, reliance, trust and the moderating role of health literacy as research subjects. [3] This shows that source-clicking can be studied directly. The readable material in this round, however, contains no participant count, health task, experimental condition, click coding, direction or effect size. It cannot answer who is more likely to click in a health task, or who returns to the original website in general search. The existence of a research topic does not make the result needed here available.
Together, these two leads remind editors that a plausible explanation remains an explanation. Someone may be looking for a particular fact in the answer; someone else may need to compare options, take an action or check an ongoing state. The risk and cost of error may differ too. Until there are stratified source-opening records, however, these differences can only be pre-registered variables. The coding can allow categories to be added or removed, multiple task labels to be applied, and "cannot be classified" to be retained. It should not be packaged as a set of mutually exclusive, naturally occurring and exhaustive user types.
Research design must also acknowledge confounding. The task, interface, prominence of the source, a user's existing trust in a brand and the device selected for the sample may all affect an opening together. Even if a future study finds an association between one variable and source opening, that cannot immediately be upgraded to "this kind of person is a durable clicker". The denominator must be specified; so must the exposures with a source marker, the presence of a comparable no-summary condition or different link position, and what happened after the opening. If the opening rate does not change, the direction is opposite to the hypothesis, or an agent has already completed the task, the result should remain unknown rather than being forced into the original category.
For now, the "who" in the title can be answered only at the level of a behavioural event: in a specific and limited set of exposures, some next URLs matched a result or a source. To move towards "who, why and when", the record would at least need to include the user's task and the page conditions, while keeping openings, anchor views, authorised actions, task completion and returns separate. Without those fields, any polished audience portrait may simply be product intuition given a new name.
3. Evidence, action and relationship are page responsibilities to test, not click laws
Publishers still need to decide which pages are worth continuing to build, but the unit of decision should be page responsibility rather than a presumed clicker. The judgement can be written as a conditional chain: has the answer already closed the task; if not, is the gap verifiable evidence, live state and an authorised action, or an account, consent and responsibility handover? Each gap should have a page asset and an independent event measure, together with an alternative route that would make the hypothesis fail. The chain is an analytical and experimental framework, not an industry rule proved by this round's material.
The answer is already complete. If the user wants only an acceptable one-sentence answer and needs no further evidence, state, permission or responsibility, continuing to pursue an opening may simply repackage the answer as a longer page. The team can test answer coverage and session end, but cannot interpret session end as satisfaction or add a task-irrelevant barrier to manufacture a click. If shorter evidence with a precise location makes users leave sooner, that departure may still represent task completion; page value should be judged by the task result, not by whether the click was prolonged.
Missing verifiable evidence. When an answer gives no locatable original text, source identity or version, a page can try to provide attributed excerpts, version history and verification anchors. The minimum condition is not that the content is "professional", but that the user genuinely needs to judge whether the answer can be traced. Events should separately record verification_anchor_viewed, source opening and task_completed after verification. If the user does not view the anchor, or the summary already supplies enough of the original excerpt and its date, the evidence asset may not result in an entry; an opening may not mean that verification occurred.
Missing live state or an authorised action. If an answer cannot obtain a current state, or cannot represent the user in logging in, confirming stock, submitting a form or making a payment, a page can provide live data, interactive tools and a clear handover of permissions. The corresponding measures should be task completion, the start of an authorised action and its completion, rather than only a landing page view. If an agent already has API, login or payment authority and can perform the task reliably, the user may not need to open the page at all. That is not evidence that the website has "lost the user"; responsibility for the task may have moved to another execution entry point. An opening is a path to test only when the site must carry the final confirmation and the agent cannot obtain or reliably express that state.
Missing an account or responsibility relationship. Some tasks may require an identity, consent, payment, an audit record, human support or confirmation of responsibility. A page can provide an account flow, confirmation records and support channels while recording authentication steps, completed confirmation and support contact separately. Yet visiting a page does not mean that a relationship has been established. If the user does not complete confirmation, nobody takes responsibility, or the user switches to a phone call, email or another channel, the original website cannot present the visit as a relationship outcome. The easiest mistake is to treat "someone opened it" as "someone trusted it", then treat trust as revenue or retention and skip the responsibility chain in between.
The four gaps may appear together or not at all. One page may have both version evidence and a live tool; on another, the answer may already be enough and every extra component may only add cost. Task, risk, permission, trust and source presentation should support multiple labels and unknown values rather than a fixed portrait. Reports should pre-register denominators, proportions, confidence intervals and interaction terms. If a result does not change the opening rate, runs in the opposite direction, or an agent completes the task directly, the conclusion should be an unconfirmed page responsibility, not an announcement that a class of "durable clickers" has been found.
4. Counterexamples should first break the most pleasing explanations
Treating evidence, action and relationship as potential entry points is still not enough. Every explanation needs a counterexample stress test, or publishers will turn a guess into a product roadmap.
The first counterexample is "the summary is already sufficient". If a summary provides the original excerpt, source identity, date or a locatable reference at the same time, even users who care about evidence may decide not to open it. The page then has a stronger evidence asset but may receive fewer direct entries. Without measuring how often the excerpt is sufficient, it is impossible to say that evidence always preserves a click; a fall in openings cannot be used to assert that users no longer need factual verification.
The second counterexample is "high risk but no verification". A user may trust the answer, underestimate the cost of an error, turn to a different product or service, or open the original without reading the key passage. An opening by itself does not prove an intention to verify, and viewing a verification anchor does not automatically prove task completion. Risk perception, the verification action and remediation after an error must be coded separately; otherwise "high-risk users will come back" is just an untested reassurance narrative.
The third counterexample is an authorised agent acting directly. If an agent has a usable interface, a logged-in state, payment confirmation or permission to access live stock data, it may compare, fill in a form or transact directly. A page may be the next entry point only when it carries final confirmation or human judgement, or when the agent cannot reliably obtain the state. Even when an entry occurs, agent action, site confirmation and the user's final consent must be distinguished. Counting an agent call as a web click, or a web click as a transaction, would enlarge the evidence beyond what was measured.
The fourth counterexample is an unattributable return. Direct brand visits, email, notifications, bookmarks and returns from an existing subscription can all produce a page visit without showing whether a particular AI answer triggered it. Pew's browsing context did not establish a matching chain from AI exposure to long-term retention. Site-wide returns therefore cannot stand in for a missing exposure_id, and a later registration or subscription by the same user cannot be attributed backwards to that summary.
The point of counterexamples is not to deny a website's value, but to make a value proposition falsifiable. Each experiment should state in advance what result would invalidate the explanation: a more complete summary with no change in openings; people viewing verification anchors without completing the task; agent completion rising while web visits fall; or returns appearing only when no source marker exists. When results land in one of these cases, the page-responsibility hypothesis should be updated rather than searching for a label that preserves the idea of a clicker.
5. Separate control from the validation path
Before testing "who will still visit", a publisher must answer whether it can change the answer interface. Different control means different identifiability. Mixing the two paths turns an observational association into a platform experiment.
Path A: an organisation operates its own answer interface, or works with a platform while actually controlling the experiment. Only under this condition can the same task or query be randomised across summary excerpts, source visibility and link position, with a no-summary control. Each treatment should use exposure_id to connect answer exposure, source-link opening, traditional-result opening, session end, verification-anchor view and task completion. The primary window can be registered in advance as 60 seconds, with five minutes as a sensitivity window, and the report can add events before the same search session ends. Count only the first opening; verification-anchor views, authorised actions and task completion must not be folded into click.
Even if Path A is randomised successfully, the estimate is limited to the controlled interface and participating sample. It can compare whether a form of source presentation changes entry or task completion, but it cannot turn a treatment difference in one interface into a clicker category that applies to every website. Change the answer content, recruitment or task mix and the result needs to be interpreted again.
Path B: an ordinary publisher has no control over an external AI interface. An ordinary publisher cannot change the summary excerpt, source visibility or link position in Google or a chat product, and cannot create a no-summary control. The work should be called observational analysis, not a platform randomised trial. A publisher can pre-register the use of naturally occurring variation in summaries, source display or link position where it can lawfully obtain logs with source markers, then match landing-page events, verification_anchor_viewed, authorised actions and task_completed to task coding.
The limitations of Path B must appear in the results, not be hidden in a footnote. Without an exposure_id, source marker or comparable naturally occurring variation, the result should be reported as unattributable. Site-wide traffic, search ranking or one isolated return cannot replace this chain. Even with matchable events, query, interface-selection and sample-selection confounding may remain, so the result can only be reported as a pre-defined observational association. For an ordinary publisher, the most valuable output may be knowing which fields are missing rather than rushing to produce a click rate.
This requires editorial, analytics and product teams to maintain an event dictionary together before discussing a redesign. The dictionary should specify each event's trigger, deduplication rule, permitted unknown values, attribution window and exit route. Reports should show exposure count, matchable-exposure count, opening rate, verification-anchor view, authorised-action start, task completion and session end together. The purpose is not to make every behaviour trackable, but to stop an easily obtained metric from standing in for an outcome that has not been measured. If a source link has no stable identifier, the report should retain "cannot be determined" and explain which visits are background only and cannot enter the post-answer entry denominator.
Page assets should undergo the same minimum-condition check in editorial decisions. An evidence module should state which locator or version gap it fixes; a live tool should state which kind of state it can obtain; an account flow should state who takes responsibility for confirmation and when. If none of those conditions exists, adding components only expands the maintenance surface and does not automatically create a reason to visit. Conversely, if a lightweight page lets a user complete verification or a handover, a low direct-opening rate alone cannot declare it a failure. The comparison should be between pre-specified task outcomes and responsibility events, with the click as only one entry metric.
The same records should retain why an exposure was not matched, rather than only a total of unopened cases. No source marker, outside the window, cross-device, direct address entry and missing event can be listed as separate unknown states. These labels describe data quality, not user types. The analysis should report each denominator first and then decide which states enter a sensitivity analysis. If one version merely adds a trackable link without increasing verification or task completion, a higher attributable rate cannot be treated as higher page value. If fuller event records instead show that most tasks end at the answer layer, the team should accept "no necessary entry" as the result.
For an editor, the minimum deliverable is not a list predicting who will click, but a decision table that the next round of data can overturn. It should state the answer gap, the responsibility carried by the page, the corresponding events, the primary window, the alternative route and the treatment of unknown values. Content updates, tool development and relationship flows can then be compared against the same task outcome. Even if openings do not increase, the team can tell whether the answer was already complete or the attribution chain is still broken. Calling both situations "traffic is doomed" would only make the next experiment answer the wrong question again.
Technical documentation can help mark this measurement boundary, and no more. Google Search Central's AI Features documentation describes the AI Overviews and AI Mode interfaces and their supporting links. [4] Google's robots documentation says that the nosnippet rule concerns control of a summary or preview. [5] OpenAI's crawler documentation distinguishes three technical roles: OAI-SearchBot, GPTBot and ChatGPT-User. [6] These materials contain no user sample, treatment assignment, opening-event definition or effect size. They cannot prove clicks, citations, conversions, revenue, licensing rights, copyright conclusions or competition-law conclusions, and they should not be treated as independent evidence of user behaviour. What they can do is help a team distinguish search display, summary control, crawling and user-triggered access, so that crawler logs are not misreported as reader entries.
Conclusion: there is no reliable "who" yet, only limited entry evidence and testable responsibility hypotheses
As of the material found and verified in this round, there is no reliable list of who will still visit the original website. What can be confirmed is that, in Pew's specific US Google browser sample, an opening was observed under the rule of whether the next URL after the results page matched. When a summary appeared, the traditional-result match rate differed from the no-summary case, but this is an observational association: it identifies neither the user nor a causal effect of the summary. Kaiser's material provides only an abstract-level clue about generative search, while Degachi's material shows only that source-clicking is a research variable. Neither is enough to provide a clicker, a motive or a conditional effect size.
That does not mean that the original website can only wait for traffic to fall. Publishers can start by asking whether the answer closes the task, then check whether the page still carries source verification, live state and permission actions, or an account, consent and responsibility relationship. These remain analytical hypotheses that can be added, removed, multi-labelled or disproved; they are not established durable-clicker categories. Each page responsibility should have independent events, a clear denominator and an alternative route. If an agent completes the task directly, a summary is sufficient, a user does not verify, or a return cannot be attributed, the result should remain unknown.
The direct answer to the title is therefore not "a particular kind of person will stay", but "there is not yet enough evidence to name one". The next question is who controls the answer interface: a controller can run a bounded comparison, while a publisher without that control can only pre-register natural variation and matched observational records with source markers. Only by separating clicks, verification, action, responsibility confirmation and long-term relationships can a publisher learn when the original website remains a necessary entry point. Until then, calling any citation, return or conversion a click will make the answer sound more certain than the evidence.