1. Recommendations narrow the candidate set first, but do not necessarily reduce the whole search burden
“Making a choice easier” includes at least three distinct things. First comes discovery: finding a few apparently relevant options among thousands of possibilities. Second comes comparison: looking at price, specifications, time, risk and other people’s experience to establish exactly how those options differ. Third comes deciding: setting one option alongside one’s own budget, commitments and consequences, accepting it and giving up other possibilities. Recommendations usually intervene directly in the first stage and can affect the order of the second; the third is difficult to replace with a string of past clicks.
This is easy to see in ordinary situations. To order a familiar takeaway at the weekend, a person may need to choose only from among the restaurants they use regularly; to replace the fridge at home, the first three products that appear are only the beginning. The latter also involves dimensions, energy use, after-sales support, noise, budget and delivery conditions. Even if a recommendation happens to predict a brand preference, it may not know the width of the kitchen door, much less deal with the consequences if the appliance proves too large after installation. Calling both situations “search” obscures the difference in the judgements they require.
Research on candidate lists supports a contingent benefit, not a universal efficiency myth. In a film-recommendation experiment, Martijn C. Willemsen, Mark P. Graus and Bart P. Knijnenburg found that a highly diverse list of five items could deliver satisfaction comparable to lists of ten and twenty while producing less choice difficulty; when participants were required to choose, they viewed about 80% of the items in the five-item list and about 45% in the twenty-item list. This suggests that a small, diverse list may allow people to look more thoroughly at what is before them. It does not show that five is the best number in every context, or that all recommendations save time.[1]
The information on a list likewise determines whether “saving effort” is real. In a randomised experiment at a US online retailer, Xing Fang, SunAh Kim and Pradeep K. Chintagunta held the underlying recommendation algorithm and product match constant across conditions; what changed was the presentation of product information cues. The experiment covered about 260,000 product lists. The results showed that, as information cues were reduced, users made more non-recommendation search clicks and browsed lists more; when people could not assess fit from a recommendation card, search had not disappeared but shifted to keywords or other entry points.[2]
The difference is not simply a matter of how many minutes are spent. If a recommender first offers ten items that are hard to tell apart, people must devote attention to guessing the differences. If it offers a small number of candidates and also states their key attributes, they can rule out unsuitable items more quickly. The former interface looks just as simple, but leaves the cognitive work with the user; only the latter may turn a compressed candidate set into a genuinely useful starting point for comparison. The value of recommendations, then, is not that they replace thought, but that they reduce pointless searching for the material needed to make a judgement.
This also explains why the same person may say that a home page is convenient while opening several pages before buying. Recommendations can remove the discovery cost of starting from zero, yet hide comparison work further along: opening details, checking the returns policy, confirming dimensions and looking for alternatives. If an interface offers only a polished product image or a single title line, the candidate set may be smaller but the material for judgement is inadequate; people must keep looking to make up the shortfall. Conversely, a card with sufficient information may not lead to an immediate purchase, but it can show whether a suggestion deserves to enter the next round of comparison.
Research on choice overload provides a more reliable scale. After aggregating 99 observations involving 7,202 participants, Alexander Chernev, Ulf Böckenholt and Joseph Goodman identified choice-set complexity, decision-task difficulty, preference uncertainty and decision goals as important moderators of choice overload; item count does not determine choice overload on its own.[3] Ten very similar, poorly specified options may therefore be harder than twenty with clear boundaries; someone who already knows that silence and the warranty period are the only things that matter may also find a long list easier to handle than someone whose goals have not yet formed.
Recommendations may therefore ease the burden of where to start looking, rather than all search. For low-cost, repeated choices, they can shorten the route to a useful starting point; once an item requires explanation, comparison or the acceptance of irreversible consequences, users should still separate “what is in the list?” from “what else do I need to know?” The latter question has no ready-made ranking answer.
2. Ranking turns behavioural traces into the priorities before us
Recommendations can easily seem like judgement because they do not merely provide a catalogue: they determine its order. The same thousand songs or several hundred shops may still be in the catalogue, but the items on the first screen are given the chance to be seen, clicked and compared first. There is no need to treat this order as a conspiracy; any finite screen needs ranking. What matters is seeing that ranking relies on traces the system can read, rather than all of a person’s reasons.
Take video services. YouTube’s help page publicly lists watch history, search history, subscriptions, likes, dislikes and “Not interested” among the signals that affect its home page and recommendations.[4] TikTok’s explanation identifies user interactions, video information, device and account settings as the main factors affecting “For You”, and says that different signals do not carry the same weight.[5] These materials are useful for showing the types of input involved; they are public accounts by the platforms, not independent audits, and cannot reveal the full weight behind each ranking.
Behavioural traces have one advantage: they can be processed quickly. Listening to one kind of music repeatedly, searching again and again within one price range, or moving back and forth between several routes all leave computationally usable clues for a system. For someone who simply wants to find similar material quickly, such proxies are often useful. But they also have an inherent gap. A person may play cartoons continuously to accompany a child, search for a certain device for work, or look at the same city only because of an unexpected trip; such actions show what happened, but not necessarily what they will want in future.
A Netflix engineering paper once described a home page that was not decided in a single pass by one all-purpose algorithm, but organised multiple recommendation tasks and algorithms around different page positions and usage situations.[6] It is a company report about its 2015 system and cannot establish that every platform works in the same way today. Still, it points to an easily overlooked fact: “recommended for me” often contains several local rankings. Continue-watching material, similar titles, popular content and entry points for new releases do not necessarily answer the same question.
This is also why a home page can sometimes feel internally contradictory. One row may infer similarity from something just watched, another may preserve an entry point for new content, and a third may seek to maintain familiarity. They need not conflict with one another; they serve different interface tasks. If users treat every row as the same diagnosis of their stable preferences, they mistake a page assembled from contexts for a single, complete portrait of the self.
The sense that a home page fits needs to be translated into a plainer sentence: using the material it can see, the system has assigned priorities for this moment. That is far weaker than “the system understands me”, but closer to the real relationship of use. A person’s goals also include conditions that have not been captured in the data: whether they are tired today, willing to spend more, inclined to try something unfamiliar, required to explain a decision to family, or would rather have fewer choices than miss a crucial risk. Those conditions are not evidence that recommendations have failed; they simply remind us not to mistake recorded proxies for preference for a complete ranking of values.
This distinction also changes how accuracy should be understood. A recommendation can predict the next click with great accuracy while still not answering, “Is this what I most need to do now?” It may even get a short-term taste right while missing a long-term goal. Faced with such a list, a more useful question is not to demand an abstract “truth” from a platform, but to ask: which visible signals do the items placed first chiefly reflect? Have the conditions I care about most right now been included in the comparison?
3. Recommendations can broaden or narrow the range of exploration
Once it is clear that ranking affects what is seen first, it is easy to slide into another overstatement: repeated use of recommendations will lock people into an ever-narrower world. The claim has intuitive appeal, but goes further than the available evidence. The diversity of recommendation lists, the diversity of actual consumption, whether a group becomes more similar, and whether an individual really follows a suggestion are all different questions. Folding them into the single phrase “recommendations create a filter bubble” leaves the discussion without a testable claim.
When T. T. Nguyen, Pik-Mai Hui, F. Maxwell Harper, Loren Terveen and Joseph A. Konstan used MovieLens data to study the relationship between recommendations and content diversity, they found a result worth taking seriously, but only on its own terms: across all users’ top fifteen recommended films, mean pairwise distance fell from 25.02 to 24.67. The change was statistically significant, but small in magnitude. This number describes a specific system, a specific algorithm and a specific recommendation-list measure; it is not a measurement of one user’s entire viewing activity.[7]
The same study did not equate “lists becoming more similar” with “people consuming more similarly”. The authors also compared films users had rated, consumption-related measures for different groups, and user experience; each measure answers a different question.[7] More importantly, the researchers explicitly acknowledged that they could not determine whether users actually followed MovieLens recommendations, and did not include discovery sources such as friends, editors, search engines or other services.[7] This is not a methodological detail that can be passed over lightly: if people move between multiple entry points, the log of one service cannot stand in for their whole environment of exploration.
Other empirical evidence points in a different direction as well. In their paper’s abstract, Kartika Hosanagar, Daniel Fleder, Dokyun Lee and Andreas Buja reported no consumer fragmentation; the abstract describes personalisation as both broadening consumer interests and making conditional product combinations more similar.[8] “Range of interests” and “which bundles were bought under given conditions” are not the same measure, so the two changes can occur at once. Because the verified material here is a publisher’s abstract, this article can only report that conclusion faithfully; it cannot add unreported sample structure or individual causal mechanisms.
Put differently, diversity is not a scale with only a high and a low end. A list can be more spread across genres yet still direct a user towards a few high-frequency brands; someone may encounter works they had never searched for because of recommendations, yet ultimately choose the most familiar kind of content. Researchers therefore need to say whether they are measuring recommended items, actual choices, category distance or overlap between groups. Everyday users, too, need not rush to apply a popular label to their experience; they can separate “what else did I see this time?” from “what did I finally choose?”
These studies do not eliminate the concern; they return it to the right scale. If someone starts every time from the same home page and the same sequence of historical signals, certain categories are more likely to recur: that is a visibility structure produced by ranking. But it does not directly show that their world has narrowed, still less that they have lost autonomous judgement. Some people will continue along a recommendation, while others use the home page as a warm-up before turning to friends’ lists, offline experience, specialist reporting or active search. Exploration has more than the two extremes of being pushed by a system and searching entirely by hand.
The genuinely useful reminder is that the first screen is always only the first screen. It is well suited to relieving the uncertainty of starting from zero, but should not automatically count as evidence that all important alternatives have been seen. For choosing a film for one evening, treating the first screen as sufficient carries little cost; for changing career, enrolling on a course, choosing insurance, making a major purchase or dealing with a health matter, what may be missed is not simply another product, but another standard of evaluation. In such situations, opening one more entry point that does not rely on the same behavioural history is not diagnosing a “filter bubble”; it is acknowledging that the home page may have omitted something unknown.
4. Autonomy means retaining avenues to scrutinise, reject and reshape
The easiest mistake in discussing autonomy is to imagine it as “having no external help at all”. That places the convenience of recommendations and human judgement in opposition. The standard used here is narrower: when using recommendations, does a person still have the opportunity to understand broadly why a suggestion appeared, see or look for meaningful alternatives, reject it, and use feedback to affect subsequent ranking? It is a position from which to exercise judgement, not a psychological state automatically delivered merely because a feature exists.
That position matters not because everyone needs to understand a model, but because people at least need to know where they can begin to intervene. Thao Ngo, Johannes Kunkel and Jürgen Ziegler interviewed ten frequent Netflix users and found substantial uncertainty and confusion about recommender systems among participants; the paper also discusses users’ lack of clarity about controllability and interaction points.[9] This is a small, single-platform qualitative study, so it cannot estimate everyone’s level of understanding. It does at least show that a row of recommendations in an interface does not mean users naturally know how the row was formed or how it can be changed.
To make “transparency” concrete, one can first ask not whether a system has made all its code public, but what information users lack when they make a decision. In a simulated setting involving high-stakes clinical recommendations, Eric S. Vorm and Andrew D. Miller asked twenty-two participants to rank questions they considered valuable, including “Why is this recommendation the best option?”, “Which factors did the system consider, and how were they weighted?”, “What information about me does the system know?”, “Are there other options?”, and “Can feedback I give affect the system?”[10] These questions come closer to actual judgement than the general request to “explain the algorithm”: people need to know what reasons they are accepting, and which boundaries those reasons do not cover.
The limits of this study must, however, remain in view. It took place in a clinical simulation and a high-risk context, with a very small sample; it identifies information that users may want to ask about, rather than proving that displaying it would make users more autonomous or trusting, or lead them to make more accurate decisions. Inflating the finding into “transparency has resolved the problem of control” would repeat the substitution this article has sought to avoid: using an apparently complete label in place of a real process of judgement.
The controls supplied by actual products also come in layers. TikTok’s public explanation says users can use feedback such as “Not interested” to affect subsequent recommendations.[5] This proves only that the platform describes a feedback entry point; it does not prove that the point is prominent enough, that its effects are predictable, or that it covers every pattern a user no longer wishes to see. But it suggests that “rejecting” and “reshaping” are not the same thing: dismissing one item says no to the current item; being able to understand and adjust the next ranking comes closer to setting conditions on the relationship itself.
In consumer interfaces, controls are often deliberately lightweight: an X, a “Not interested” label, an unfollow action. Lightweight controls are not inherently bad; the problem is that, if users do not know what an action will change, they cannot easily judge whether it is enough to correct the system’s misreading of them. For important choices, being able to return to search, rewrite the conditions or temporarily do without personalised ranking is also part of control, rather than a remedial procedure activated only after failure.
Autonomy therefore does not require people to bypass recommendations every time, nor does it require platforms to compress all complexity into one explanation. It asks for a more limited, more realistic ability: when a suggestion matters, people can ask why it is ranked first; when it does not fit, they can find another route; when a system persistently misreads them, they need not rely only on repeated clicks to correct it. Whether those conditions exist is the normative yardstick used here to judge whether recommendations remain a tool rather than a substitute for a person’s decision, not an empirical claim about the effects of any one product.
5. Recommendations can simplify the start of a choice, but cannot make the final judgement
Manual search has costs too. Every additional page opened and every parameter compared consumes attention; when preferences are still unclear and a task is complex, adding candidates does not necessarily make a decision easier. The choice-overload meta-analysis makes precisely this point: candidate-set size must be understood alongside complexity, task difficulty, preference uncertainty and decision goals.[3] The reasonable conclusion is therefore not “never use recommendations again”, but not to mistake saving the first step for outsourcing the whole task.
A more practical approach is to allocate the judgement one needs according to cost and reversibility. For low-cost choices that can easily be reversed — what to listen to tonight, which coffee to drink on the commute — recommendations can set the starting point, followed by a very brief check of one condition that matters to the user. In that situation, spending half an hour looking for the theoretically optimal option may not be freer than accepting a sufficiently good suggestion; freedom also includes keeping attention for more important things.
For medium-cost choices, writing down two or three personal criteria is more useful than endlessly refreshing the home page. When buying a device to use for several years or enrolling on a course that will require weeks of commitment, one can first establish a price ceiling, compatibility, timetable or after-sales requirements, then compare the recommendation list with at least one independent entry point. The entry point need not be mysterious: it may be official specifications, independent reviews, a friend’s real experience, or a different way of searching that does not reuse the same historical signals. The point is not to manufacture information, but to fill in the questions most likely to be absent from the current ranking.
For high-cost choices that are hard to reverse, recommendations are better suited to the role of “give me a few categories and keywords” than “reach a conclusion for me”. Career changes, long-term medical plans, insurance and major spending often involve personal goals, risk tolerance and other people’s interests, which any single trajectory of clicks is unlikely to cover. Expanding the comparison, verifying sources and consulting professionals do not reject convenience; they reclaim the parts that recommendations have no standing to decide for a person.
Layering also means allowing oneself to stop. Some people, afraid of missing out, scroll through a recommendation list to the very end, believing that seeing a few more options makes them more responsible; others, to avoid weighing trade-offs, accept the first answer that looks appealing too early. The former add information that may not be relevant, while the latter give up necessary comparison. A better stopping point is not a fixed number, but the point at which one’s key conditions have been checked, the main alternatives have been seen, and the benefit of continued search is known to be small.
This layering is not a universally applicable process proven by experiment, but a personal practice proposed on the basis of the evidence in this article. Small, diverse candidate sets could reduce choice difficulty in a specific film-recommendation experiment, showing that people need not believe that more is always better.[1] The online retail experiment also showed that, when candidate cards lacked the information needed for judgement, users continued searching outside recommendations; seeing less saves effort only when the information is adequate.[2] Taken together, these points support bounded use, not a moral allegiance to either recommendations or manual search.
After handing a decision to recommendations, what truly becomes easier is where to start looking and which items to compare first; what does not become easier is working out what one is willing to pay, wait for, give up or risk. Good recommendations can arrange the books on the table more sensibly, but cannot decide which one is relevant to a reader’s life. Leaving room for one check, one rejection and one opportunity to reshape the next round of recommendations does not assume that recommendations have already drawn the boundaries of judgement; it tests, when needed, whether they have omitted conditions that still require comparison.