AI agents
When the buyer has an agent
Patrick Collison expects personal AI agents to act as a structural subsidy for product quality. They can reward only the quality they can verify, and early experiments find that what an agent buys depends on what it saw first.
- Patrick Collison expects personal agents to act as a structural subsidy for product quality, because they make a better product cheaper to find and to check.
- Cheaper search has moved prices before. Comparison websites helped cut US term life insurance prices by 8 to 15 percent between 1995 and 1997, and sellers in other markets answered by making their prices harder to compare.
- In about 26,000 tests at Wharton, what an agent bought depended on what it saw first. One review screenshot moved a model’s choice by up to 99 percentage points, and one planted sentence of user “memory” redirected most models.
- Whether agents create superstars or spread demand across niches may depend on whether they all read the same sources.
- Airline booking screens and Google’s shopping results both favoured the companies that owned them. An agent’s reward for quality is only as good as the evidence it reads and the interests it serves.
On 9 October 2026, Patrick Collison, the cofounder and chief executive of Stripe, posted a list of things he had been “mulling over regarding agents in the economy.” His first example is the insurance offered at the end of a booking, which he says is “typically priced substantially above what a regular policy would cost.” Companies, he wrote, “are in some sense the original superintelligence.” They can often use their “amortized cognitive surplus” against certain customers, who by comparison can be “inattentive” or “underinformed.”
Personal agents could narrow that gap. In Collison’s words, they will “act as a kind of structural subsidy for product quality.” He asks how markets would look “if individuals always spent at least 10 hours researching their purchases.”
I work on how to evaluate AI agents in science and medicine, where an agent can reach the right answer from evidence that does not support it. Collison’s argument runs into the same problem at the scale of an economy. An agent that spends those 10 hours can reward only the quality it is able to check, and it rewards that quality only if it is working for the person who pays.
Inattention is a business model
Collison invokes Herbert Simon here. In a 1955 paper, Simon replaced the all-knowing chooser of economic theory with one who has limited information and limited capacity to compute, and who stops searching once an option is good enough. The idea became known as bounded rationality, and companies have learned to price it.
When Britain’s Financial Conduct Authority studied insurance sold as an add-on in 2014, it found that add-on buyers were less likely to shop around and less sensitive to price. Their attention was on the main purchase, and many bought cover they did not need or understand.
Subscriptions profit from a related mistake. Stefano DellaVigna and Ulrike Malmendier followed 7,752 members of three US health clubs over three years. Members who paid a flat monthly fee of more than $70 went 4.3 times a month on average, which came to more than $17 a visit, when a 10-visit pass would have cost $10 a visit. On average they forwent about $600 of savings over their membership. Monthly members also kept paying for an average of 2.31 full months after their last visit, which cost them about $187. The authors’ leading explanation is overconfidence, which leads members to overestimate both how often they will go and how likely they are to cancel. Collison’s shorthand is that “people who forget to cancel subscriptions subsidize those who don’t.”
Coupons rely on the same limit from the other side. Chakravarthi Narasimhan modelled coupons in 1984 as a way to offer a lower price only to customers whose time is cheap enough to spend clipping them. Collison asks “how do coupons work if every agent hunts fastidiously for them?” An agent’s time costs almost nothing, so the sorting stops working.
Cheaper search has cut prices before
George Stigler treated the search for a lower price as a cost like any other in a 1961 paper. A buyer keeps checking sellers until the expected saving from one more check falls below its cost, which is one reason the same product sells at different prices in the same market. An agent pushes that cost toward zero.
The internet has already run a version of this experiment. Jeffrey Brown and Austan Goolsbee studied term life insurance in the 1990s, as comparison websites spread. A 10 percent increase in the share of a group using the internet reduced that group’s average prices by as much as 5 percent. Rising internet use did not lower prices for kinds of insurance the sites did not cover, or in the years before the sites came online. Overall, the authors estimated that the rise of the internet from 1995 to 1997 cut term life prices by 8 to 15 percent.
Sellers adapt. Glenn Ellison and Sara Fisher Ellison studied small retailers selling computer parts through Pricewatch, a price search engine popular with savvy buyers, where demand for the cheapest memory modules became “tremendously price-sensitive.” The retailers responded with what the authors call obfuscation, “practices that frustrate consumer search or make it less damaging to firms.” The retailer they studied listed bare processors at very low prices, then tried to sell buyers a processor with a cooling fan attached.
Who pays when every buyer pays attention
Collison admits he is “not sure what the incidence of all of this will be.” Economic theory offers part of an answer. In Xavier Gabaix and David Laibson’s model of shrouded add-ons, firms sell a base product cheaply and earn their margin on add-ons that myopic buyers fail to anticipate. One of their examples is the ink for a printer. Sophisticated buyers take the cheap base product and avoid the add-ons, so, in the authors’ words, “two kinds of exploitation coexist.” Firms exploit myopic buyers, and sophisticated buyers exploit the firms’ pricing. When the add-on has close substitutes, educating customers does not pay, because an educated customer avoids the add-on and keeps buying the cheap base product.
Agents could make every buyer sophisticated. In the model, the add-on price would then fall to what it costs a buyer to avoid the add-on, and the discount on the base product would shrink with it. Buyers would gain on average, because shrouding wastes resources. The gains would go to today’s inattentive buyers, and part of them would come from today’s careful shoppers, who would pay more for the base product.
A subsidy for quality
Collison’s most hopeful idea starts from an old argument in corporate strategy about distribution and product quality. Perhaps you have made something better, he writes, “but does anyone know about it? Won’t people just continue to buy the ACME Corp Mousetraps?” His answer is that agents tilt the argument toward quality. “On the margin,” he writes, “focusing on making a product better is, I think, going to become a more effective strategy for companies.”
Collison’s phrase for what agents do to markets. The subsidy is indirect. An agent lowers the cost of finding a better product and of checking that it is better, so the same improvement wins more customers than it used to. The subsidy is only as large as the part of quality an agent can observe.
That limit is old. In The Market for “Lemons”, George Akerlof showed in 1970 that when buyers cannot tell good used cars from bad ones, they pay only for average quality, and owners of good cars withdraw from the market. Quality that buyers cannot observe earns no premium. Agents widen what buyers can observe, and the rule holds at the new boundary.
Collison draws the corollary himself. If he is right, information about quality “will become more important and impactful,” and “agents will have insatiable thirst for such signals.” Some of that information already exists. Price is visible, and battery life can be checked against a standard test. Evidence about durability arrives years after a purchase and is rarely collected in one place, although an agent could read thousands of repair records in seconds if someone kept them. The subsidy is likely to be largest where someone other than the seller already measures quality.
A signal that steers millions of agents is also worth gaming. Aounon Kumar and Himabindu Lakkaraju showed in 2024 that adding a carefully crafted string of text to a product’s page made a language model more likely to list that product as its top recommendation, in a catalogue of fictitious coffee machines. In ACES, a framework for auditing shopping agents, a seller-side agent making “simple, query-conditional description tweaks” could drive significant gains in market share. If agents read pages that sellers write, the subsidy can flow to whichever seller writes best for machines.
What an agent buys depends on what it saw first
The Wharton Generative AI Labs tested how stable these choices are. In Agentic Shopping is Complicated and Contingent, published on 27 August 2026, a six-person team including Anushka Kumar and Ethan Mollick ran roughly 26,000 tests of six frontier models acting as a shopping assistant. The task was to choose a fitness watch for a user with no stated requirements.
Shown only a grid of eight products, each model settled on a stable favourite. Context moved it. A single screenshot of Wirecutter’s pick sent Gemini 3.5 Flash to that watch in essentially every run, a shift of 99 percentage points. When three sources, Wirecutter among them, arrived bundled in one tool call, GPT-5.5 moved 53 percentage points toward Wirecutter’s pick. When the same sources arrived one at a time, it barely moved.
In another test the team made one product objectively best, a $29.99 watch with a perfect 5.0 rating, against rivals costing at least $359 with worse reviews. Then they planted one sentence of user “memory” in the prompt. With “I love hiking!” added, Claude Opus 4.8 switched from its usual pick to a hiking-friendly Garmin in three-quarters of runs. Only Gemini 3.5 Flash chose the best product at baseline and kept choosing it whatever the team injected.
Researchers at Microsoft found a related weakness in Magentic Marketplace, a simulated market where assistant agents shop for customers and service agents sell for businesses. Every model it tested showed a “severe first-proposal bias,” which gave businesses “10-30x advantages for response speed over quality.”
The hiking result is the hardest to read. A real hiker might well want the Garmin, and then following the sentence is good service. The agent cannot tell a preference its user stated from one that someone planted unless it knows where each line of its memory came from. Following a preference counts as fidelity only when the preference came from the user.
Wirecutter may well be right about the watch. Even so, an agent whose choice flips on one screenshot has reached a good outcome through a process that anyone who controls its inputs can steer. In evaluation terms, the answer was right and the reasoning was fragile.
The Wharton team tested a single product category on simulated pages, for a user with no stated requirements. The tests show how far a choice can move when the context changes. They do not show how often deployed agents buy worse products than their users would have bought alone.
Superstars or niches
Collison leaves one question open. Agents might strengthen superstar effects, if “rational agents all settle on the same product.” Or they might spread demand out, because “everyone has slightly different tastes and preferences, which agents are good at eliciting.”
Studies of recommender systems have found both. Erik Brynjolfsson and two coauthors compared the internet and catalogue channels of one retailer that sold the same products at the same prices in both. Internet sales were significantly less concentrated, and the use of search and recommendation tools went with a larger share of niche products. Daniel Fleder and Kartik Hosanagar modelled and simulated common recommenders such as collaborative filters, which rank products by past sales and ratings. Such recommenders “push each person to new products, but they often push similar users toward the same products.” Each person’s purchases can become more varied while sales across the market become more concentrated.
Early audits of shopping agents point toward concentration. The ACES audits found that agents “can exhibit choice homogeneity, often concentrating demand on a few ‘modal’ products while ignoring others entirely,” and that a model update could “drastically reshuffle market shares.” The Wharton results suggest one reason. Wirecutter’s pick dominated the choices of both Claude models. If millions of agents read the same review site, they could converge on the same product because they share a source. The result would look like a superstar effect while saying more about the source than about the product.
Which effect wins will differ by category. Where one product is best for nearly everyone, convergence is the right outcome. The worry is convergence driven by a shared source when buyers’ needs differ.
Travel agents had this problem first
Collison names the assumption under all of this. “Much of what I’m writing hinges on the assumption that personal agents will be on the consumer’s side.” He suspects that will be the winning strategy, and adds that whether it is true “involves a lot of other questions.”
Airline booking went through a version of this once. Before online booking, travel agencies sold most airline tickets, using reservation systems that airlines had built, such as American’s Sabre. The owner airlines had, in the Department of Transportation’s words, “the incentive and ability” to use the systems to give their own flights “an undue preference,” and the Civil Aeronautics Board adopted rules against display bias in 1984. When the department reviewed the rules in 2002, it explained why such bias could work. “Travel agents tend to book the first flight displayed by a system. Their customers depend on them to extract information from the system display, which consumers do not view themselves.” The department had tightened the rules in 1997, largely over concerns that United had caused Galileo, a system it partly owned, to create displays that prejudiced its competitors. Alaska Airlines put its own lost revenue at about $15 million a year.
An AI agent sits where the travel agent sat, and its user sees even less of the screen. The bundled tool calls in the Wharton tests are a modern version of display order, a lever that only the platform controls and no user would ever see.
Search repeated the pattern. In 2017 the European Commission fined Google €2.42 billion for placing its own comparison-shopping service at or near the top of its search results while demoting rivals, the best-placed of which appeared on average only on page four. The Commission’s evidence showed that the top generic result received about 35 percent of clicks. The EU’s Court of Justice upheld the fine in 2024.
Collison grants that paid placement carries some information. “Willingness to pay is itself a kind of signal, which is valuable to buyers,” he writes. The airline reservation systems did something else. In the department’s words, they “also hid the extent of their bias.”
Economists call this the principal-agent problem. The person who delegates and the agent who acts can want different things, and the agent knows more about what it did. A personal AI agent can also know its user’s habits and budget in detail. That knowledge makes it more useful, and it would make it a more effective salesperson for anyone else who pays it.
Grading the decision
A completed order measures execution. The Wharton tests measured which watch an agent chose and what moved that choice. Grading a buying agent means grading the decision behind the order, which has at least four parts.
- Outcome. The purchase served the user’s needs at the price paid, and buying nothing was one of the options.
- Evidence. Its claims rested on sources that support them, with independent tests weighed above sponsored claims.
- Preference fidelity. It followed preferences the user actually gave.
- Stability. It would survive changes that should not matter, such as the order of sources or the format of a tool call.
Two of the Wharton manipulations test the fourth directly. Shuffling the order of the sources and changing the tool-call format leave the evidence unchanged, so a stable agent’s choice should stay put. Evaluation in science and medicine already separates a right answer from a well-grounded one. Buying agents face the same test, repeated across millions of purchases a day.
The market’s backpropagation
Collison ends with an analogy from machine learning. The market, he writes, is a “highly imperfect value attribution machine,” and agents “will change the nature of the backpropagation such that the rewards for better products increase.” He hopes that companies “that make things that they’re proud of will feel that the universe is a little more partisan in their favor.”
Backpropagation adjusts each weight in a neural network according to its share of the error, so a network learns whatever its loss function measures. A market learns from purchases in a similar way. If agents’ choices track verified quality, money and talent will flow toward better products. If their choices track whatever the agent happened to read first, firms will learn to optimize that instead.
I share Collison’s hope, with one condition attached.
Agents will raise the reward for quality only where quality can be verified independently of the seller.
Further reading
- Patrick Collison, on agents in the economy (9 October 2026), the post this essay responds to.
- Kumar et al., Agentic Shopping is Complicated and Contingent (Wharton Generative AI Labs, 2026), the 26,000 tests of shopping agents.
- Allouah et al., What Is Your AI Agent Buying? (2025), audits of position bias and choice concentration in shopping agents.
- Gabaix and Laibson, Shrouded attributes, consumer myopia, and information suppression in competitive markets (Quarterly Journal of Economics, 2006), on who pays for hidden add-on prices.
- Brown and Goolsbee, Does the Internet make markets more competitive? (Journal of Political Economy, 2002), on comparison websites and life insurance prices.
- George Akerlof, The market for “lemons” (Quarterly Journal of Economics, 1970), on quality that buyers cannot observe.
- US Department of Transportation, Computer reservations system regulations (Federal Register, 2002), on display bias in airline booking systems.