Why is an AI agent's budget a wall?
An AI shopping agent given a price limit treats it as a hard wall. On Cromanion's agent bench, three model families placed 231 orders and none was above the limit. They rarely even opened a product priced just above it. A limit worded as soft did bend, by single-digit percentages, in a smaller study on one model family.
Last updated
What did the bench measure about budgets?
Whether agents ever spend above the limit they were given. Across the first four studies, on three model families, 231 orders were placed and none exceeded the limit. 223 recommendations sent back to a person named a product, and none exceeded it either.
The bench is a fictitious headphone store, driven by models from Anthropic, OpenAI and Google, each through its own computer-use interface. Each run carried a mission with a price limit, and usually a target price below it.
Every limit in these studies was stated explicitly. Nothing tested a vague budget, a limit the agent had to infer, or a basket of several items.
Do agents even look above their limit?
Almost never. Products priced just above the limit, at €99, €179 and €349 against limits of €90, €160 and €300, were opened by 0 of 8, 1 of 7 and 0 of 8 of the agents that saw them. Agents walk up the price list and stop.
The highest price an agent opens is so close to its limit that it can be read back from the visit. In one study (71 runs, three families), the limit was guessed right 86% of the time from the first three pages alone, against 34% by chance.
Where inside the limit do agents buy?
Between the target and the limit, not below the target. An agent told to aim at one price but allowed up to another shops in that band. When a visibly better product sat inside it, the agent went up to it.
At a €300 limit, all 7 buying agents took a €279 product, 93% of the limit, over a €241 twin.
| Where the twin was priced | Agents that saw it | Opened it | Bought it |
|---|---|---|---|
| Below the target | 23 | 26% | 9% |
| Between the target and the limit | 24 | 88% | 38% |
| Above the limit | 18 | 0% | 0% |
Does a budget limit ever bend?
Only when it is worded as soft. In a study of 12 runs on one model family, a hard limit was never broken, even when obeying it meant buying nothing. A limit phrased as approximate was broken by 2.5% to 9.9%, to buy at all or to arrive sooner.
The store never sees how the limit was worded. "About €150" and "no more than €150" produced opposite behavior at the same number.
| Situation | Ordered | Went over the limit |
|---|---|---|
| Nothing fits, hard limit | 0 of 3 | 0 of 3 |
| Nothing fits, soft limit | 3 of 3 | 3 of 3, by 9.9% |
| Fast delivery fits, hard limit | 3 of 3 | 0 of 3 |
| Fast delivery fits, soft limit | 3 of 3 | 2 of 3, by 2.5% |
Will an agent pay more for faster delivery?
A little, and it depends on the premium. In a study of 12 runs on one model family, same-day delivery at a 3% premium was chosen 3 times of 3, at 13% once of 3, and at 29% never. When the mission required same-day delivery, the agent paid without hesitating.
Does the agent tell the store its budget?
Rarely, and not reliably. When a form asked for it and the mission said the budget was confidential, 0 of 84 agents answered. When it could be shared, 24 of 55 answers understated it, giving the intended spend rather than the limit. The limit shows in what the agent opens, not in what it says.
Common questions
Is 231 orders with none above the limit true of every AI agent?
No. It is what three model families did on one fictitious store, with limits stated explicitly in the mission. Real agents with vaguer instructions, or shopping for several items at once, were not tested and may behave differently.
Why are some results from one model family only?
The later studies ran through a free subscription driver that uses one model family. They were designed to find mechanisms cheaply. Each result from them says one model family, because it is a statement about that family, not about agents in general.
See it on your own site
Paste one tag. Cromanion crawls your site, watches real sessions in Learn mode, and only acts when you switch it on — with a permanent 10% holdout proving what it caused. Free to start, no credit card: a 14-day or 1,000-session live trial, then keep measuring for free.