Agent Visit Optimization

How do you measure agent traffic with a holdout?

A holdout is a randomly chosen share of visits where a system deliberately does nothing. For AI agents, it is the only fair comparison: agent visits where the system acted, against agent visits where it held back at the same moment. Before-and-after comparisons fail for agents in ways they do not for people, because agents change and remember.

Last updated

Why not compare agent visits before and after a change?

Because too much else changes at the same time. Agents are updated between releases. Some keep memory across conversations, so a second run is not a fresh visitor. And agent traffic is still small. A before-and-after gap can come from any of these.

Cromanion's own lab shows the trap. On 2026-09-25, four agent runs on its test store looked better after a change than four runs before it. Those runs are anecdotal. There was no concurrent control, and the agents keep memory across conversations. It is exactly the comparison a holdout exists to replace.

How are agent visits assigned to the control group?

At random, by visitor rather than by visit. A returning visitor keeps its group every time. Agent visits are assigned by the same rule as people's, so the control group holds the same mix of agents as the acting group, on average.

Assigning by visit would be wrong. A visitor who comes back five times would count five times, and a few frequent visitors could make the two groups look different when they are not.

Cromanion's holdout is permanent. It is never switched off to recover the traffic, because the comparison is only as good as its control group is continuous.

What exactly is compared?

Agent visits that reached a moment where the system could act. In the acting group, the system acted. In the control group, it would have acted, and held back. Each group's result is how many of its visits reached checkout.

A visit that never reaches such a moment is in neither group. Counting it would dilute both sides with visits the system never had a chance to change.

The acting group counts what the system decided, not only what reached a screen. The control group can only be "would have acted", so this is the fair match. It is called intent to treat.

Recognize the agent

By a verified signature, a verified bot category, or behavior.

Reach a decision moment

The visit gets to a point where the system could act.

Act, or hold back

The visitor's group decides. The control group is left unchanged.

Count the outcome

Whether the visit reached checkout, per group, over the last 24 hours.

How many agent visits are enough?

Enough that a gap is not noise, and that takes a while for agent traffic. Cromanion shows raw counts from the first visit, and a rate only once each group has 30 agent visits. Even then it shows the two rates side by side, never a lift. A handful of visits produces large gaps that mean nothing.

What can make the comparison wrong?

Anything that sorts agents into a group by something other than chance, or that miscounts them. Late recognition, split visits and unrecognized agents are the three seen so far. Each is a limit to state, not a result to correct silently.

Late recognition

An agent recognized by behavior is recognized after a few seconds. Its first pages are counted before the system knows it is an agent. On 2026-09-25, Muse's first catalog pages were seen before it was recognized.

Split visits

One Muse visit sometimes landed as two sessions. A split visit can count twice, once without the checkout.

Unrecognized agents

An agent that signs nothing and moves like a person is counted as a person, in both groups. The comparison covers recognized agents only.

Common questions

Why keep a control group for agents that would have bought anyway?

Because without it there is no way to tell which ones would have. An agent that reaches checkout after the system acted may have reached it without any help. Only the held-back group shows how often that happens.

Is the agent holdout the same as the one for people?

It is the same assignment, by visitor. The comparisons are read separately, because agents and people behave differently. Mixing them would let one population's behavior move the other's number.

See it on your own site

Paste one tag. Cromanion crawls your site, watches real sessions in Learn mode, and only acts when you switch it on — with a permanent 10% holdout proving what it caused. Free to start, no credit card: a 14-day or 1,000-session live trial, then keep measuring for free.