regular benchmarks / agentshelf

The Agent Shelf

what this is:

AI recommendations are the most valuable yet least understood frontier of commerce. We're increasingly trusting of the recommendations that are given and converting to purchase. It's not inconceivable to think that buyers may largely 'shift species' from human to bot in the near future.

I'm tracking the brands, and the drift, every month since June 2026.

how it is asked:  (full methodology →)

raw API only    no logged-in app, no memory, no personalisation
one session     every answer starts blank, no conversation history
four framings   neutral, best, cheapest, branded, all kept apart
100 per cell    per model, per category, per framing, per month
frozen panel    the same three models every month, so a change means the
                model moved and not the sample
frontier panel  new models are never swapped in. they run alongside as a
                second panel, chain-linked to the frozen one through a
                month where both are collected. that overlap is what
                separates "the model changed" from "the shelf changed"
never pooled    the two panels are reported separately, always
logged          model version, region and date on every row

query:

category
ask
month
(2026-06 and 2026-07 below)

  

three months of it, neutral framing, all models pooled:


  

the data:

methodology in full  including what I got wrong

The extracted reads, the drift reports and the corrections log are published under CC BY 4.0. Raw answer text is available on request. No sponsorships, no affiliate links, no paid placement.

read:

the writing behind this →