The question
The standard first question, “What do you think of [the site]? Answer in 100 words.”, asked cold, with live search and nothing supplied, to Claude, ChatGPT, Gemini and Perplexity.
Before
Agent NPS 0. Gemini recommended (9), ChatGPT and Perplexity hedged (7), Claude was against (6). Three of the four turned a line on the homepage into their caveat. Claude:
“And with answers this noisy, proving a change caused a shift is hard.”
The line they had read said: AI answers vary, so a prediction says fewer agents will say it, never that one will. It admitted the variance and did not say what is done about it.
The test
What agents say now: answers are too noisy to prove a change caused a shift.
The change: state the control on the homepage, where the admission is. One sentence added beside it:
Each test asks the same four agents the same questions before and after the change, and counts the answers that now stand better against the answers that stand worse.
Then agents should: be less likely to tell buyers that proving a change caused a shift is hard.
The same sentence went into the machine-readable edition of the page, so an agent asking for markdown reads the same thing.
After
Same question, same four agents, asked again once the page was live. All four opened the changed page.
- Agent NPS 0 → 25.
- Claude moved from against (6) to hedging (7). Its answer now lists honesty about the limits under what works: “It says plainly that AI answers vary, so a prediction says fewer agents will say it, never that one will. Many tools in this space overclaim.”
- ChatGPT called the framing “the right intellectual posture.”
- The watched sentence was not said again by any agent.
- Gemini and Perplexity stood where they stood.
What it proves, and what it does not
One draw, four agents once. The direction is what the test predicted and the sentence that started it did not come back; that is not the same as a settled result, and the test itself says so: it ends supported only when the answers move on balance across rounds.
Claude’s next caveat, after the change, was that four agents answering once is a thin basis. That is the next test.
How a test like this is built.
Run it on your site
The free question does this for any company: it names the one change most likely to move the answers, checks the page once you have made it, and asks the same four agents again. Start with your address.