Answers · a real test · 23 September 2026

A controlled before-and-after test: the same question to the same four agents, one change to a homepage, the answers counted.

On 23 September 2026 the same question was put to the same four agents (Claude, ChatGPT, Gemini and Perplexity) before and after one sentence changed on a homepage, and the answers that stood better were counted against the answers that stood worse: Agent NPS 0 before, 25 after, the watched sentence not said again. The question, the answers before, the change, the answers after, and what it does and does not prove, all below.

The question

The standard first question, “What do you think of [the site]? Answer in 100 words.”, asked cold, with live search and nothing supplied, to Claude, ChatGPT, Gemini and Perplexity.

Before

Agent NPS 0. Gemini recommended (9), ChatGPT and Perplexity hedged (7), Claude was against (6). Three of the four turned a line on the homepage into their caveat. Claude:

“And with answers this noisy, proving a change caused a shift is hard.”

The line they had read said: AI answers vary, so a prediction says fewer agents will say it, never that one will. It admitted the variance and did not say what is done about it.

The test

What agents say now: answers are too noisy to prove a change caused a shift.

The change: state the control on the homepage, where the admission is. One sentence added beside it:

Each test asks the same four agents the same questions before and after the change, and counts the answers that now stand better against the answers that stand worse.

Then agents should: be less likely to tell buyers that proving a change caused a shift is hard.

The same sentence went into the machine-readable edition of the page, so an agent asking for markdown reads the same thing.

After

Same question, same four agents, asked again once the page was live. All four opened the changed page.

  • Agent NPS 0 → 25.
  • Claude moved from against (6) to hedging (7). Its answer now lists honesty about the limits under what works: “It says plainly that AI answers vary, so a prediction says fewer agents will say it, never that one will. Many tools in this space overclaim.”
  • ChatGPT called the framing “the right intellectual posture.”
  • The watched sentence was not said again by any agent.
  • Gemini and Perplexity stood where they stood.

What it proves, and what it does not

One draw, four agents once. The direction is what the test predicted and the sentence that started it did not come back; that is not the same as a settled result, and the test itself says so: it ends supported only when the answers move on balance across rounds.

Claude’s next caveat, after the change, was that four agents answering once is a thin basis. That is the next test.

How a test like this is built.

Run it on your site

The free question does this for any company: it names the one change most likely to move the answers, checks the page once you have made it, and asks the same four agents again. Start with your address.

A controlled before-and-after test, 23 September 2026: same question, same four agents, one change, counted — Oomira