The second score
What is Agent NPS, and how is it scored?
Net Promoter Score has one question behind it: would you recommend this company. Companies have asked it of customers for twenty years, and the number every operator already knows how to read is promoters minus detractors. Agent NPS asks the same question of the four agents a buyer actually consults before they ever contact you, cold, and scores each answer by what it does to the person reading it. Minus 100 to plus 100. Every answer is a promoter, a passive or a detractor, and the score is promoters minus detractors over the answers scored. The rubric is printed in full below, because a number nobody can score by hand is a number nobody should believe.
Four agents asked cold: Claude, ChatGPT, Perplexity, Gemini. No memory of you, no system prompt, nothing supplied by us, one live search tool each.
One answer, scored 0 to 10 by what it does to a buyer reading it, carrying the verbatim clause that set the score.
Promoters minus detractors over the answers scored, on the same minus-100 to plus-100 scale as the customer metric.
Why ask an agent a customer question
Because the agent is answering it already, to people you will never meet. A buyer who asks ChatGPT about your company gets a recommendation, a description or a warning, and that sentence does the work a reference call used to do. Sentiment scores tell you the mood of that sentence, which is not a thing anybody can act on or budget against. Promoters minus detractors is a number an operator has read for twenty years and knows immediately how to feel about, and it is only computable if you hold a judged score per answer with the clause that produced it. That is what a run produces, and it is what a sentiment label cannot be turned into afterwards.
Who answers, and under what conditions
The four are asked in parallel, each by its own API, each with its own live search tool, and each is given the same verbatim question with no system prompt and nothing supplied by Oomira. The model id actually sent is printed so a quiet downgrade to a cheaper tier would be visible: Claude is claude-opus-5, ChatGPT is gpt-5.6-terra, Perplexity is sonar-pro, Gemini is gemini-pro-latest. Every search each one chose to run and every page it opened is kept with the answer, so a score can be traced back to what produced it rather than argued about.
The rules that decide the hard edges
Most answers are easy to place. These are the rules for the ones that are not, and the first is the one that matters most: without it the metric measures the presence of nuance rather than what an answer does to a buyer, and a market leader described as excellent but expensive at scale comes back a detractor, which is the fastest way for a reader to stop believing the number.
- A fit caveat is not a trust caveat. Guidance about who a product suits (dearer at scale, more technical than simpler tools, compare it for your own volume) is the shape of an accurate recommendation. A reader who fits is untouched by it, so it scores 7 to 8, or 9 when the recommendation is unreserved for the buyer it suits.
- A trust caveat questions whether to believe or buy at all (unproven, uncorroborated, verify it is legitimate, wait until it matures). No amount of surrounding praise survives it, because a reader cannot self-select out of a doubt about the company itself. That is 0 to 3, or 4 to 6 when it is hedged.
- A conditional trial is read by what it gates. "Test it against your own knowledge before relying on it" gates belief and lands in 4 to 6. "Run a pilot on one workflow first" is ordinary purchase diligence and does not lower an otherwise recommending answer.
- The answer is scored, never the company. An accurate description of a young company with no caveat attached is a 7 to 8, not a punishment.
- Every score carries the verbatim clause from that agent’s own answer that set it, so the number can be checked against the sentence it came from.
- An answer that found nothing is not scored. An agent with no account of you has neither recommended nor warned, and the surfaces say how many answers were scored for exactly that reason.
Four into one number
Each answer lands in a band, and the score is the share of promoters minus the share of detractors, with passives counting toward the denominator and never toward the difference. So four answers with one promoter and three detractors is minus 50, and one changed mind is worth 25 points. Your respondents are not thousands of customers. They are the four agents a buyer actually asks, so a single answer is worth 25 points and one of them changing its mind moves a quarter of the score. That is a real limitation of a four-respondent panel and it is stated rather than hidden, because the alternative is a reader working it out and concluding we hoped they would not.
What a finished number means
The scale is borrowed whole from the customer metric, so its reading is borrowed too rather than invented for this. These are the four bands a finished score is read against:
- -100 to 0, needs work: At least as many answers cost you the deal as win it
- 0 to 30, good: A caveat is reaching buyers
- 30 to 70, great: Clearly positive, the odd answer only describes you
- 70 to 100, excellent: Recommended outright, unprompted
Why it moves, and why that is the point
Ask again tomorrow and the wording will differ. What is held constant is everything we control: the same verbatim question, no system prompt, no memory of you, nothing supplied by us, and no personalisation, so what moves is the web the agent searched rather than the way it was asked. That is why the report records every query it ran and every page it opened, in order. A number you can see the inputs to is worth more than one that never moves, and the gaps it finds do not change between days even when the sentences do. How far the deterministic half moves when the same site is scanned again is measured weekly and published at oomira.com/answers/would-the-same-scan-give-the-same-result.
What is free and what is not
The number is cheap to produce and the derivation is the work, so that is where the line sits. A free scan returns both scores and the four answers they came from. The report is how each score was reached: the clause that set every band, the searches and pages behind each answer, and the specific pages that have to change to move the next one. Nothing about the rubric is behind that line, which is why it is printed on this page rather than described.
Questions this helps answer
- What is Agent NPS?
- How is an AI agent’s answer about my company scored?
- Why is a positive answer with a caveat counted as a detractor?
- Which AI models does Oomira ask?
- What is a good Agent NPS score?