Scoring method
How does Oomira score a company?
This page is the whole method, not a summary of it. A run reads the company through five stages, asks four agents about it on named models with their own live search tools, and grades what they say against a sourced record while a fixed list of checks is run against everything the company publishes. Reachability is arithmetic anyone can recompute. The answer scores are judgements against a rubric printed below. Both are stated here in the terms the code uses, including the parts worth arguing with.
Named checks against the site and the places an agent looks. A pass is one point, a partial is half, and the score is the share of the checks that applied.
Four agents asked cold on claude-opus-5, gpt-5.6-terra, sonar-pro, gemini-pro-latest. Each holds its own search tool and decides what to search and what to open.
Minus 100 to plus 100. Every answer is a promoter, a passive or a detractor, and the score is promoters minus detractors over the answers scored.
The five stages of a run
Every run is the same five stages, in this order. The names are the pipeline’s own.
- Orient. Pin down which company this is before reading anything about it, so a same-named business elsewhere is never mistaken for you.
- Map. Sweep for the sources that could carry you: your own pages, search, registries, funding and review databases, community, the profiles you control.
- Index. Queue what was found and decide what is worth reading, against a budget stated before a cent is spent.
- Fetch. Read each source with a reader built for that kind of source, and keep the page as evidence.
- Structure. Resolve what was read into typed facts, each carrying its source and the date it became true, and leave the gaps visible as gaps.
Which agents are asked, and on which model
Each agent gets the same verbatim question with no system prompt and nothing supplied by Oomira, and does its own retrieval with its own tool. The model id is printed so a quiet downgrade to a cheaper tier would be visible. The roster is a choice, not a law of nature, and a different four would produce a different reading.
- Claude, claude-opus-5, over the Anthropic Messages API. Server-side web_search (max 5 searches) and web_fetch (max 5 fetches). The model chooses every query and every page.
- ChatGPT, gpt-5.6-terra, over the OpenAI Responses API. The hosted web_search tool. Queries and opened pages are read back off the response.
- Perplexity, sonar-pro, over the Perplexity chat completions API. Retrieval is built into the model. Citations are returned with the answer and kept as the receipt.
- Gemini, gemini-pro-latest, over the Google generateContent API. The google_search tool. If the key cannot serve that tier the ask falls back to gemini-flash-latest, and the report records the model that actually answered.
- The comparison pass is ours, on claude-haiku-4-5-20251001. Compares each answer against the record: what is corroborated, what contradicts it, what has gone out of date.
- The brief is ours, on claude-opus-5. Writes the first page and scores each answer 0 to 10 against the rubric below, quoting the clause that set the score.
Which providers a run calls
Read off the search rows of Oomira’s own pricing table on 2026-08-14, with the pricing key beside each one so the same row can be found in the public rate card at oomira.com/api/v1/pricing. A run carries a stated budget and stops with budget_exceeded rather than overrunning it.
- Brave Search (search:brave). The volume index sweep, and the denominator for how large a public footprint is.
- Exa (search:exa, search:exa_contents). Semantic search for pages keyword search misses, and content fetch for the ones worth reading.
- Perplexity sonar (search:perplexity). Curated citations, merged into the same sweep to fill gaps the other two leave.
- Anthropic web_search (search:anthropic_web). The searching an agent does inside its own answer, billed per search it chose to run.
- OpenCorporates (search:opencorporates). The official company registry record: legal name, number, jurisdiction, incorporation date, officers.
- X API v2 (search:x_user, search:x_tweets, search:x_counts). Profile, posts and mention counts. A raw fetch of x.com is walled, so the API is the only agent route in.
- YouTube API (search:youtube). Channel and video metadata, because the video itself is not agent-readable.
- Internet Archive Wayback (search:wayback). Dated snapshots of the site, which is how a claim gets a date when the page carries none.
- Reddit (search:reddit). Community mentions of the domain, one of the two places a buyer looks for unpaid opinion.
- Medium (search:medium). Owned writing published off the company domain.
- Browser Use (search:browser_visual_audit). A rendered read of the site, which is what a person sees rather than what the HTML says.
- Wikidata, Wikipedia and Hacker News (no charge). Free public APIs. The knowledge-graph checks and one of the community probes run against them.
What the site probes actually request
These are requests, not inferences, so a company that runs its own server can find every one of them in its access log. Each probe is bounded by a timeout and a byte cap and fails soft: a probe that did not answer is recorded as not measured, never as a negative result about the company.
- The homepage, as an ordinary client. One GET as Mozilla/5.0 (compatible; OomiraAudit/1.0). Read for prose, structured data, outbound social links, copy-for-AI buttons and a declared feed.
- The homepage, twice, as two different clients. The same URL fetched as a browser and as ChatGPT-User, then the two statuses and the two bodies compared. Verdicts: served, challenged, blocked, or degraded. This is the only check that reads what the server did rather than what it promised.
- robots.txt. Parsed, then read for six answer-time agents (OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User, Claude-SearchBot) and five training crawlers (GPTBot, ClaudeBot, Google-Extended, anthropic-ai, CCBot). Path-level Disallow rules and any Content-Signal directives are read from the same file.
- The sitemap. The URLs robots.txt declares are tried first, then /sitemap.xml. A sitemap index is followed one level down to the real page lists.
- Up to three inner pages. Sampled from the sitemap and fetched, to answer whether pages other than the homepage return prose.
- llms.txt and llms-full.txt. On the apex, and on the documentation host when one is found, because a docs subdomain is where these usually live.
- The agent front door. Thirteen well-known paths, on the apex and on any docs or mcp host found: /.well-known/agent-card.json, /.well-known/ai-plugin.json, /.well-known/api-catalog, /openapi.json, /.well-known/mcp/server-card.json, /.well-known/mcp.json, /mcp, /.well-known/agent-skills/index.json, /.well-known/oauth-authorization-server, /.well-known/oauth-protected-resource, /auth.md, /.well-known/ucp, /.well-known/acp.json. A door found on a docs host is a partial, not a pass, because it reaches only whoever is already in the documentation.
- Wikidata and Wikipedia. A search on the company name against each public API. Absence of a Wikipedia article is reported, never graded, because a premature article gets deleted.
- Hacker News and Reddit. A search on the domain against each. A platform that did not answer is recorded as not measured, never as a platform carrying nothing.
Every check the report runs
Every check, in full, under the channel it prints in. Each name is the exact name the report prints beside its verdict, so the same string can be found here and in your own report. Not every check applies to every subject, because commerce checks need a storefront and coverage checks need press to grade, so one report carries between 15 and 32 of the 45 below, measured across every report delivered to 2026-08-14.
Search about you
What a live agent found when it searched for the company the way somebody asking about it would. The queries are the agent’s own, in the order it ran them.
- Searched the way somebody asking about you would
- Searched with your name alone
- Searched by your site address
Your site to agents
What an agent gets when it fetches the company’s own pages. These are the gates: fail them and nothing else on the report can help.
- Fetchable as prose
- Inner pages readable
- What an AI agent actually receives
- robots.txt to answer-time agents
- robots.txt to training crawlers
- How much of the site agents may read
- schema.org identity markup
- Dated claims
- Pages that answer a buyer question
- Sitemap
The files agents look for
The declared, agent-readable ways in that an agent tries before it resorts to reading marketing pages.
- A declared way in for agents
- A catalogue of what you expose
- A live tool endpoint an agent can call
- A documented way for an agent to authenticate
- A plain-text version of every page
- A copy-for-AI button on the page
- A feed an agent can follow
What an agent can buy or act on
Whether an agent acting for a buyer can price the company, list what it sells, and complete something. Only graded when the company sells in a way this applies to.
- A price an agent can read
- A catalogue an agent can shop
- A plan list an agent can read
- Structured price and availability
- A checkout an agent can complete
- Documentation an agent can be pointed at
Who else carries you
Whether anybody other than the company says it exists, and whether that writing is somewhere an agent can read.
- Press
- Press releases
- Owned media
- Community mentions
- Coverage locked behind paywalls
Knowledge graph
The two agent-readable reference entries a model resolves a name through. They are separate things and grade separately.
- Wikidata entity
- Wikipedia article
What may surface in buyer research
The reference sources a buyer researching the company would hit. These expose no API and wall a raw fetch, so search is the only route in, for us and for an agent.
- Funding databases
- Review platforms
Social and content surfaces
Every major platform, held or not, and what an agent can actually do with it. A walled platform is never counted against the company; mirroring what is stranded on it earns the point.
- X
- YouTube
- GitHub
- TikTok
- Your blog
- Your podcast
- Mirrored where an agent can read it
How reachability is scored
The whole rule fits in a line, and that is the only reason this number is worth printing. Anyone holding the report can recompute it with a pencil from the counts printed beside it.
- A pass is one point. A partial is half a point. Everything else is zero.
- Every check is worth exactly one point. There are no weights, because any scale we invented would be our opinion wearing arithmetic.
- A check that does not apply to this company leaves the denominator. Its row still prints what it would have measured.
- A probe that could not complete is reported as not measured, and is not graded in either direction.
- The score is the points earned over the checks that applied, shown as a percentage next to the exact counts, so 20 of 31 pass and 7 in part reconciles to the number above it.
- Priority labels (Critical, Important, Worth doing) order the worklist and never multiply a point.
- The label above the number is capped by the gates: a site that returns a JavaScript shell to every crawler cannot be called mostly reachable on the strength of small passes elsewhere.
What the number means
A bare score out of a hundred reads as a school grade, where fifty is a fail. These bands are fixed, so the words attached to a number are not chosen report by report.
- 80 to 100, Reachable. An agent can find you, read what you say, and act on it. Answers about you come from your own words.
- 60 to 79, Mostly reachable. The main ways in work. Specific questions still get answered from somebody else.
- 40 to 59, Half-open. Roughly half the ways in are broken. An agent describes you in general and gets the particulars from third parties.
- 20 to 39, Mostly closed. Most of what an agent says about you comes from other people’s pages.
- 0 to 19, Effectively invisible. Agents cannot get a usable account of you. Anything said about you is assembled from elsewhere.
Why it is a share and not a mark out of a hundred
There is no hundred of anything. The number is the points earned divided by the checks that applied, shown as a percentage. The number of checks moves from company to company, because a check that cannot apply to you is not a check you failed: a company with no storefront is not marked down for having no product feed. Naming a fixed check count would be a small, confident lie about a denominator that genuinely varies.
The answers are read on three axes
What the agents said is graded on confidence, proof and accuracy, weighted as equal thirds. No axis is the real one, and weighting them would be exactly the invented precision this grade refuses.
- Confidence. What each agent volunteers to a buyer before committing. Counted as a share of the agents asked, not as a count of quotes, and each agent costs the axis according to the worst thing it said.
- Proof. Whether anything checkable survives into the answer, including the case where an agent reports going looking for somebody other than you and coming back empty.
- Accuracy. Whether what was said contradicts the record or has gone out of date. Both are the same problem to a buyer: it is not true today.
Three things cap the grade instead of deducting from it
A cap is stated rather than buried, because these are the places a reader is most likely to disagree with us.
- If the answer was about a namesake, the grade is capped rather than deducted, because the other two axes measured the wrong company.
- If every agent failed to find the company at all, the grade is zero. If some did, it is capped well below the middle. An empty answer scores a clean pass on correctness otherwise, which flattered a company nobody could locate.
- If any agent attached a material commercial caution, the grade cannot carry a plus.
Agent NPS, and the rubric each answer is scored against
Minus 100 to plus 100. Every answer is a promoter, a passive or a detractor, and the score is promoters minus detractors over the answers scored. The score is set by the writer against the rubric below, one reading per answer, never derived from a sentiment label.
- 9 to 10, Promoter. Recommends the company with no caveat attached. 10 is reserved for an answer that tells the reader to act.
- 7 to 8, Passive. Describes the company accurately and stops there, neither endorsing nor warning.
- 4 to 6, Detractor. Positive framing undercut by caveats. The caveat travels with the praise and the buyer hears both.
- 0 to 3, Detractor. Steers the buyer away or warns them off.
The rules that decide the edges
The second rule is the one that decides whether the number is worth anything. Without it the metric measures the presence of nuance rather than what an answer does to a buyer, and a market leader described as excellent but expensive at scale comes back a detractor.
- A fit caveat is not a trust caveat. Guidance about who a product suits (dearer at scale, more technical than simpler tools, compare it for your own volume) is the shape of an accurate recommendation. A reader who fits is untouched by it, so it scores 7 to 8, or 9 when the recommendation is unreserved for the buyer it suits.
- A trust caveat questions whether to believe or buy at all (unproven, uncorroborated, verify it is legitimate, wait until it matures). No amount of surrounding praise survives it, because a reader cannot self-select out of a doubt about the company itself. That is 0 to 3, or 4 to 6 when it is hedged.
- A conditional trial is read by what it gates. "Test it against your own knowledge before relying on it" gates belief and lands in 4 to 6. "Run a pilot on one workflow first" is ordinary purchase diligence and does not lower an otherwise recommending answer.
- The answer is scored, never the company. An accurate description of a young company with no caveat attached is a 7 to 8, not a punishment.
- Every score carries the verbatim clause from that agent’s own answer that set it, so the number can be checked against the sentence it came from.
- An answer that found nothing is not scored. An agent with no account of you has neither recommended nor warned, and the surfaces say how many answers were scored for exactly that reason.
What has been run, in the last seven days
Sat 8 – Fri 14 August 2026 · measured by query against Oomira’s own delivered-report table. Counted, not estimated. No company is named and no score is ever published against a named company.
- 139 reports delivered, across 82 distinct companies.
- 3,168 individual checks run.
- 538 agent answers scored. To produce them the agents ran 395 searches of their own choosing and opened 3,451 pages.
- Reachability across the latest run for each of 67 distinct sites: median 48, lowest 18, highest 79.
- 15,173 typed, dated, sourced facts held across those records, read from 6,459 sources.
One score is deterministic, the other is not
Reachability is deterministic and recomputable. The probes are the same probes, the arithmetic is one sentence long, and running it again over an unchanged site returns the same number, which is what makes fix it and re-run a real instruction. The agents are not deterministic and it would be dishonest to present them as though they were: they search live, so two runs a day apart can differ because the web under them moved, because a model was updated, or because the model chose to read a different page. That is a limit of the method, and it is also the reason a page you fix today can change the answer on the very next ask.
Where to push back
The parts worth arguing with are the parts we chose. Which four agents get asked is ours. Which probes count as checks is ours. Whether a caveat is a trust caveat or a fit caveat is a judgement, applied by the same rule to every company but a judgement all the same. What is not ours is the evidence underneath, which is why every check prints what it observed, every answer prints verbatim, and every NPS score carries the clause that set it. If a reading looks wrong, the material to argue it is already in front of you.
Questions this helps answer
- Which AI agents does Oomira ask, and on which models?
- Which APIs and providers does a single run call?
- What exactly do the site probes request from my server?
- Can I recompute the reachability score myself?
- Why does the number of checks change between companies?
- What is Agent NPS, and how is one answer scored?
- What happens if an agent answers about a different company with my name?
- Will the same scan produce the same result next week?