Your badge screen lists the requirements and says which ones a category clears. This page explains each one in the same words the screen uses, and why it is there. All of them must hold at once, on fresh evidence, for a badge to be issued.
Every category has its own chance level — how often any brand gets named there, given how many brands plausibly answer that need. In our round that ranges from about 30% to 43% depending on the category. Being named a lot in a category where everyone is named a lot is not a finding.
Clearing chance by a whisker is not evidence either; it is noise that happened to fall the right way. So each assistant has to beat its own chance level by a required margin, measured on the lower bound of a confidence interval rather than on the raw rate. And at least two assistants have to do it.
This used to be two rows on the screen, and they were the same test. Both asked whether two assistants clear the margin, so they could never disagree — the second one just looked like a different requirement. They are one row now.
The row has three states. Met shows how many assistants clear it and your best margin. Missing means no assistant beats its category baseline at all. And in between there is a faint check — not a tick, not a cross, not a dash.
That faint check is the most common state we measure, and the most misread. It does not mean your margin is thin: in our round, 15 of the 25 categories in this state had a margin above the required 15 points — some above 50. What is missing is the second assistant, not a bigger margin. The screen says so in those words, because “work on your margin” and “you need another assistant to name you” are different jobs.
This one is not one of the criteria, and it is on the list because leaving it off was a lie by omission.
Two assistants above chance makes a category a candidate. A badge needs the number the round declares — three. Without this row, a category could show every other requirement met and still carry no badge, which reads as arbitrary or broken. The threshold comes from the sealed round itself, never from a copy written into the screen.
We ask each buying need in five different wordings, then split them into two halves and compare. If your rate collapses between halves, the signal belongs to a phrasing rather than to the need.
This is not hypothetical. In our round one brand is named in 91.7% of answers to two wordings of the same need and 55.6% of the other three — a 36-point gap, with a permutation p of 0.0005. Same category, same round, same brand. The screen shows how much your rate moves between wordings, because “inconsistent” with a 7-point spread and “inconsistent” with a 36-point spread are not the same news.
A category asked in only one wording cannot be split, so this requirement is not computable there — which is a result, not a pass.
A signal that traces back to one repeated citation is fragile, and it is also the cheapest thing to manufacture: one affiliate page, quoted everywhere. We measure how concentrated the citing domains are, and flag a category where they are too concentrated.
A blank here means not evaluated, not clean. If we could not compute it, the screen says so rather than showing a tick. One known limit: one of the three assistants returns citations as opaque redirect links, so its source diversity is not reliable — we treat its grounding as present and its domains as unresolved, and never rest a badge on it alone.
A recommendation for something a shopper cannot buy is not a recommendation. This requires at least one citation that resolves to a real retail channel in the measured market. The screen lists what it found.
A requirement about us, not about you. An assistant only counts if it returned enough usable, grounded answers with enough wordings in that round. If too few did, the category is not certifiable regardless of how it looks — and that is our gap, not yours.
Records run on evidence within a 90-day window. The screen shows the age of yours and the limit.
Two clocks, and they are not the same. The 90 days above govern whether the record still qualifies. Your badge on a store is current for 30 days from the measurement date; after that it keeps showing but dated — never a negative mark. So the badge goes dated long before the record stops qualifying, and paying keeps the mark current without changing what the record says.
There is a requirement about stability between rounds: a record should not swing from one measurement to the next. It is evaluated from the second round onward, and today there is one.
We deliberately do not list it. Putting it on your screen today would put on your list something that depends entirely on us measuring again — which is not a requirement you can do anything about. When the second round exists it appears here, and then it is genuinely yours.