Nobody was measuring this in Portugal. So we started.
What AI assistants say about Portuguese brands was in no study at all: the figures in circulation come from American vendors about American markets, and serve as direction rather than measurement. We ask the questions here, in Portuguese, and publish what comes out, including when it does not flatter us.
Five rules, each born from a mistake.
None of these is an intention. All of them are implemented in the code that produces the numbers, with tests, before being written on this page.
Questions are fixed before there are answers
What counts as a mention, and what counts as a recommendation, is written before the first collection run. Any criterion invented after seeing results is a decision contaminated by what you want them to say.
Numbers that go out come from code, not from a conversation
Our first study was counted by hand and got three of fourteen denominators wrong. Today a study's numbers come from a script anyone can re-run, with the denominators defined in one place, and the result lives in git so a future run shows the difference as a diff.
Human judgement is stored as data
Reading an answer and deciding whether it recommends a brand or merely lists it cannot be derived. That reading is done by people, recorded line by line, and stays auditable. It does not live in the head of whoever read it.
Brand names that are also common words are counted by reading the sentence
In Portuguese, NOS is also a pronoun, ERA is also a verb, Livre is both a party and an adjective. Each of these traps has a test in our code, because one collection run where the brand is absent is enough for a pattern to invent presence that was never there.
We say what the numbers do not say
A rate measured over thirty answers is not the same as one measured over eight hundred, and one week does not prove a trend. Every study states its base and what falls outside it.
The mistakes stay in plain sight.
A house that publishes primary data will get things wrong. The difference between research and marketing with charts is what happens next. When one of our numbers is wrong, we say so here, with the old number and the reason.
In the study on political parties, our first count put one party in 57% of one assistant's listings and 88% of another's. Thirty points apart, on the country's second largest party.
The difference was ours. The detector had been written imagining prose and the assistants answer in lists. Once the shape was fixed, all three read 100% and the difference disappeared entirely. A difference between assistants is a hypothesis about our own code before it is a finding about them.
We once wrote that one assistant did not search the web in 49% of its answers.
Withdrawn. Those rows did not come from the API we claimed to be measuring, and we had the means to know that before asserting it. Every stored answer now declares how it was collected, and a claim about an assistant is checked first against latency, token counts and who wrote the row.
Have a question nobody has measured?
Three cases worth writing about. Journalists who want a cut of the sector they are writing on, with the denominators explained before they publish. Universities and research centres: the collection machinery exists and the data is open. Companies that want a question from their market measured properly and published, with the note that whoever pays for a study does not choose its result, and that this is stated in the methodology.