The Nuance Index measures one thing: how clearly AI understands a consumer category. Every Friday we ask one AI model the questions real shoppers ask — fifty separate conversations, fresh each time — and score what comes back. Not which brands the model likes: whether its picture of the category holds together at all.
Baseline means first measurement — no arrow yet. Arrows appear when we come back.
The verdict sits on two axes: Agreement, and Truth — Accuracy and Alignment averaged. When a card says "76 truth," that's this.
The brands the model names most, and how often each one actually leads an answer — is named first — because being visible isn't being chosen. Bands, not ranks: when two brands' results overlap, we say so instead of inventing a number one.
The reality check: the model's ranking against the real-world one, measured with a yardstick declared on each edition's face — bank assets one week, store shelves another. It changes by category, and we always tell you which one it is and what it leaves out. (When the yardstick is store shelves, marketplaces like Amazon stay out of the count on purpose: everyone is "carried" there by somebody, so it measures nothing.)
Reading the sign: positive means the answers place a brand above its real-world rank; negative, below. One worked example: Arrid sits 9th on the mass shelf and 21st in the answers — twelve places apart in a 24-brand set, a gap of −50. And brands the answers never name enter at the floor — a floor, not a measurement.
When a brand shows up zero times, that's not a finding until it survives three tests: does the model think the brand belongs in this category, does it know the brand at all, does it appear when you change the question. Absences that fail get cut — visibly, on the page.
It scores the category, not the brands. Brand positions are evidence.
It is not a visibility score. A brand can be highly visible inside a category the model fundamentally misunderstands. That's worse — and a frequency metric can't see it.
Scores never compare across categories. A 70 in one category says nothing about a 70 in another. The through-line of the season is the method, not the numbers.
One model, stated on every run line.
We publish the ingredients, not the recipe. The question types, run counts, inclusion rules and vetting tests are public. The exact wording and weighting are not.