The Measurement Was the Easy Part
AI visibility became a commodity in about a year. AI legibility did not.
Twelve months ago, finding out how ChatGPT described your company was a small research project. You wrote prompts by hand, ran them a few times, screenshotted the answers, and argued with a colleague about what they meant.
That period is over. Otterly starts at $29 a month. HubSpot gives away an AEO grader. Semrush folded AI visibility into a platform your team already pays for. At the top of the market, Profound raised $155M at a billion-dollar valuation and sells to the Fortune 500, and the category took in more than $300M between mid-2025 and spring 2026.
Measurement went from scarce to abundant in about a year. This is good. It is also the end of measurement as a service.
The dashboard produces a finding, not an instruction
Open any of these tools and you get a version of the same sentence. You appear in 12 percent of relevant answers. Your competitor appears in 41 percent. Here are the sources being cited instead of yours.
Now what.
The honest answer, and the better vendors give it, is that this is where the product stops. Profound reports where you are cited and where competitors appear; it does not publish or update anything. Scrunch is diagnostics-first: it explains where and why you show up, not what to build. Peec is monitoring, full stop. One industry review put it more bluntly than any of them would: data without execution is overhead.
This is not a complaint. Instrumentation is a real business and these are good instruments. It means only that the interesting problem sits downstream of them, and that nobody has yet turned it into a reliable system, because the missing layer is not generation. It is judgment.
What actually has to be fixed
Consider a shape that is common in any distribution-heavy industry.
A woman owns a retail store. She also runs a distribution company that imports several brands into a region. She teaches a certification course for professionals. She is opening a clinic. She is launching a curation mark for products she has vetted.
Five entities. Ten possible relationships between them. To a customer standing in front of her, this is one person and one reputation. To a language model, it is five weakly connected strings, several of them confused with unrelated companies of similar name, most with no confirmed relationship to each other at all.
No dashboard will tell you which of the five should be canonical. That is a decision about the business: which entity carries the authority, which ones inherit it, which are described as subsidiary, which are deliberately kept apart. Get it wrong and you spend a year reinforcing the weakest node. The work is closer to corporate structuring than to content marketing, and it needs someone who understands both the industry and how a model assembles a picture of an entity.
What building the machine teaches you
I built a recommendation layer for beauty retail: a system that takes a person's inputs and returns products, deployed in actual stores, making actual recommendations to actual customers.
What you learn quickly from that side of the interface is that the system does not know brands and does not rank them by fame. It resolves entities, looks for corroboration, and then produces something it can defend. A brand mentioned everywhere but resolving ambiguously loses to a brand mentioned less often that resolves cleanly, because only the second one can be attached to a claim without risk.
Most companies believe they have a volume problem. They have a resolution problem. They are not unknown. They are unusable.
Legibility is the difference between the two: whether a model can identify an entity correctly, place it in relation to the entities around it, and attach a claim to it without taking on risk.
Nothing holds still
Here is the part the category does not like to say out loud.
The industry is slowly discovering that the object it measures is stochastic. Recent work from SparkToro found meaningful variation in AI brand recommendations across identical prompts. The consequence is severe: a single measurement, from any vendor, at any price, is closer to weather than to climate. Point-in-time numbers describe volatility at least as often as they describe position.
This applies to my numbers as much as to anyone else's. A screenshot of a good answer proves almost nothing. Only a fixed prompt set, run repeatedly, against named model versions, with language and market held separate, produces something you can reason about.
And a real gain does not stay bought. Models retrain. Competitors publish. Sources rot. The structure that made a company legible in March is quietly degraded by August with nobody having touched it.
That makes this maintenance, not a project. The right analogy is not a website redesign, which ends. It is security, which does not.
The field has no standards yet
There is no agreed metric. No shared test set. No published protocol that two practitioners could run independently and compare. No convention for how many runs constitute a result, or what size of change is worth calling a change.
Worth saying plainly, because the confidence in this market currently exceeds the evidence in it. Anyone who tells you exactly what moves a model, and by how much, is describing a hypothesis in the tone of a finding.
What does exist is a set of practices that can be written down, run in the open, and attacked: publish the prompt set before the result, name the models and their versions, separate languages and markets, measure at fixed intervals instead of once, keep a control, and log everything that changed in between.
That is a smaller claim than the market is making. It is also the only one that survives contact with a second person trying to reproduce it.
Questions
What is the difference between AI visibility and AI legibility?
Visibility measures how often a brand is mentioned in AI answers. Legibility measures whether a model can identify the entity correctly, place it in relation to the entities around it, and attach a claim to it without taking on risk. A brand mentioned often but resolving ambiguously loses to a brand mentioned less that resolves cleanly.
Why can AI visibility tools not fix what they find?
The tools report where a brand is cited and where competitors appear. The repair is usually a decision about entity structure: which entity should be canonical, which ones inherit authority, which are described as subsidiary. That is a judgment about the business, not a generation task, which is why no dashboard closes the gap.
Is a single measurement of AI visibility reliable?
No. Model answers vary across identical prompts, so a point-in-time number can describe volatility rather than position. A usable measurement requires a fixed prompt set, repeated runs, named model versions, and languages and markets held separate.
Does AI legibility work stay done?
No. Models retrain, competitors publish, sources decay. A structure that made a company legible in one quarter degrades in the next without anyone touching it, which makes this maintenance rather than a one-off project.