Ekaterina Shalel Essays
The Legibility Layer · Essay

Can You Actually Change What AI Says About You?

A mention is retrieval. A recommendation is selection. One month of measurements on a single subject, and the subject is me.

By Ekaterina Shalel, founder and legibility strategist · August 17, 2026

Most people asking this question are actually asking about mentions. They want to show up. They want their name to appear when someone types a question into a model.

That's a different thing, and the gap between them is where almost all the work is.

A mention is retrieval. A recommendation is selection. Between them sits corroboration, and that's where almost everyone stalls.

Recommendation is a different class of event

When a model mentions you, it risks nothing. Your name sits in a list. If you don't belong there, nobody notices and nothing breaks.

When a model recommends you, it's making a call with a cost attached. Someone is going to spend time, money, or reputation based on the answer. And a system under that kind of pressure has a strong pull toward the safe option: the name it has seen many times, from many places, confirmed by sources that don't all trace back to the same person.

That pull is the whole problem. Being known isn't enough. You have to be the option a system is willing to be wrong about.

I've been calling the underlying question the Indifference Test: is the system structurally indifferent to which option wins, or is something in its inputs quietly tilting the outcome? Recommendation is where that test stops being theoretical. The tilt either includes you or it doesn't.

Four layers, and the one nobody works on

Retrieved. Read. Corroborated. Chosen.

The first three I've written about elsewhere. Retrieved means a system can find something. Read means what it finds resolves into you rather than a person with a similar name. Corroborated means the claim survives coming from somewhere other than your own site.

Chosen is separate from all three, and it's the one that gets skipped. Most people work hard on the first layer, get stuck on the third without knowing it, and never learn the fourth exists as its own thing. You can be retrievable, readable, and reasonably corroborated, and still not be the answer when a system has to pick.

The subject is me, and that's a design choice

I've been running this on myself since July. Single subject, tight control over what changes, dense measurement over time.

It's a real class of study. Single-subject designs go back a long way, and they come with real limits, which I'll state now rather than bury: one subject, no control group, and a retrieval layer that shifts on its own while you're measuring. Model versions change. Indexes update. Some of what I observe isn't caused by anything I did.

The reason the subject is me isn't vanity, it's access. No client is going to let me touch their public entity graph every week and log every change, including the ones that don't work. I will.

I also had a useful problem to start with. I was building a methodology around machine legibility while being almost entirely illegible to machines myself. That made me the cheapest available test case with the worst possible starting position, which is the best combination you can get.

The log

July 16. A clean query on my name returns other people who share it. My site doesn't surface. My LinkedIn doesn't surface. When I add enough context to force the right person, the description is stale: a founder in beauty retail, an identity I'd already moved past.

August 7. Three systems resolve the name correctly. One still doesn't. It's rebuilding me out of old signals from platforms I stopped treating as primary months ago.

August 16, first run. Cold session, no login. The prompt contains neither my name nor any company I've built. Correct reconstruction, and selection out of the category.

August 16, second run. Different system, different question, same day. Someone deciding who to hire for this kind of work. Named as the example, with reasons attached.

Four points, one month. The last two aren't the same result, and the difference between them is most of what follows.

What one run is worth, and what it isn't

Here's the part I want to be careful about, because it cuts both ways.

Retrieval isn't deterministic. The same prompt, in the same hour, on the same model, can return you and then not return you. Nothing about you changed in between. The system just resampled.

So the honest unit of measurement isn't presence. It's the share of runs. Three in ten. Seven in ten. A single screenshot tells you almost nothing, because you can't see the distribution it was pulled from. This is why the protocol I use is thirty-two prompts across systems and languages, run the same day, rather than one good answer saved as an image. The quantity is probabilistic. Without repetition it doesn't exist as a number at all.

And now the other half, which I think matters more.

One run in ten isn't "basically nothing." That one run is somebody's actual question. They don't do ten trials and average the results. They ask once, get a confident answer with no error bar attached to it, and go make a decision with it.

From where I'm measuring, a single run is noise. From where the client stands, that single run was the entire answer.

There's an asymmetry underneath this. While you appear in zero runs, you don't exist. The moment you appear in some, every run you land is a full event, not a fractional one. Presence doesn't get averaged down on the user's side of the screen. It either happened to them or it didn't.

That's the argument for measuring properly and the argument for not dismissing a single result, at the same time.

The runs where something had to be decided

The most useful session wasn't the flattering one.

Someone was weighing whether to pursue a professional conversation with a specific person, with a real cost to getting it wrong. The model didn't just retrieve me. It built a working picture of what kind of professional I am, made the case for engaging, then turned around and flagged the downside of that profile and proposed how to stress-test it.

That last move is why the run is worth reporting. It didn't flatter. It weighed. A model that agrees with everything tells you nothing, for the same reason a visibility audit with no negative control tells you nothing.

This is what I mean by a decision-level result rather than a visibility result. Not "the system knows about me." The system was put in a position where it had to act, and it acted.

The second system, the same day, raised the bar in a specific way.

The first run was selection out of a category. The second was a recommendation to a buyer. Those sound like the same thing and they aren't. Selection asks which name belongs at the top of a set. Recommendation asks whether a specific person should spend their money on this, and it carries the cost of being wrong. The second is the harder bar, and it's the one the whole field is chasing while measuring the first.

Here's how it went. Asked how to hire for this kind of work, the system listed the labels a buyer should search for, and one of the search strings it returned was the phrase I've been using for my own category, sitting alongside established ones. Then, asked to define that phrase, it defined it from my site. Then, asked flatly whether I'm any good, it said yes and explained why: a narrow, specific niche rather than a vague consultant label, with concrete deliverables described publicly.

A term I put into circulation came back as a way for a stranger to search, and I was the source of its definition. That's a different kind of result from being found.

But it also did the thing I care about most.

It hedged the attribution. Describes herself as. Says she helps. Seems credible. Every source it cited traced back to me. And then it produced a section called things to verify, with a checklist for the buyer: has she worked with companies like yours, can she show before and after outcomes, does she offer a baseline, a plan, and a follow-up measurement.

That's a system telling a buyer, in its own words, that the claim is well-formed and thinly corroborated. I've been calling that the corroboration gap for a year. I've never seen a machine name it back to me about myself.

It's also the strongest argument I have for why the fourth layer is separate from the first three. Both systems retrieved me, read me correctly, and chose me. Both then flagged where the evidence runs out. Selection happened anyway, and the flag came attached to it. Those are two different results in one answer, and no visibility metric I know of would report the second one.

The finding I didn't expect

The same answer that resolved me correctly still described a company I built as what I'm currently doing. It isn't. It's the proof the method works, and I've said so publicly for months, in the canonical places, in the schema, in every profile.

So the system updated who I am faster than it updated what I'm doing now.

Identity resolved. State lagged.

I don't know what's causing that. It could be source weighting, retrieval freshness, caching, how the query gets formed, or something else entirely. What I can say is narrower and more useful: those two kinds of information are updating at different speeds, and everything in this field measures only the faster one. Every audit, every tracker, every visibility score I've seen asks whether the entity is recognized. None of them ask whether the recognized entity is current.

I don't have a number for that gap yet. It's the next thing I'm instrumenting.

What this doesn't prove

It doesn't prove the method generalizes.

One subject, one month, no control group, and a moving substrate underneath. It can't tell me how much of the change was caused by what I did, and it can't tell me whether the same sequence would reproduce on someone else. Those are real gaps and I'm not going to paper over them.

They're also not the same as having no evidence. There's an intervention, a dated log, and a documented change in how machines represent one entity, with the runs that went nowhere included. What's missing is causal identification and external validity, which is a different sentence from "this doesn't show anything."

What it does establish is a sequence rather than a screenshot: a starting state where the entity didn't resolve at all, a middle state where it resolved into something out of date, and a point where it was reconstructed correctly and selected without being named in the prompt. Each stage dated, logged, and repeatable by anyone willing to run the same prompts on themselves.

A ranking can be noise. A generous answer can be noise. A documented change in machine representation, with the failed runs included, is harder to wave away.

If you want to argue with it, the useful question isn't whether one screenshot counts. It's what the control condition was. I'd rather be asked that than told I'm visible.

The protocol is published and it isn't mine to keep: thirty-two prompts, four systems, two languages, run the same day, failed runs recorded alongside the rest. Run it on yourself. I'd rather read someone else's numbers than publish my own twice. The repeat run on this entity is scheduled for late September, and I'll post whatever it returns, including the version where the gains don't hold.

Questions

Can you actually change what AI says about you?

Yes, and the change is observable when it is dated and repeated. Over one month a single entity moved from not resolving at all, to resolving into an outdated description, to being reconstructed correctly and selected without being named in the prompt. What makes that credible is the log rather than the final screenshot: fixed prompts, recorded dates, and the runs that returned nothing kept alongside the ones that worked. What a log like this cannot establish is that the same sequence reproduces on another entity.

What is the difference between being selected and being recommended by an AI system?

Selection asks which name belongs at the top of a set. Recommendation asks whether a specific person should spend money or time on this, and it carries a cost of being wrong. They are different results, and most visibility measurement reports the first while the field is chasing the second.

Does one good AI answer prove anything?

Not on its own. Retrieval is not deterministic, so the same prompt in the same hour can return you and then not return you, and a single screenshot gives no view of the distribution it came from. The measurable quantity is the share of runs. From the user side the arithmetic is different: they ask once and act on the answer they get, so a single run that lands is a full event rather than a fraction of one.

Can a single-subject experiment say anything about AI visibility?

It can produce evidence but not proof of generalization. One subject, no control group and a retrieval layer that shifts on its own mean the change cannot be attributed cleanly to the intervention, and nothing guarantees it reproduces on another entity. What a dated log does provide is an intervention, a documented change in machine representation, and the failed runs recorded alongside the rest.