Ekaterina Shalel Essays
Method · Companion to the canonical buffer

The Canonical Buffer Audit

Five questions that reveal the state of your evidence environment before you build anything.

By Ekaterina Shalel · July 29, 2026 · Concept: Every Model Has a Different Diet

A canonical buffer is the controlled layer between what is true inside a business and what AI systems can prove outside it. It holds the current, dated, verifiable version of an entity: its one-sentence definition, its facts, its terms, its history of updates. External platforms distribute and confirm that version. The buffer is where it is kept current.

Before building or fixing one, audit the environment it has to work in. Five questions, in order. Each produces a written artifact, and the artifacts together are the specification for your buffer.

1. What questions should AI systems be able to answer about you?

List the ten commercial and identity questions that matter: who you are, what category you belong to, what you are the answer to. Not what you want said about you. What a system assembling an answer for a stranger must be able to support.

Artifact: the question set. It becomes your fixed test set for every future measurement.

2. Where do the engines currently forage when they answer those questions?

Ask the same questions in ChatGPT, Claude, Gemini, and Perplexity. Record which sources each one cites. Expect little overlap between engines; published citation studies suggest the majority of cited domains appear in only one engine. That divergence is normal, and the record of it is your map.

Artifact: a source map per engine, dated.

3. Which version of you wins, and is it the current one?

Look for the strongest contradiction: an old bio, a stale profile, an outdated description that keeps surfacing. Engines tend to reproduce the version most consistently supported by the sources they can retrieve, not necessarily the version you published most recently. Find what is outproving you.

Artifact: a contradiction list, ranked by how often each outdated source appears.

4. What lives in your canonical buffer, and when was it last dated?

Check whether your site actually contains the current one-sentence definition of the entity, dated facts, and your terms, in a form a machine can extract. If your positioning exists only in your head and your pitch deck, it does not exist in the evidence environment.

Artifact: a gap list between what the buffer should hold and what it holds today.

5. What is confirmed only by you?

A claim that appears solely in your own materials is weak evidence. Identify the two or three facts that most need independent confirmation, and go earn it there: a client's page, a publication, a registry, a dataset someone else maintains.

Artifact: a confirmation backlog, with an owner and a target surface for each claim.

Honest limits: this audit describes and organizes your evidence environment; it does not guarantee citation or ranking. No published technique has demonstrated a durable, cross-platform, causal effect on AI visibility. Findings apply within the query set, markets, and languages you test, and engine behavior changes over time, so the audit is repeated against the same frozen question set, not performed once.

The audit is the diagnostic. The Legibility Sprint is the seven-day protocol that acts on it: llms.txt, structured data, entity seeding, machine-readable pages, and a before-and-after measurement across ChatGPT, Claude, and Perplexity on your fixed question set.