A Plainpaper playbookPremium

AI search visibility

Freeze a set of buyer questions, record how each AI engine answers them today, fix what keeps you out of the answers, and re-run the same audit monthly to prove movement.

By Plainpaper v1.1.1 CC-BY-4.0

A live board built with this playbook, running the Plain & Paper demo shop's own campaign — hit Explore to move around inside it, open cards included.

Use this playbook

This one is already in Plainpaper. Create a board, pick AI search visibility from the template list, and the phases, card types and starter cards are there waiting.

A premium playbook, included with every paid plan.

Open Plainpaper

Plainpaper is a shared canvas where an AI agent drafts a marketing campaign as cards on a board you review and approve. It works with Claude, ChatGPT or any other MCP client, and you can start for free, no credit card required.

What this playbook is for

Ranking first on Google while being invisible in ChatGPT is the new normal, and there is no rank tracker to buy your way out of it: the answers are generated fresh, vary run to run, and cite sources you may never have optimised for. What works is an audit you can repeat. This playbook freezes a set of questions your buyers actually ask an assistant, records what each engine says verbatim, diagnoses WHY the answer skips you (the reason is usually a source the engine trusts and you are absent from, not a tag you forgot), ships the fixes, and re-runs the exact same audit next month. The board is the series: run over run, the same questions, honest numbers. Call it GEO, AEO, or AI visibility; the work is the same.

The board it builds

Every playbook sets up a board with named phases running left to right, the card types that belong in each, and starter cards that show your agent the shape of the work. They are the first of each kind, not the last: your agent writes as many more as the campaign actually needs. Here is what AI search visibility lays out:

Phases
BaselineGapsFixProve
Card types
Question set Run record Gap Cited source Fix Note Result
Starter cards
8 pre-written cards, already linked

The method

This is the written method your agent works from when it fills this board, published here in full rather than kept behind the product.

What this audit measures

A ranking is a position on a list, and it holds still long enough to be tracked. An AI answer is prose written fresh each time from a handful of sources the engine decided to trust, so there is no position to hold and no rank tracker to buy your way out of the problem. What you can hold is a repeatable measurement: the same questions, the same protocol, month after month, against a record nobody edited.

The one rule: freeze the question set after the first baseline. Reword a question and you have started a new series rather than continued a trend. Everything this board is worth comes from cycle three being comparable to cycle one.

How to think about it

Questions first, then runs, then diagnosis, then fixes. Reversing that order produces the familiar failure: somebody ships an FAQ page, the answers do not change, and nobody can say whether the page was wrong or the wait was short.

  • An engine's answer is only as good as its sources, and those sources are visible. Every run record lists them. That list is the real map of the work, and it usually does not point at your website.
  • The third-party gap is normally the biggest one. Best-of and comparison questions get answered from review sites, community threads and roundups. Being present where the engine already looks beats persuading it to look somewhere new.
  • Answers are non-deterministic. Ask the same question three times and you can get three different source sets. A single run is weather; the median of three, repeated monthly, is climate.
  • Entity hygiene is cheap, boring and worth doing once. One consistent description everywhere, a real about page, current public pricing, schema. Expect no drama from it and do it anyway, because the alternative is an engine that is unsure what you are.
  • Fixes lag. Content has to be crawled, indexed, and then chosen. Judge a fix at the second cycle after it shipped, not the first.

Failure modes worth naming: mass-generated answer pages, which are exactly what the engines and every recent core update exist to filter; a question set built only from questions you already win; calling a good single run evidence; and quietly rewording a question because it kept losing.

Numbers to hold it against

These are the shape of normal rather than targets, and this field moves fast enough that your own run records outrank anything published, including this.

  • Expect a low baseline. Most brands appear in a minority of the answers to their own frozen set on the first cycle, often well under a third, and in fewer still with a link attached. Track mentioned and cited-with-link as two separate numbers: a mention persuades, a link converts.
  • Run-to-run variance is large. The same question asked three times commonly returns overlapping but different source sets, so treat a swing of under roughly a fifth in share of voice as noise until two cycles agree with each other.
  • Citations concentrate. In most categories a small number of sources supply the majority of citations, and community threads and review aggregators are over-represented relative to brand sites. Read the run records for that concentration before deciding anything.
  • Assistant referral traffic is small and unusually good. For most ecommerce sites it is a low single-digit share of search traffic, and it arrives with the question already answered, which is why it converts at a multiple of ordinary organic. Judge it on conversions, never on sessions.
  • Weeks, not days, between shipping and moving. Several weeks is normal. Two cycles before a verdict is the honest rule, and it is the rule people break first.
  • 15 to 25 questions is the working size: enough to survive variance, small enough to run three times per engine per cycle without the audit eating somebody's month.

What moves all of these: the category (regulated and high-consideration categories get cited very differently from commodity ones), whether anyone in the category publishes original data, and how much of the category's conversation happens inside one or two communities.

What has to be true before a card asks for approval

Run records. Actual engine output, pasted, never summarised from memory. Date and cycle stated, protocol identical to last time, three runs per question, every miss quoted verbatim. A fabricated run poisons every decision downstream of it and there is no way to tell later which run it was.

Gap cards. Classified as content, third-party, entity or freshness, carrying the pasted evidence that shows it and the question numbers it affects. A gap with no evidence is a hypothesis and says so on its own face.

Fix cards. The gap it closes, the question numbers it should move, and, where it ships copy, that copy in the exact form it will publish. An answer page answers the question in its first two sentences, under a heading phrased the way the question is phrased, with a visible date on it.

All of it. No mass-generated pages, no fake reviews, no prompt-injection text in your own markup. Those lines hold for a practical reason as much as an ethical one: the whole strategy here is being the kind of source an engine is right to cite.

The seed cards are the first instance, not the ceiling

Seven cards show the shape of one cycle. A running audit is far denser than that, and it is meant to grow every month.

A healthy Baseline holds the question set plus one run record per engine per cycle, so three engines across four cycles is twelve run cards and not one of them is ever edited. Gaps normally carries six to twelve gap cards and five to fifteen cited-source cards, one per source the engines keep returning to. Fix holds one card per fix, which in a first cycle is usually six to fifteen rather than a single ladder. Prove holds the share-of-voice card plus whatever per-engine detail is worth reading on its own.

If the Gaps phase holds one card after a full baseline, the baseline was read rather than mined.

The layout is one arrangement of many

The columns here run by stage of the loop: baseline, gaps, fix, prove. That fits the first cycle well. Once several cycles exist, some teams arrange by engine and some by cycle, keeping each month's run records together as a block, and both read better than the original once the board is a year old. Rearrange freely. The freeze rule and the never-edit rule survive any layout, and they are the only two things that have to.

When this is the wrong playbook

If the site is not indexed, or holds no content that answers any buyer question, this board will spend a month proving it. Fix the fundamentals first. If nobody in the category asks an assistant anything yet, the audit will be honest and dull, and the honest thing is to say so. If the reason engines cannot describe you is that nobody can, that is a positioning problem and no amount of schema will repair a company that has not decided what it is. And this board does not replace a competitor analysis: it tells you who the engines name instead of you, not why those competitors' customers stay.

How a playbook stays safe

A playbook can only ever propose: every card arrives as a draft and nothing leaves Plainpaper until you approve it, which is enforced by Plainpaper rather than by the playbook. Approvals & control covers the rules in full. Playbooks by Plainpaper are written and maintained by us.

geo aeo ai-visibility ai-search chatgpt perplexity seo

More playbooks like this

  • Competitor analysis: Let buyers define the competitive set, mine what rivals' customers actually say, and turn it into dated, living battlecards instead of a deck nobody reopens.
  • Positioning sprint: Walk April Dunford's positioning components in order, from real alternatives to a messaging hierarchy every other board can consume.
  • Ecommerce campaign: The end-to-end campaign board: brief, audience, strategy, the emails and creative themselves, then results measured back against the goal.

Browse every research & positioning playbook, or the full library.