Use this playbook
This one is already in Plainpaper. Create a board, pick Paid social test-and-learn from the template list, and the phases, card types and starter cards are there waiting.
Open PlainpaperPlainpaper is a shared canvas where an AI agent drafts a marketing campaign as cards on a board you review and approve. It works with Claude, ChatGPT or any other MCP client, and you can start for free, no credit card required.
What this playbook is for
Paid social gets treated as a content treadmill: more creative, faster, and no memory of what was already learned. This playbook makes it an experiment log instead. Each test starts as a written hypothesis with a decision rule attached, produces variants that differ on exactly one axis so the result means something, and ends in a Learning card that survives the campaign it came from. The point is compounding: by the tenth test you should be able to tell a new agent what works for this brand and why, from the board alone.
The board it builds
Every playbook sets up a board with named phases running left to right, the card types that belong in each, and starter cards that show your agent the shape of the work. They are the first of each kind, not the last: your agent writes as many more as the campaign actually needs. Here is what Paid social test-and-learn lays out:
The method
This is the written method your agent works from when it fills this board, published here in full rather than kept behind the product.
What a test programme is for
To end up knowing things. Not to produce creative faster, and not to fill an account with ad sets. The output of this board is the Library phase: statements about what works for this brand, each with the spend behind it and the scope it applies to, written so somebody who was not here can act on them.
The one rule: the decision rule is written before the spend and never edited afterwards. Primary metric, minimum difference worth acting on, minimum spend before reading, and what you will do at that threshold. Written afterwards, every ambiguous result becomes a win, and a programme of wins that changed nothing is the most expensive way to run ads.
How to think about it
- One axis per test. Angle, format, hook, offer or audience. One. If the copy and the format and the audience all change, the winner tells you nothing you can reuse, and reuse is the entire point.
- Write the control down. An unwritten control drifts: the thing that "currently works" quietly becomes three things, and the comparison stops meaning anything. Restate it as an angle, with the numbers it holds today.
- More executions of the same idea is production, not testing. Both are necessary; only one belongs on this board, and confusing them is how a team spends a year busy and learns nothing.
- Order the axes by leverage. Creative and offer move cost per acquisition far more than bidding settings or placements. Test in that order, and be suspicious of a quarter spent on the last of those.
- Overlap contaminates. Audiences that share people bid against each other and blur the read, so exclusions are part of the design rather than an afterthought at launch.
- Inconclusive is a result. It stops the same idea being re-tested in three months by somebody who does not remember. Record it with the same care as a win.
Failure modes worth naming: calling a test at two thousand impressions because the gap looked big; splitting a small budget across six ad sets so none of them learns anything; reading day two as a verdict; and an empty Library at test ten, which means the programme ran experiments and kept none of the answers.
Numbers to hold it against
Read these as the physics of the channel rather than as targets. They decide what a budget can actually resolve, which is the question most test plans skip.
- Roughly 50 optimisation events per ad set per week is the working threshold for stable delivery. Below it the system is still exploring, and what you are reading is its exploration rather than your audience's opinion. It is the most useful number here, because it turns a budget into a maximum number of simultaneous tests.
- Resolution is expensive, and worth knowing before you promise a result. Separating a modest relative difference in click-through at a one percent baseline takes tens of thousands of impressions per arm; separating one in purchase rate takes hundreds of conversions per arm. Most declared winners in small accounts are coin flips read confidently.
- Run at least a full week, commonly ten to fourteen days. A shorter window reads one part of the weekly cycle, and reading before roughly day four reads the delivery system.
- Cold ecommerce link click-through commonly sits between 0.8 and 1.8 percent, varying more by offer and category than by execution. Use it to catch a broken setup, not to grade creative.
- Auction prices move by season far more than by creative. Fourth-quarter impressions routinely cost well above a first-quarter baseline, so a November test read against a February control compares two markets rather than two ideas. Always run the arms concurrently.
- Expect most tests to be losses or draws. Something like one clear winner in four or five is a normal, healthy rate. A programme reporting a winner every time has a reading problem, not a creative advantage.
- Frequency is a clock on every result. Performance in a fixed audience decays as the same people see the same ad, so a variant that wins in week one and fades in week three usually fatigued rather than failed. Record frequency next to the result or the learning comes out wrong.
What moves all of these: audience size, price point, purchase frequency, and how much of the account is retargeting. A small warm pool produces flattering numbers on tiny volume and resolves almost nothing.
What has to be true before a card asks for approval
The hypothesis card. A belief stated so it could be wrong, the single axis named, the primary metric, the minimum difference worth acting on, the minimum spend before reading, the window, and the call. Filled in before anything goes live, and not edited once spend starts.
The angle cards. Control restated explicitly with the numbers it currently holds. Variant differing on exactly the named axis, with the list of what is deliberately identical. If that list cannot be written, the test is not controlled.
The creative. The first frame carries the message muted and uncaptioned, because that is how most of it is seen. The claim matches the claim on the landing page, in the same words where possible. The destination is the page for this offer. One axis differs and nothing else does. No placeholder survives into awaiting_approval.
The test plan. Tracking verified end to end before launch: a broken pixel invalidates the whole spend and is discovered late every single time. Exclusions applied, including other live tests. Placements identical across arms.
The result. Read against the rule copied verbatim from the hypothesis card, with confounds listed honestly: seasonality, stock, a promotion running elsewhere, audience overlap, one arm stuck in learning.
The seed cards are the first instance, not the ceiling
Nine cards show one test. A programme is many tests, and one phase is meant to grow forever.
Hypothesis holds one card per open test plus its audience card, and a running programme usually has two or three open at once, budget permitting. Angles holds at least a control and a variant, often three or four when the angle itself is the axis. Produce holds one creative per angle per format, so four to eight cards rather than two: the same angle as a static image and as vertical video are different executions. Run holds one plan per test, Read one result per test, including the ones that separated nothing.
Library is the phase that never resets. By test ten it should hold ten to twenty pinned learnings, and it should be able to brief a new agent on what works for this brand with no other context. If Library is thin while Read is thick, the programme is generating results and discarding knowledge, which is the specific failure this board exists to prevent.
The layout is one arrangement of many
The columns run through one test cycle: hypothesis, angles, produce, run, read, library. Five of the six cycle; Library is cumulative, and that asymmetry is deliberate. Teams running several tests at once often prefer a column per test with a standing library beside them, and teams working one funnel sometimes arrange by stage. Rearrange it freely, but keep the library separate from whatever cycles: a learning filed inside a finished test is a learning nobody will find.
When this is the wrong playbook
Below a few thousand a month the account cannot buy resolution, and the honest programme tests offers and audiences, where effects are large, rather than creative nuances, where they are not. If tracking is broken, fix that first: until then every test measures the measurement. If the site converts its current traffic badly, more clicks is an expensive way to find out. And during a scaling push, when the job is spend efficiency rather than knowledge, running formal tests alongside it usually delivers neither: pause the programme, scale, come back with the budget the tests need.
How a playbook stays safe
A playbook can only ever propose: every card arrives as a draft and nothing leaves Plainpaper until you approve it, which is enforced by Plainpaper rather than by the playbook. Approvals & control covers the rules in full. Playbooks by Plainpaper are written and maintained by us.
More playbooks like this
- Ecommerce campaign: The end-to-end campaign board: brief, audience, strategy, the emails and creative themselves, then results measured back against the goal.
- BFCM sprint: A compressed Black Friday board: decide the offer and the cut list early, then run a five-touch sequence you have already written.
- AI search visibility: Freeze a set of buyer questions, record how each AI engine answers them today, fix what keeps you out of the answers, and re-run the same audit monthly to prove movement.
Browse every paid media playbook, or the full library.
Related reading
- Claude + Meta Ads: plan and approve paid social creative before anything spends.
- The Meta Ads MCP server in Claude: add Meta's official ads connector with no code and no API key.
- Create ad creative: creative shapes, mockup staging, and what approval requires.
- Templates in Plainpaper: how a playbook becomes a board, and how to save your own.
- Install Plainpaper in Claude: connect as a Claude connector, no install, about two minutes.
- Install Plainpaper in ChatGPT: the same board, reached from ChatGPT instead.