How to measure AI-written marketing

When drafting is expensive, output is a reasonable proxy for effort, and "we published twelve posts" feels like a report. When drafting is nearly free, output stops meaning anything. Twelve posts is not a result. It's a receipt for electricity.
So the measurement step needs rebuilding. This is the fourth step of the loop we run, and the one most often skipped — partly because it's the only step that can tell you the previous three were wasted.
Table of Contents
- The arithmetic that explained our own numbers
- Four numbers worth tracking
- Measure against the audit, not against zero
- How long to wait before judging
- The reports that mislead
The arithmetic that explained our own numbers
For a long stretch, this site produced no leads. Not a low number. Zero. The instinct was that something was broken — a form, a mail provider, a tracking script.
Nothing was broken. The arithmetic simply never allowed for a lead.
Search Console said the site was getting roughly 33 clicks a week. Of those, about two landed on a page that could actually convert — a services page, a page with a booking link. The rest arrived on blog posts answering informational questions, read the answer, and left, which is exactly what someone searching an informational question does.
Two commercial visitors a week is eight a month. At a generous 5% conversion rate on a page like ours, that predicts well under one lead a month. Zero is not an anomaly in that model. Zero is the expected outcome rounded down.
That single calculation was worth more than any dashboard, because it told us the problem wasn't the form or the copy on the booking page. It was that essentially nobody with commercial intent was arriving. You fix that upstream, not at the form.
It also reframed every piece of content we'd published. The posts were doing their job — ranking, earning impressions, answering questions. They just weren't connected to anything. Nobody had checked whether the traffic they earned could get anywhere.
Four numbers worth tracking
Volume metrics (posts published, emails sent, creatives made) belong in a work log, not a report. These four belong in the report:
1. Commercial-intent sessions. Not total sessions. The count of visits that reached a page where a purchase or an enquiry is possible. This is the number that scales into revenue, and it's usually a small fraction of the number in your analytics headline.
2. Impressions on terms you targeted. For content, the earliest honest signal. Rankings move before clicks do, and clicks move before revenue does. If impressions on your target terms aren't moving after a quarter, the content isn't landing, regardless of how much of it there is.
3. Flow revenue share. For email, the share of revenue coming from automated flows rather than campaigns. Campaign revenue is a function of how often you send. Flow revenue is a function of whether the system you built works, which is what you're actually trying to measure.
4. The gap between impressions and clicks on your best pages. This is the cheapest win in the list and almost nobody looks at it. A page with thousands of impressions and a poor click-through rate is already ranking — it's the title and description that are failing. Rewriting those is an afternoon's work against traffic you have already earned.
Measure against the audit, not against zero
"Did it work?" is unanswerable, which is why it gets asked — an unanswerable question can't produce bad news.
The answerable version requires a baseline, which is what the audit step is for. You counted every asset and its current numbers before you started. Three months later you count again and diff. Now the question is specific: these eleven posts targeted these terms; impressions on those terms moved this much; of the resulting clicks, this many reached a commercial page.
That question has an answer, and sometimes the answer is that the work did nothing. That outcome is still worth buying. The alternative is running the same play for another year because nobody defined what failure would look like.
How long to wait before judging
Judging too early is as expensive as not measuring. Rough windows we use:
- Email flows — two to four weeks for enough sends to be meaningful, longer for low-traffic stores.
- Campaigns — immediate, but per-send numbers are noisy; look at trailing four sends.
- Content — a full quarter before any conclusion. Anything sooner is reading noise.
- Ads — days for creative fatigue, weeks for anything structural.
The mismatch worth naming: AI can produce a quarter of content in a week, but it cannot make that content take less than a quarter to prove itself. Speed of production and speed of feedback are unrelated, and confusing them is how brands end up abandoning approaches that were about to start working.
The reports that mislead
Three specific traps:
- Last-click attribution on a long consideration cycle. If your buyer takes six weeks and four touches, last-click will tell you branded search is your best channel. Branded search is a scoreboard, not a channel.
- Aggregate site conversion rate. It blends people who arrived intending to buy with people who arrived to read a definition, and it moves when the mix changes even if nothing improved.
- Open rates. Since privacy-protection features started pre-fetching images, opens measure inbox providers, not humans.
FAQ
What's the minimum measurement setup for a small store?
Search Console, your email platform's flow revenue report, and one spreadsheet holding the audit baseline. That's enough to answer the four questions above. Everything else is refinement.
How does this work if a tool is doing the publishing?
It should close its own loop — measuring what it published rather than handing you a send log. That's what the last step of Tilly is: it publishes into Klaviyo, Omnisend, Shopify Email or Meta, then reports what happened, and that report is the input to the next audit. A tool that only tells you what it sent has automated three steps out of four and left you the hardest one.
What if measurement says the work didn't do anything?
Then you saved yourself the next quarter of it. That's the entire return on the measurement step, and it's why the loop closes rather than ending.