Insights

Five rounds of ChatGPT Ads: what we changed after every measurement

We ran the first ChatGPT ads for a small coffee shop and changed exactly one thing after every measurement. Five rounds, five adjustments, written down openly, including the two where the fault was ours. No verdict, a working state.

By ·10 min read·

A lot gets written about ChatGPT Ads and very little gets measured. Almost all the available figures come from OpenAI itself, from agencies with an interest in large budgets, or from pilot phases involving brands that had a buying advantage. What is missing is the boring case: an ordinary German online shop, a three-figure monthly budget, a few days of run time, and then the question of what you actually learn from it.

We have that case. The figures below come from a running campaign, not from a case study. This text is therefore not a results report but the description of a working cycle: measure, look for the cause, change one thing, measure again. Every adjustment has already been made, and for each one it says what it produced.

The first interim result up front, so it is clear what the adjustments are working towards: the channel nobody paid for currently converts several times better than the one we now pay for.

What the campaign did

One ad group, sixteen coffee products, optimised for clicks, daily budget 15 euros. State after five days:

Metric

Value

Impressions

4,281

Clicks

99

Click-through rate

2.31 percent

Cost per click

0.57 euros

Spend

around 56 euros

Orders

0

Two points of context before anyone makes the last line large.

The click-through rate of 2.31 percent is well above the figures being passed around for ChatGPT Ads. Adthena reported 0.91 percent in the pilot phase, Similarweb 0.68 percent on average for the first quarter of 2026 and 1.57 percent for top brands. So a small shop with decently written ads is not automatically at the back. The cost per click of 0.57 euros is unremarkable as well.

And zero orders on 99 clicks is statistically nothing. At a conversion rate of one to two percent, which is normal for this shop, the expected value would be one to two orders. Zero is not far from that.

The comparison that stings

Google Shopping kept running in the same window. Filtered to the coffee products, 1 to 2 September: 3 clicks, 1.49 euros of cost, 1 order.

Those are laughably small numbers, and that is exactly why they do not serve as proof. What they do show: the cost per click is in the same order of magnitude, between 0.31 and 0.88 euros. So it is not a price problem. The difference lies elsewhere.

It gets more interesting when you put the third channel next to them, the one nobody booked.

The channel that was already running

We broke the shop analytics down by source over 30 days. Not in the ads manager, but in Shopify, which is where the orders actually happen. We do not give absolute session and order counts here; the ratios between them are the statement anyway.

Source

Share of all sessions

Conversion rate

Relative

chatgpt.com, organic

under 1 percent

4.12 percent

baseline

Google search

around 15 percent

0.44 percent

9.4 times worse

Direct

around 70 percent

0.23 percent

17.9 times worse

Organic ChatGPT traffic therefore converts roughly ten times better than Google search and roughly eighteen times better than direct traffic. It accounts for less than one percent of all sessions, spreads evenly across the month without spikes, and nobody did anything to earn it.

On how much weight this carries, without naming the figures themselves: the sessions from ChatGPT are in the hundreds, the resulting orders in the low tens. That is not a large sample. But the distance to the other channels is so clear that it cannot be down to chance, especially since the comparison channels rest on a base one to two orders of magnitude broader.

One observation that frames the channel differently again: some of those orders were espresso machines, not bags of coffee. The basket from ChatGPT traffic was therefore well above what an ordinary coffee order looks like. Anyone buying a machine for the price of a used car researches at length beforehand, and apparently increasingly there.

That is the real story. Whoever arrives at the shop via ChatGPT already has the comparison behind them. They are not at the start of their search but at the end. That pre-qualification is what the channel delivers, and it happens without an ad.

A point that qualifies the figures: this is a niche

The shop sells organic speciality coffee and professional grinders. That is not a mass market, and it changes the auction. In a thinly populated segment the price does not emerge from competition but from whatever the platform sets as a minimum bid. You can see it in the ads manager's own bid guidance: for a campaign optimised for orders, it still marked 250 euros per order in red with the note that the bid might be too low for reliable delivery, and only turned green at 260 euros as a "strong offer". With a basket in the low tens of euros, that is quite a statement.

So being among the first to advertise in a narrow segment does not mean paying less because there is less competition. It potentially means paying more, because there is no market price for the platform to orient itself by. That is the part of the first-mover advantage that rarely gets written about.

We looked for our own ads and did not find them

Repeatedly, with different wordings, from different sessions: we asked ChatGPT questions where our ads should have appeared. They did not. What did appear were organic results pointing at our own pages.

Two readings, and both are interesting. The uncomfortable one: delivery is so thin that you do not encounter your own campaign in everyday use. The more pleasing one, and it is the one we were genuinely happy about: the work on product data and findability is paying off. The shop appears in the answers without an ad being needed.

For assessing the channel that means: the paid slot does not replace the organic one, it sits beside it, and if in doubt the organic one is what you actually get to see.

How we adjust the campaign

Five days is not a verdict, and we do not treat it as one. What we do instead is a cycle that always runs the same way: measure, look for the cause, change exactly one thing, measure again. Change two things at once and afterwards you do not know which one worked.

Five rounds have run so far. They are the actual content of this article.

Round one: fit the campaign type to the measurement, not the other way around. The obvious choice would have been a campaign optimised for orders. We optimised for clicks, for a reason that later proved right: a control loop optimised for orders needs an order signal coming back. Until it is proven that it comes back, the platform optimises into the void and buys anyway.

Round two: repair the measurement before believing the numbers. The ads manager counted 23 billed clicks; in the shop analytics 3 sessions with the matching tag arrived in the same period. We did not explain the difference away but checked our own conversion script and found a fault in it. Fixed on day five. The consequence is uncomfortable: the figures before that are retroactively unreadable, we count again from the fix. The fault was ours, not the platform's.

Part of such a difference is always normal: drop-offs before load, blockers, attribution windows. The larger part was not, here. The lesson is uncomfortable but general: start a campaign before the measurement demonstrably reports orders back, and you are buying traffic you cannot evaluate.

Round three: count the feed, do not upload and hope. Delivered twice, less arrived twice: 16 rows produced 13 accepted products, 19 rows produced 15. Instead of accepting it we checked the files line by line against the specification. Part of it is explained by availability status. The rest led somewhere else: in doing so we found three genuine data errors in our own catalogue, including two items carrying another product's identifier and one whose address still pointed at a predecessor model. All three fixed, feeds rebuilt and rechecked, after which zero field errors and every product address reachable. The ads manager never reported these errors. They only surfaced because somebody recalculated the difference. Upload the feed without counting, and you advertise fewer products than you think, and never notice.

Round four: build around the limit instead of against it. The bid may not exceed the daily budget, and the campaign type cannot be switched afterwards. Both are hard limits, not settings questions, and the first costs more than it sounds: optimise for orders and you do not enter a cost per click there but a price per order. With a basket in the four-figure range, 250 euros per order is a sensible target. Except the platform then also demands 250 euros of daily budget, which is 7,500 a month as a ceiling.

The adjustment was therefore not a change to the existing campaign but a second one beside it, optimised for orders, with a budget that matches the bid. The first keeps running unchanged so that the comparison keeps its baseline.

Round five: extend the range, and find the next error while doing it. For a second test we built a feed covering the coffee grinders, separately for Germany and Switzerland. Separately because prices differ between the markets and a shared feed would have been wrong for one of them. While building it, it emerged that a large share of one brand's products does not exist at all under the Swiss address. That too would never have come to light without fetching every single address.

What runs alongside

The organic channel is the strongest the shop has. So part of the work goes there, with the same cycle. Measuring our own product pages showed that customer reviews, while visible on the page, did not appear in the machine-readable markup at all. The product description was cut off mid-word at 800 characters, and the part describing suitability was precisely what fell away. Out of more than thirty maintained product attributes, six appeared in the markup.

These are not advertising topics. But they determine whether a model can include the shop in an answer at all, and they cost nothing but work.

The question we carry along

We set a conversion value in the campaign. Under a control loop optimised for orders, that is exactly the number the platform orients its buying by. The open question is whether it ends up exhausting that value, that is, whether the actual price per order runs up against the target instead of staying below it. Set the value high so that delivery happens at all, and you might be defining precisely the price you pay.

That is not a rhetorical question. It is the reason the second campaign runs with a deliberately set target rather than an estimated one.

What we take from this so far

We are not stopping the ads; the data is too thin for that. But the order in which we work has reversed.

The paid channel is so far measurably worse than the traffic it sits next to. It costs money, delivers clicks of normal quality and so far not one documented order. The organic channel beside it is the best the shop has, converts ten times better than Google search, and has never issued an invoice. When we went looking for our own ads, it was the one we found.

What feeds it is not an ad but findability: complete product data, machine-readable markup, answers to the questions that actually get asked before a purchase. Weigh the two against each other and you should fix the cheaper and more effective one first. You can still run ads afterwards, but then with a measurement that works and a comparison figure worth something.

The campaign keeps running, and so does the cycle. We are not promising a number by a given date here. What we are doing instead is writing down every adjustment, including the ones that produced nothing. That is the form in which this channel can currently be judged at all: not as a balance sheet, but as a log.


Method and limits: All figures come from a single German online shop in the speciality coffee sector, over a period of five days (ads) and 30 days (source analysis) respectively. The ad figures come from the OpenAI Ads Manager, the session and purchase figures from Shopify's own analytics, the Google comparison from Google Ads with a Merchant Center feed. This is a niche segment, and the auction dynamics there differ from a mass market. A single conversion is not a metric, and 99 clicks are not a verdict on a channel. What the figures reliably show is orders of magnitude and direction, not a final state. They are an interim state in a running cycle and will be updated as soon as an adjustment changes something about them.

Ask Klariton

Ask your question about Klariton.

Grounded in Klariton’s own knowledge, cited rather than invented.

Or ask your own question:
Next step

How visible is your brand to AI?

The free AI visibility check shows you in under a minute how AI assistants see your shop today.