Shopify App Store Ads: what 310,000 impressions say about Relevance, bids and dead keywords
Published 20 August 2026 by Adot Technologies Inc, the team behind Arvio. Data window: June 2025 – June 2026, plus one 7-day window in August 2026. Every figure comes from keyword-level exports out of our own Shopify Partners ad dashboard; the full method is in the next section.
Key takeaways
- Shopify's Relevance label predicts click-through rate very well — and install rate not at all. Click-through spans 3.6× across the five bands; click-to-install stays flat at 24–31%.
- The auction sends volume away from relevance. Keywords Shopify labelled "Very high" were 58.4% of our list but took 6.4% of impressions. "Very low" was 2.8% of the list and took 24.2%.
- Most keyword slots never serve at all: 49.0% had zero impressions over a full year.
- The one bid-suggestion signal that separates a live keyword from a dead one is simply whether Shopify shows a range at all — not how low the range is. Keywords with a range got 3.18× the impressions per slot (95% CI 3.00–3.37).
Is there any public benchmark data for Shopify App Store Ads?
Not that we could find, which is why this post exists. Shopify shows you suggested bids and your own keyword metrics inside the Partner Dashboard, and publishes setup guidance for advertisers, but it releases no aggregate figures for the channel. We went looking for a keyword-level dataset published by a real advertiser before we ran our own numbers. There wasn't one.
So here is one account's ground truth: 310,520 impressions, 1,844 clicks and 488 installs over 365 days, broken out by the Relevance label Shopify itself assigns each keyword. It is one advertiser, not the industry, and the limits are listed in full at the end. It is still more than zero, and the arithmetic is reproducible from the tables below.
Why you should be sceptical of this post, and read it anyway
We build Arvio, an AI store operator for Shopify — it audits a store's SEO, product content, prices and stock, drafts the fixes, and holds every change for the merchant to approve. Dataset B below is Arvio's own ad account; Dataset A is our sister app, AiLead. We buy this inventory ourselves, so we have an obvious interest in how you think about the channel, and you should weigh the post accordingly.
So don't trust us. Check the method and the sample instead, and the section at the end that lists what would break these conclusions. Every number below comes from keyword-level exports out of the Shopify Partners ad dashboard. Nothing is modelled, estimated or extrapolated.
We're publishing it because if you run Google Ads or Meta you can go and look up what a normal CTR is, and here you can't. Every advertiser on this channel is calibrating against their own account and nobody else's.
Method
Sample. Two datasets, from two apps in two different categories, both ours:
| Dataset A | Dataset B | |
|---|---|---|
| App | AiLead (sales chatbot) | Arvio (store operations) |
| Rows | 7,620 keywords | 7,609 keyword slots |
| Window | Last 365 days (Jun 2025 – Jun 2026) |
Last 7 days (Aug 2026) |
| Volume | 310,520 impressions · 1,844 clicks · 488 installs | 8,673 impressions · 42 clicks |
| Match type | 7,478 Broad · 142 Exact | 7,599 Broad · 10 Exact |
| Geo | Not split — all rows are All |
Same |
The windows are different and we do not merge them. Dataset A answers "what does Relevance do to performance". Dataset B answers "which keyword slots serve at all". Any sentence below is about one dataset or the other, never both.
Selection bias, stated plainly. These are keywords we chose to buy, mined from our own search-term reports. This is therefore the distribution of a mid-size advertiser's keyword list — not the distribution of App Store search demand. If you want to know what merchants search for, this dataset cannot tell you.
How the rates are computed. CTR per band is impression-weighted (total clicks ÷ total impressions in the band), not the mean of per-keyword CTRs. That choice matters, so we checked it: the unweighted per-keyword mean gives 1.088% / 0.851% / 0.646% / 0.491% / 0.301% — same order, same monotonicity. The finding is not an artefact of weighting.
What we excluded, and what happened when we did. Both datasets are overwhelmingly Broad match, with a small Exact tail. Dropping the Exact rows entirely moves the Dataset A CTR curve from 1.121 / 0.811 / 0.656 / 0.565 / 0.310 to 1.121 / 0.808 / 0.656 / 0.565 / 0.310 — one digit in one band. We kept them.
Right-censoring. In Dataset B, 7.0% (536/7,609) of bid suggestions are capped at the display ceiling $75.00+. Any average of suggested bids is therefore biased low, and we don't publish one.
Timing caveat. Bid suggestions are read at export time. They are not locked to the metric window, so a suggestion shown next to a 7-day impression count was not necessarily the suggestion in force during those 7 days.
Account size, so you can calibrate. This is a small advertiser: annual spend on this channel is in the low four figures, not five or six. Every rate below should be read as a small account's numbers — we have no way to know whether they hold at ten or a hundred times the budget, and we'd be surprised if all of them did.
What we deliberately don't publish. Absolute spend and absolute cost per install. Those are commercial. Everything cost-related below is expressed as a ratio within the same dataset, which preserves every conclusion and leaks nothing. Given the impression, click and install counts above, anyone determined to estimate our costs can get close; the ratios are what we're standing behind.
1. Relevance predicts clicks. It does not predict installs.
Dataset A, 310,520 impressions, split by the Relevance label Shopify assigns each keyword:
| Relevance | Impressions | Clicks | CTR | Installs | Click → install | Cost per install (indexed) |
|---|---|---|---|---|---|---|
| Very high | 19,982 | 224 | 1.121% | 64 | 28.6% | 1.00× |
| High | 41,815 | 339 | 0.811% | 105 | 31.0% | 1.20× |
| Medium | 74,207 | 487 | 0.656% | 124 | 25.5% | 1.47× |
| Low | 99,321 | 561 | 0.565% | 137 | 24.4% | 1.60× |
| Very low | 75,195 | 233 | 0.310% | 58 | 24.9% | 1.83× |
| All | 310,520 | 1,844 | 0.594% | 488 | 26.5% | — |
The CTR column is perfectly monotonic across all five bands, a 3.6× spread from top to bottom. As a predictor of whether a merchant clicks your card, the Relevance label works.
Now read the next column. Click-to-install is 31.0% for "High" and 24.9% for "Very low", and it isn't monotonic: "High" beats "Very high". Across a 3.6× swing in click-through, the probability that a click turns into an install barely moves.
That shows up in the cost column. A "Very low" keyword costs 1.83× per install compared with "Very high", and the whole of that penalty is paid getting the click rather than converting it. Merchants who arrive from a badly-matched keyword install at about the same rate as everyone else. There just aren't many of them.
So if you have been reading Relevance as a proxy for traffic quality, this data says it isn't one. It predicts click-through, and nothing past the click.
2. The auction sends volume away from the keywords Shopify calls relevant
This is the part we did not expect. The same dataset, counted by keyword instead of by impression:
| Relevance | Keywords | Share of list | Impressions | Share of impressions |
|---|---|---|---|---|
| Very high | 4,453 | 58.4% | 19,982 | 6.4% |
| High | 997 | 13.1% | 41,815 | 13.5% |
| Medium | 1,168 | 15.3% | 74,207 | 23.9% |
| Low | 785 | 10.3% | 99,321 | 32.0% |
| Very low | 217 | 2.8% | 75,195 | 24.2% |
Nearly six in ten of our keywords were labelled "Very high" relevance by Shopify, and together they captured 6.4% of the impressions the account received. At the other end, 217 keywords — under 3% of the list — pulled almost a quarter of all impressions.
We can't see the auction, so we can't tell you the mechanism with confidence. The shape is consistent with the obvious explanation: a keyword being highly relevant to your app says nothing about how many merchants type it, and the terms that describe your product precisely tend to be the terms nobody searches. Relevance is a match score, not a demand signal, and it is easy to read it as both.
The operational version: a keyword list that looks excellent in the Relevance column can still be starved of volume, and the dashboard will not flag this. Impression share by relevance band is not a view Shopify gives you. You have to build it.
3. Half of the keyword slots never serve
Dataset A, over a full year: 3,736 of 7,620 keywords — 49.0% — received zero impressions.
Dataset B, over 7 days: 5,819 of 7,609 slots — 76.5% — received zero impressions.
The two aren't in conflict: a keyword that serves rarely shows zero in a short window and non-zero in a long one. We quote both because neither window makes the dead half of the list look alive.
We're not saying that's the norm for the channel. But if you have never counted, the number is probably bigger than you assume, and "we have 7,000 keywords" describes a list, not a reach.
One concession, made here rather than buried at the end. Section 4 shows that on keywords where Shopify offered a bid range, our bid sat below the bottom of it 84.5% of the time. So an unknown share of these silent slots is not the channel refusing to serve them — it is us not paying to enter the auction. We can't split the two apart with observational data. Read "dead slot" as "produced nothing for us at our bids", not as "has no demand".
4. The only bid signal that separates live from dead is whether a range exists
Shopify shows a suggested bid next to each keyword. Sometimes it is a range ($16.00 – $75.00+), sometimes a single flat figure. We assumed for months that the useful signal was how expensive the suggestion was, and that cheap suggestions meant available inventory.
That was wrong. The useful signal is whether there is a range at all.
Dataset B, 7,609 slots over 7 days:
| Bid suggestion | Slots | Impressions per slot | Clicks per slot |
|---|---|---|---|
| Shows a range | 4,800 | 1.53 | 0.00813 |
| Shows a flat figure | 2,809 | 0.48 | 0.00107 |
| Ratio | — | 3.18× (95% CI 3.00–3.37) | 7.6× (95% CI 2.4–24.6) |
And the flat suggestions are barely a price at all: 93.1% of them (2,615 of 2,809) are exactly $1.00, which is the floor. Our reading is that a flat $1.00 means there is no auction to price — nobody bidding because nobody searching — so Shopify falls back to the floor.
A note on that second column, because the confidence interval is doing real work. The whole 7-day window contains only 42 clicks. The direction is not in doubt: if clicks were distributed in proportion to slot counts, the flat group should have received 15.5 of those 42 clicks; it received 3 (exact binomial, one-sided, p = 1.0 × 10⁻⁵). But the magnitude — "7.6×" — has a 95% confidence interval running from 2.4× to 24.6×, and anyone quoting 7.6× as a fact is over-reading it. The impressions-per-slot figure of 3.18× is the one to use; it is built on 8,673 impressions rather than 42 clicks and the interval is tight.
One more figure from the same dataset, which surprised us: on keywords where Shopify did show a range, our bid was below the bottom of the suggested range 84.5% of the time (4,058 of 4,800). We were not outbid; we were not in the auction.
What we'd do differently, if we were starting this account again
Not advice, just what this data would have told us eighteen months earlier:
- Build the impression-share-by-relevance-band view on day one. It's the table in section 2, it takes ten lines of code, and it isn't in the dashboard.
- Use the presence of a bid range as the first filter on a keyword list, before looking at the suggested amount. A flat $1.00 is not a bargain, it's an empty auction.
- Stop reading Relevance as traffic quality. It's a click-through predictor. Judge quality on click-to-install, which in this account was roughly flat across every band.
- Count dead slots monthly. Half a keyword list can go quiet without any dashboard changing colour.
The first two are a groupby and a division — the FAQ at the bottom has the recipe, and you don't need anything from us to run them. If you'd rather compare notes than build it, we're easy to find: we're the team behind Arvio, and we'd take a second account's numbers over another blog post any day.
What we still don't know
- One advertiser, two apps, no control. We cannot separate "this is how the channel behaves" from "this is how our two apps behave in their two categories."
- Installs are Shopify-attributed. We cannot verify them independently at merchant level, so the install column inherits whatever Shopify's attribution does.
- No geo split. Both exports report
All. Country-level effects, which are large in every other channel we run, are invisible here. - We haven't tested the causal version. Everything above is observational. We have not taken a set of keywords and moved them between bid levels to see what happens — so read section 4 as "these two groups differ", not "showing a range causes impressions".
- Bid suggestions aren't time-aligned with the metric window (see Method).
If you run App Store Ads and your numbers disagree with ours, we'd genuinely like to know — a second account would roughly double the amount of public data on this channel.
FAQ
What is a good CTR for Shopify App Store Ads?
In this account, across 310,520 impressions over a year, the blended CTR was 0.594%. By relevance band it ran from 1.121% down to 0.310%. We'd caution against treating one account as a benchmark, which is exactly why we published the band-level split rather than a single number.
Does a higher Relevance score get me cheaper installs?
Cheaper, yes — 1.00× versus 1.83× per install between the top and bottom bands in this data. But the saving comes from click-through, not from install rate, which was roughly flat at 24–31% across all five bands.
Why do my "Very high" relevance keywords get no impressions?
Ours mostly didn't either — 58.4% of our keywords carried that label and they took 6.4% of impressions. Relevance measures how well a keyword matches your app, not how many merchants search it.
What does a flat $1.00 bid suggestion mean?
In our data it is overwhelmingly the floor shown when there is no auction to price — 93.1% of all flat suggestions were exactly $1.00, and those slots served at roughly a third the rate of keywords with a range.
Should I bid above the suggested range?
We can't answer that from observational data. We can say we were below the bottom of the range 84.5% of the time on keywords that had one, which is worth knowing before you conclude that a keyword "doesn't work".
How many keywords should I add?
This data doesn't support an answer, but it does undercut the premise: 49% of ours produced nothing in a year. Adding slots is not the same as adding reach.
Is Broad or Exact better?
We can't tell you — our list is 98% Broad in both datasets, so we have no meaningful Exact comparison. Anyone claiming a Broad-vs-Exact benchmark for this channel should be asked for their Exact sample size.
Can I reproduce this?
Yes, from your own account. Export keyword-level performance with the Relevance column, group by band, and compute clicks ÷ impressions and installs ÷ clicks per band. The whole analysis is a groupby and two divisions. If it doesn't replicate on your account, that's a more interesting result than this post.
Written by the team behind Arvio: AI Store Operator — the AI agent that fixes Shopify store SEO, product content, prices and stock in bulk, with every change waiting for your approval. We're a small app and we buy this ad inventory ourselves; treat this post as one advertiser opening its books, not as an industry study.
Does the Relevance score predict installs, or only clicks?
Only clicks, in this dataset. Across 310,520 impressions, click-through rate is perfectly monotonic across all five Relevance bands with a 3.6× spread from top to bottom, while click-to-install stays between 24% and 31% and is not even monotonic — "High" beats "Very high". Treat Relevance as a click-through predictor, not a traffic-quality score.
Is there any published keyword-level data for this channel?
We could not find one before publishing this. Shopify publishes no aggregate benchmarks, and the figures that circulate come from individual practitioners quoting their own accounts without the keyword-level breakdown. The tables above are our attempt to add one public sample. If you run App Store Ads, yours would be the second.
Related: Your SEO app found 400 missing meta descriptions. Now what? — the same approach applied to an audit report: which of the flagged items are actually worth fixing, and the four bulk routes for the ones that are.
