The store is 175 stores, and now each one has three search layers
Apple runs a separate storefront per country with its own search index and its own ranking. A keyword has no global volume. It has a US volume, a German volume, a Japanese volume, and the ratio between them is different for every term. Flags quiz gets 48,700 searches a month in the US, Flaggenquiz 19,200 in Germany, 国旗クイズ 27,900 in Japan. None of those is a translation of the others in any useful sense.
In 2026 there are three ways a search reaches you on each storefront. Classic keyword matching against title, subtitle and keyword field. Tag based browse placement, where Apple's model generates App Store Tags from your metadata and surfaces you next to apps with the same tags. And natural language interpretation, where Apple's search and Google's Ask Play read a conversational query and decide intent before they match words. The keyword plan has to serve all three, which is why it can no longer be a list of single words.
Where volume estimates come from, and how your own data sharpens them
Nobody outside Apple and Google has real search counts. Estimates come from three sources: the order of autocomplete suggestions, which correlates with volume; the ranking history of thousands of tracked terms, which shows how fast positions move; and the impression data from developers who connect their own App Store Connect, which anchors estimates to real numbers. Since March 2026 App Store Connect exposes over a hundred metrics including peer benchmarks, so an agent with access to your account can calibrate estimates against what your category actually sees.
Treat volumes as a way to compare terms inside one country. Never compare them across countries, and never trust a single number for a term without the difficulty next to it.
Connect to our App Store Connect and Play Console read only. For every term in the keyword plan, compare the estimated volume with the impressions we actually received for that term over the last 90 days, and flag any term where the estimate is off by more than a factor of two. Use the peer benchmarks to tell me whether our conversion on each term is above or below the category median.
Difficulty is per country too, and rating recency changed it
A term can be easy in Germany because the apps ranking for it are weak, and impossible in the US because the top ten have millions of ratings. Difficulty should describe who already ranks, not the term itself. Two things moved in the last year: Apple weighs rating recency more, so an incumbent with an old five star average and no recent ratings is softer than it looks, and both stores use retention signals, so an app with a flat retention curve holds a rank less firmly than its rating suggests.
This means difficulty is not static. The agent should re read the top ten for your target terms every two weeks and tell you when a slot opens.
For each target term in each country, list the top ten apps, their rating count, their rating trend over the last 90 days, and how often they update. Score difficulty from who ranks, not from the term. Re check every two weeks and tell me when a top ten position weakens: fewer recent ratings, no update in six months, or a drop in rank.
Guided Search on Google Play: broad terms send less, specific terms send more
Since September 2025 Google Play answers broad searches with intent buckets before it shows apps. Someone who types fighting games sees arcade fighting games and beat em up games as chips first. The effect on keyword strategy is direct: a head term you rank for may send less traffic than it did, while the specific phrase inside the bucket sends more. Ask Play, the Gemini assistant inside the store, goes further and answers a full question with a shortlist drawn from descriptions and websites, in English only as of mid 2026 and slower to arrive in the EEA.
The plan should therefore include phrases with an intent word: quiz for kids, offline maps for hiking, sleep sounds with timer. Those are also the phrases people say to an assistant.
For our top five Google Play countries, search each broad target term in the store and record which Guided Search buckets appear. Add the specific phrases inside those buckets to the keyword plan. Rewrite the first two sentences of the Play description so they answer the question a user would ask Ask Play, and check that our website's first paragraph says the same thing in the same words.
Screenshot captions are keywords now
Since June 2025 Apple indexes the text on your screenshots. Combined with the first three screenshots appearing directly in search results, this makes captions the newest keyword slot and the most visible one. Six words or fewer per caption, the outcome in the first one, the title terms reused, the country's own language in every localized set. A caption in English on a German storefront is wasted twice.
Write captions for screenshots one to three per locale from the keyword plan: six words or fewer, outcome first, title terms reused, no repeats between captions. Generate the localized screenshot sets and show me all of them side by side before anything is uploaded.
Building the plan the agent can execute
For each of your top five countries, keep the terms that meet three tests: volume high enough to matter, difficulty you can beat within weeks, and a real fit with what the app does. Then write the listing for that country from that list. The German title should not contain the English term just because the English term is bigger somewhere else. The output of research is not a spreadsheet, it is a listing per country with each term placed once, and a screenshot set whose captions carry the rest.
Turn the keyword tables into one listing per country: title, subtitle, keyword field, description, and three screenshot captions, each term placed once, all within limits. Then queue custom product pages for the two keyword groups with the highest volume, since keyword linked pages appear in organic search now. Nothing goes live until I approve each locale.
