The most repeated statistic about Amazon keyword tools right now claims that 62% of high-intent queries never show up in Helium 10 or Jungle Scout. Google quotes it in AI Overviews. Sellers forward it to each other. We looked for the study behind it and found only one agency blog post dated April 25, 2026, with no sample size, no category, no time period, and no second source.
That is roughly the state of the conversation about keyword research this year. The people telling you your tools are broken have not audited their own numbers.
It gets stranger. Amazon retired the Rufus name on May 13 and moved its shopping assistant into the main search bar. Two months later, most of the articles ranking for Amazon keyword research still discuss Rufus as a live product. One of the better guides in the set was published in June and does not mention the assistant at all.
So this is not another piece arguing that your tools are lying to you. Your tools are mostly fine. This is a records check on the advice you are getting about Amazon keyword research and on the three things that actually changed.
The tools are not lying to you. They are counting different things.
It is often discussed that keyword tools disagree by 30 to 50 percent. But none explain why, leaving sellers with a vague sense of betrayal and no way to act on it.
Here is the mechanism. Helium 10 reports normalized search volume. Amazon Search Query Performance and Jungle Scout report denormalized volume. Helium 10’s VP of Education has published this. It is worth noticing because the vendor argues against the simplest reading of its own accuracy.
Denormalized counts search for events. If one shopper types “yoga mat” four times while comparing options, that counts as four searches. Normalization accounts for that repetition and instead estimates distinct shopping intent. So when your tool says 8,000 and Amazon says 24,000, neither number is necessarily wrong. They answer different questions. Comparing them is like comparing ad impressions to unique visitors and concluding your analytics platform is broken.
This has a practical cost. The standard advice is to run the same seed through two tools and treat any keyword where they diverge sharply as suspect. If you follow that without understanding normalization, you will discard good keywords because much of the divergence is structural. You are not detecting errors. You are detecting a unit conversion.
The genuine accuracy question is narrower and easier to answer: inside one tool and one category, does the rank order hold? That is testable, and we cover how to do so below.
| WARNING
Backend search terms are indexed by bytes, not characters. Accented vowels, umlauts, and non-Latin characters can consume up to four bytes each, so a 249-character string can quietly exceed the limit. Amazon does not truncate the field when that happens. It ignores the entire field, with no error and no notification. Amazon’s Seller Central documentation sets a 250-byte limit for the search terms attribute. |
Check the number before you build a strategy on it
Go back to the 62% claim, because it is a clean specimen.
One post, one date, one publisher. It claims that 62 percent of high-intent Amazon queries in 2026 never appear in Helium 10 or Jungle Scout because the assistant rewrites them before they reach the search bar. There is no study, no sample size, no categories, no time window, no method, and no replication. We searched for corroboration and found only pages citing the original post. Google’s AI Overview presents the figure as a settled fact.
Watch what happens next because that is the part worth worrying about. Once an AI Overview quotes a claim, other publishers cite the Overview. Their posts become additional sources. Within a quarter, a single unsourced assertion takes on the shape of consensus, and every article that repeats it technically cites something real. Nothing was verified. The claim simply accumulates citations until it looks like a finding. That is how one agency’s marketing becomes an industry fact, and the loop now runs faster than any correction.
The same pattern repeats across the set. An agency post from December 2025 claims intent-based keywords deliver two to three times higher sales velocity. Elsewhere on the same page, it claims four times more sales and a 30 percent profit lift. No methodology accompanies any of those numbers.
The best example is a claim by a keyword tool that third-party estimates are 30 to 50 percent off Amazon’s actual data. Consider that for a second. Off compared to what? The same article explains that Amazon does not expose real search volume through any public channel. If the ground truth is unavailable, nobody can measure the distance from it. The number is a rhetorical device wearing a lab coat.
Now look at who publishes each diagnosis. The volume trap is diagnosed by a tool that sells intent scoring. Old tools being systematically wrong is diagnosed by an agency selling the replacement approach. Tool fragmentation is diagnosed by a platform selling consolidation. Every diagnosis in this category ends in a purchase.
That does not make any of them wrong. Vendors are often right about their own category, and several pieces we read are thoughtful. It means the standard of evidence should rise, not fall, when a claim supports the seller’s product. Right now it falls because the claim arrives wrapped in a statistic and statistics feel like evidence.
| WARNING
The 62 percent figure circulating in AI answers traces to a single agency post with no published sample size or method, and we found no independent corroboration. It is currently being quoted by Google as fact. Do not plan inventory or ad budget against it. |
The name changed in May. Most keyword advice has not caught up.
On May 13, 2026, Amazon discontinued the standalone Rufus chatbot and replaced it with Alexa for Shopping. Amazon says Rufus continues to power parts of the experience behind the scenes, but the consumer-facing name is no longer visible in the US interface.
Three changes matter for keyword work, and none of them is the name.
Location moved. Rufus lived behind an icon in the corner of the page, which most shoppers never clicked. Alexa for Shopping sits in the main search bar and generates answers above the results grid.
Reach changed. Rufus was an opt-in beta that Amazon reported reaching more than 300 million customers in 2025. Alexa for Shopping is the default for every signed-in US customer, free, with no Prime membership and no Echo device required.
Output changed. Side-by-side product comparisons and generated answers now render before a shopper scrolls to a single detail page.
The consequence for keyword strategy is that the moment you win or lose has moved earlier in the session. A shopper who gets a useful generated answer may never scroll the results at all. Your listing can hold position one organically and still lose the answer slot to a competitor whose page gave the assistant clearer material to work with. Rank and recommendation are now two separate contests, and you can win the first while losing the second.
The rollout is US-first. Other marketplaces may still display Rufus branding, so international teams should check their own interface before rewriting anything.
| KEY FACT
Amazon retired the standalone Rufus chatbot on May 13, 2026. Alexa for Shopping replaced it in the main Amazon search bar, free to every signed-in US customer, with no Prime membership or Echo device required. Rufus reached over 300 million customers in 2025 as an opt-in beta before the change. |
There is a symmetry here worth naming. The industry spent two months using a name Amazon deleted. It has spent two years using a name Amazon never used.
“A10” appears across this category as though it were an official Amazon algorithm. It is not. Amazon publishes no documentation under that name. It is a seller-community nickname that repetition has laundered into apparent fact, and the more careful publishers say so plainly. One keyword tool we read attributes the semantic shift to A9 in one article and to “the A10 algorithm, which is Amazon’s latest ranking thing” in another, on the same domain. Nobody is checking.
If a platform-wide rebrand can sit unnoticed in this category for two months, and a fictional algorithm name can circulate for two years, the correct posture toward the rest of the corpus is not trust. It is verification.
What Amazon Keyword Research Tools Still Get Right
Reverse ASIN is genuinely good and has no free substitute. Cerebro and Keyword Scout do a job Amazon will never do for you: telling you which terms a competitor’s ASIN ranks for. That capability alone justifies a subscription for most brands, and none of the criticism in this article addresses it.
Seasonality data holds up. Knowing a term triples in November is directionally reliable and useful for inventory and ad planning.
Rank ordering inside one tool and one category is usually sound. If your tool reports term A gets more traffic than term B, that relationship is generally right even when both absolute numbers are wrong. Estimates built on a consistent method tend to be biased in the same direction, preserving the ordering.
Index checking at the catalog scale is real work that would take days to do by hand.
So the failure is not the tool. It is treating an estimate as a measurement. An estimate that ranks correctly but misses the absolute value is good for deciding which term goes in your title. It is not good enough to order 8,000 units against it. Use estimates for relative decisions and first-party data for absolute ones, and most frustration in this category disappears.
Where Amazon Keyword Research Tools Stop Looking
Every keyword tool on the market scores the same four fields: title, bullets, description, and backend search terms. That was the right map when Amazon search was string matching. It is now a partial map.
When a shopper asks the assistant a question, it pulls from the title, bullets, A+ Content, structured attributes, customer reviews, and the Q&A section, then synthesizes an answer. It is not scanning for your exact-match term. It is deciding whether your page contains the material to answer a specific question well.
Your keyword tool scores none of the following:
- Whether your A+ Content answers a use-case question or just restates features in a nicer font
- The text inside your images, which a machine can read
- Whether your Q&A section addresses the questions shoppers actually ask
- Whether your reviews contain specific, quotable detail or generic praise
This produces a failure that looks like nothing. A listing can be fully indexed, score green across every field in your stack, and still get routed around because it lacks the sentence that answers the question.
Take a bullet that reads “Stainless steel construction. BPA-free. Dishwasher safe.” It is indexed for all three terms but says nothing about who the product is for, what problem it solves, or how it compares to the obvious alternative. A competitor whose bullets answer those three things gets cited in the generated answer. You do not. Your keyword tool correctly reports no issues because there are none. There is a content issue no keyword tool was built to see.
You can check your exposure in about ten minutes without any tool. Open one of your top ASINs, ask the assistant the five questions your customer service inbox receives most often, and read the answers. If it answers using your listing, the content is doing its job. If it hedges or answers using a competitor below you in the rankings, you have found the gap before your quarter does. That output is the closest thing to a scorecard for the fields nothing in your stack measures.
None of this is an argument against keyword tools. It is an argument about scope. Keyword research tells you which door to stand at. It says nothing about whether you have anything worth saying when it opens, which is a copywriting and creative problem, not a research one. We work on that half, so read the last two paragraphs with that in mind.
What to do when two tools disagree
Everyone in this category diagnoses differently. Almost nobody prescribes. The closest any competitor gets is “treat that keyword with caution,” which is not a procedure. Here is the one we run, in order, and it takes about a morning.
- Check whether you are comparing normalized to denormalized. If the argument is Helium 10 against Search Query Performance, the gap may be structural rather than an error. Compare within a single measurement system before you conclude anything about accuracy.
- Rank-order against Brand Analytics Search Frequency Rank. It gives you no raw volume, which is exactly why people ignore it, and it gives you something better: a rank ordering of real Amazon search terms. If your tool’s top ten and Amazon’s top ten broadly agree on order, the tool is directionally fine for that category, and you can stop worrying. If they diverge sharply, believe Amazon.
- Run a small exact-match Sponsored Products test. One term: modest budget, exact match. Impressions tell you demand exists. Conversion indicates that demand is buying rather than browsing. Give it two weeks. Do not call after three days, because statistical noise will point you the wrong way, and you will act on it.
- Read your own Search Term Report. These are the actual queries that triggered your ads. Not an estimate, not a panel, not a model. It is the most under-used dataset most sellers already own, and it is free.
The reframe worth taking from this: stop asking which tool is right. That question has no answer and no deadline. Ask which claim you can falsify fastest, then go falsify it. Two weeks of your own conversion data outrank every estimate in this article, including ours.
| TIP
Brand Analytics Search Frequency Rank is free with Brand Registry and rank-orders real Amazon search terms. It is the cheapest reality check available on any tool’s volume estimate, and most sellers paying for the tool have never once run it. Product Opportunity Explorer is also free, needs a Professional account, and does not require Brand Registry at all, which almost nobody seems to know. |
Amazon Keyword Research at Catalog Scale
Every article on this subject is written for a seller with one product. That is not the audience here, and the omission is intentional. Tool vendors price per seat, so the per-ASIN model is the one that is discussed.
At 40 SKUs, researching every listing independently still works. At 200, it collapses because of economics, not accuracy. Proper keyword research for a single listing takes 2 to 4 hours of competent work. Multiply that by 200, then repeat next quarter as search behavior shifts. Most brands respond the same way: research the top ten revenue SKUs properly and leave the other 190 with launch-day copy.
That is a rational response to bad math and the reason catalogs stall just when they should be compounding.
The way out is to stop treating every SKU as an independent research problem. Products cluster into families sharing buyer intent. A skincare brand with eight body butter scents does not have eight keyword problems. It has one keyword problem with eight variant expressions. Research once at the family level, then differentiate only where the variant genuinely changes what the buyer types. Scent changes the search. Color usually does not. We group catalogs this way under Product Family Architecture. The mechanism is simple: find the parent-level intent clusters, do the deep work once per family, then handle variants at the field level rather than rebuilding from zero.
You compete with yourself. Two ASINs in the same family, stuffed with the same head term, split-click share, and neither ranks as well as a single consolidated listing would. Your keyword tool cannot detect this because it evaluates one ASIN at a time and has no concept of your catalog as a whole. Cannibalization is invisible inside the tool that caused it. We cover how to diagnose and fix this at the catalog level in our large-catalog optimization guide.
Quick test you can run today: list every parent ASIN and count how many have five or more children. That ratio shows how much of your optimization problem is grouping rather than keywords, and it usually surprises people.
The stack we actually run
For discovery, one reverse ASIN tool is needed. Cerebro or Keyword Scout. Pick one and stay in it. Running both to compare volume numbers is what generates the normalization confusion described above. Running both to compare keyword coverage is a different question and perfectly reasonable.
For validation, Brand Analytics Search Frequency Rank and Product Opportunity Explorer. Both are free. Both are first-party. Both are routinely skipped by people paying $900 a year for estimates of the same thing.
For ground truth, use your Search Term Reports.
For language, manual reading. Sit with your reviews and your competitors’ Q&A sections for an hour. Shoppers describe products in ways no database generates. “Does this fit a standard 30-inch cabinet?” appears in no keyword tool and is precisely the kind of thing the assistant now gets asked and answers from whichever listing bothered to address it.
For deployment, the four fields your tool scores and then the fields it does not. A+ Content written to answer questions rather than decorate. Image text a machine can parse. Q&A seeded with the real questions from step four. This is part of the work we do, and if you want it handled, our Amazon Keyword Research and listing SEO optimization team can take it from the research stage onward.
That is more work than opening a dashboard and exporting a list. It is also the difference between a listing that indexes and a listing that gets recommended, and those have not been the same thing since May.
| Before you renew anything this quarter
Run the four-step Amazon keyword research check on your five highest-revenue ASINs. It costs a morning and it will tell you whether your keyword data is the problem. If it turns out the gap is in what your listing says rather than which term you picked, that is the part we work on. Tell us what you found. |


