Methodology
How SellRadar scores a product: the methodology behind Launch, Watch and Skip
A verdict you cannot inspect is a score with better branding. This page documents exactly how a SellRadar scan turns a keyword into Launch, Watch or Skip, including the weights, the thresholds and the cases where we decline to answer.
Updated September 11, 2026 · Written by the SellRadar team
What a scan actually reads
Every SellRadar analysis starts from the same object: the top twenty organic listings that Amazon returns for your keyword, on the storefront you chose. Organic means the sponsored placements are stripped out first. Those twenty listings are the market you would be entering, because ranking on that page is the only way a new product gets found by a shopper who searched for the term.
For each listing we collect the fields Amazon exposes on the search page itself: title, price, rating, review count, badges such as Best Seller or Amazon's Choice, and, where Amazon publishes it, the bought-in-past-month figure. We also record how many products Amazon says match the query overall. Everything downstream is computed from those twenty rows. There is no hidden sales estimate, no scraped seller database and no third-party score being repackaged.
This matters for one reason. Because the inputs are the same rows you can see on Amazon yourself, every claim in the result can be checked against the page it came from. When you disagree with a verdict, you are disagreeing about a specific listing, not about a black box.
The gate: is this one market or a jumble?
Before any scoring happens, the twenty results have to pass a coherence check. Amazon does not return an empty page when a keyword is nonsense or too obscure. It falls back to loose matching and stitches together whatever shares a token with your query. A search for a phrase that does not correspond to a real product category can return kitchen knives, wool socks, a paperback and a pack of dog-waste bags on the same page.
Unrelated products are a problem for any scorer, and specifically for ours, because a page of mismatched items maximises exactly the properties that look like opportunity: wide price spread, no dominant brand, wildly uneven review counts. Scored naively, that page reads as a wide-open niche. It is not a niche at all.
So the gate measures whether the twenty titles share a category vocabulary, whether prices cluster the way a single product category's prices do, and whether Amazon's own match count suggests a real market. Fewer than eight organic results, or a match count under about a thousand products, fails the gate outright, since there is no statistic to compute over a handful of rows.
Between clearly coherent and clearly incoherent there is a mixed band, where a dominant cluster exists but a minority of the results belong elsewhere. Those pages are scored, and the written reasoning is allowed to name the mismatch rather than pretend the page is clean.
The seven signals
Each listing, and the page as a whole, is measured on seven signals. All of them are expressed on the same scale, from 0 to 100, where a high value means weakness in the incumbents and therefore room for a new entrant. Four are computed per listing; three describe the page as a whole.
Per-listing signals
- Review count. Bucketed rather than linear: a listing with no reviews scores near 100, under fifty reviews around 90, under two hundred around 70, under a thousand around 45, under five thousand around 20, and five thousand or more sits near the floor. The buckets are deliberate; the difference between 40 and 60 reviews is noise, the difference between 60 and 600 is a moat.
- Rating quality. Measured relative to the page, not against an absolute. A 4.1-star listing surrounded by 4.7s is weak; the same 4.1 on a page averaging 3.8 is strong. Shoppers compare within the results in front of them, so the signal does too.
- Listing surface quality. Heuristics on the title: length, qualifying words such as premium, professional, kit or set, and whether a brand leads the title. This is the weakest of the seven and carries the smallest weight. It is script-aware, so Japanese titles are not penalised for packing more meaning into fewer characters.
- Brand strength. Best Seller and Amazon's Choice badges, established review counts, and trademark or registered marks in the title or brand name. A badge and a mark on a listing with thousands of reviews is an incumbent you are unlikely to displace.
Page-level signals
- Demand. The average bought-in-past-month across the top five listings. This is the one signal Amazon publishes directly rather than something inferred from rank, which is why it carries the largest weight when it is available.
- Price stratification. The coefficient of variation of prices across the twenty, which is the standard deviation divided by the mean. A wide spread means shoppers are already paying different amounts for different positioning, and there is room to slot in. A tight cluster means the market has agreed on a price and you will compete on cost.
- Brand concentration. How many distinct brands hold the twenty positions, with all generic or unbranded listings counted as one shared bucket rather than as twenty tiny brands. A page owned by two names is a different proposition from a page held by fifteen.
The four per-listing signals are averaged into a weakness score for each competitor, and that score assigns a tag you see in the result: Low Reviews, Poor Rating, Weak Listing, Overpriced when a price sits more than one and a half standard deviations above the page mean, Strong when both review depth and brand strength are established, and Mid for everything else. That tagged list of twenty is the competitor map, and it is deliberately never collapsed into a single competition number.
How the signals become one score
Five components are combined into the opportunity score: demand, competition depth, price stratification, brand concentration and listing surface quality. Competition depth is the average review-count weakness of the top five ranked listings, since those five take most of the clicks. Demand enters inverted, because strong demand is a low weakness for the incumbents but a high opportunity for you.
The weights are not fixed. They depend on how much of the page carries the bought-in-past-month figure, because a demand signal computed from three listings out of twenty is not a demand signal. We call this data confidence, and it has three tiers.
| Component | High confidence | Medium confidence | Low confidence |
|---|---|---|---|
| Demand (bought past month) | 40% | 30% | 0% |
| Competition depth (top-5 reviews) | 25% | 30% | 40% |
| Price stratification | 15% | 15% | 25% |
| Brand concentration | 15% | 20% | 30% |
| Listing surface quality | 5% | 5% | 5% |
High confidence means at least 80% of the listings publish the demand figure. Medium means between 50% and 79%, which is the most common case on the US storefront. Low means fewer than half do, and in that case demand is removed from the score entirely rather than estimated. Its weight is redistributed to the structural signals, which can be read from any page. An honest zero beats an invented number.
The result is an integer from 0 to 100. It is shown on the gauge under the verdict, and it is the same number the verdict is derived from. Nothing else feeds the gauge.
What opportunity scores hide is worth reading before you lean on any single number, this one included.
Where Launch, Watch and Skip begin
The three verdicts are bands on that score, and the boundaries are fixed and published here so they cannot quietly move.
| Opportunity score | Verdict | What it means |
|---|---|---|
| 65 to 100 | Launch | The incumbents are beatable on the signals that matter and demand is present or not measurable against you. |
| 40 to 64 | Watch | A real market with a real obstacle: deep reviews, a dominant brand, or a tight price cluster. Worth tracking, not funding yet. |
| 0 to 39 | Skip | The page is held by established sellers with little price or brand room. Entering means buying your way in. |
Watch is not a hedge. It is the most common verdict on genuinely interesting keywords, and it exists because most Amazon markets are neither open nor closed. They have one specific problem. The written reasoning names it, so a Watch tells you what would have to change, and the watchlist re-scans the keyword so you find out when it does.
The boundaries themselves are tested directly in the codebase. A product scoring 65 is Launch and one scoring 64 is Watch, every time, and an off-by-one there would move every borderline keyword into the wrong band silently. That is why the numbers are pinned rather than tuned by hand.
What the AI adds, and what it is not allowed to do
The score and the competitor map are deterministic arithmetic. The written explanation is produced by a language model, and it is worth being exact about the division of labour, because the phrase AI-powered is doing a lot of vague work across this industry.
The model receives the same twenty rows, the computed signals, the score and the verdict, together with one thing the arithmetic does not have: your selling model. You choose one of eight during onboarding, private label, wholesale as an authorised reseller, arbitrage, dropshipping, handmade, bundles, books and used goods, or multi-model. The explanation is written for that seller, using the storefront's currency and its fee bands, and it names the specific listings it considers beatable and why.
The model may override the verdict, and this is the part to understand. A keyword that is a protected brand name is structurally a fine market and arithmetically a Launch. For a private-label seller it is a Skip, because they cannot list against the trademark. For an authorised wholesaler it stays a Launch. The model knows this and the formula does not, so the model is allowed to move the verdict for reasons of that kind.
- The override is clamped. When the verdict moves, the score is pulled into the matching band, so a Launch badge never sits over a score of 30 and a Skip never sits over an 80. The gauge and the badge always agree, and the override is visible as such.
- The vocabulary is constrained. The model is instructed to speak with full conviction on a coherent page and is permitted to name the mismatch on a mixed one. It is never permitted to hedge with phrases like low confidence, because that phrase pushes the decision back onto you without saying why.
- Your keyword is sandboxed. The search term is wrapped as untrusted input before it reaches the model, so a keyword crafted to manipulate the explanation cannot.
Thirteen storefronts, one calibration
SellRadar scans thirteen Amazon storefronts: the United States, United Kingdom, Germany, France, Italy, Spain, Poland, Canada, Australia, Japan, India, Mexico and Brazil. The signals are the same everywhere, but the raw numbers are not comparable across them, and treating them as if they were would make every smaller marketplace look like an open field.
A competitive Japanese listing may carry a fifth to a seventh of the review count its American equivalent would, simply because the marketplace is smaller. So review counts and demand figures are converted to US-equivalent units per storefront before they hit the buckets. A listing in Japan with 800 reviews is treated like a US listing with roughly 4,000, and crosses the same thresholds. Title heuristics are adjusted for script, and the explanation's fee bands and currency follow the storefront.
The international research guide covers how to use that in practice. The point here is narrower: the same keyword on two storefronts can honestly receive two different verdicts, and neither is a calibration error.
How to read a verdict, and how to disagree with it
Because every input is a visible listing and every weight is published above, the productive way to use a result is to argue with it. Here is the order that works.
- Start with the competitor map, not the badge. Twenty tagged listings tell you where the soft positions are. If the tags match what you see when you open the listings, the rest of the result is standing on solid ground.
- Check which confidence tier applied. If demand was dropped because the storefront publishes it sparsely, the verdict is about structure alone, and you should bring your own demand evidence before acting on a Launch.
- Read the reasoning for your model specifically. If it names a listing as beatable and you know you could not produce a visibly better version within budget, that is a real disagreement. Treat the verdict as one notch more cautious.
- Look for an override. If the badge and the arithmetic disagree, the explanation says why. Trademark, gating and category restrictions are the usual reasons, and they are usually right.
- Re-scan on a Watch rather than deciding once. Markets move. The obstacle named in a Watch is the thing to check again in a month, and the watchlist does that for you.
A method you can inspect is only useful if you actually inspect it. The validation checklist is the manual version of the same steps, and running both on the same keyword is the fastest way to learn where the tool's judgement and yours differ.
Frequently asked questions
What is the SellRadar opportunity score?
An integer from 0 to 100 computed from the top twenty organic Amazon listings for a keyword. It combines five components with published weights: demand from bought-in-past-month data, competition depth from the review counts of the top five, price stratification, brand concentration and listing surface quality. Scores of 65 and above are Launch, 40 to 64 are Watch, and below 40 are Skip.
Is the SellRadar verdict decided by AI?
The score and the competitor map are deterministic arithmetic over the twenty listings. A language model writes the explanation for your selling model and may override the verdict for reasons the formula cannot see, such as a trademarked keyword being a Skip for private label but a Launch for an authorised reseller. When it does, the score is clamped into the matching band so the badge and the gauge always agree.
Why did SellRadar refuse to score my keyword?
The twenty results failed the coherence gate, meaning Amazon returned a mix of unrelated products rather than one comparable market. Scoring such a page would produce a misleadingly high number, because unrelated products maximise exactly the signals that look like opportunity. The scan is not charged. Try a keyword that names a real product category.
What does a Watch verdict mean?
A real market with one specific obstacle: usually deep review moats, a dominant brand or a tight price cluster. It scores between 40 and 64. The written reasoning names the obstacle, and adding the keyword to your watchlist re-scans it so you learn when the obstacle changes. Watch is the most common verdict on interesting keywords because most Amazon markets are neither open nor closed.
Does the scoring work the same on non-US Amazon storefronts?
The signals and thresholds are identical across the thirteen supported storefronts, but the raw inputs are normalised first. Review counts and demand figures on smaller marketplaces such as Japan, India, Mexico and Brazil run several times lower than the US equivalent for the same competitive position, so they are converted to US-equivalent units before bucketing. Title heuristics are also adjusted for script.
Why does SellRadar drop demand from the score sometimes?
Because Amazon does not publish the bought-in-past-month figure on every listing. When fewer than half of the twenty carry it, a demand average would be computed from too few rows to mean anything, so its weight goes to zero and is redistributed to the structural signals. The result tells you which confidence tier applied so you know the verdict was about structure alone.