SellRadar

Validation

Amazon opportunity scores: useful shorthand, dangerous conclusion

Every research tool has a version of it — one number that tells you how good an opportunity is. The number is not wrong. It is compressed, and what gets squeezed out is usually the part that decides your launch.

Updated September 11, 2026 · Written by the SellRadar team

What an opportunity score actually is

Opportunity scores differ by vendor but they combine broadly the same ingredients: some measure of demand, some measure of competition, usually something about listing quality and price, folded into a single figure.

As shorthand this is genuinely useful. Screening two hundred candidates down to twenty is exactly the job a score is good at, and doing it by hand is not realistic. The problem is not the score. The problem is treating the output of a triage tool as the input to a decision.

A score is a compression of a market. Compression is lossy by definition — the only question is whether what it discarded mattered.

What the compression discards

The shape of the competition

The most important loss. A score reduces twenty competitors to one competition value, and it cannot distinguish twenty evenly-matched sellers from a leader plus eighteen weak listings. Those markets score similarly and demand opposite decisions — one is closed, the other has a soft middle you can take.

Who you are

Scores are computed the same for everyone. But whether a market is an opportunity depends entirely on the seller. A page of generic no-name listings is an opening for someone building a brand and a margin problem for a reseller. A category requiring domain knowledge is easy for someone who has it. The score cannot know any of that, so it answers a question about the market when your question is about you.

What you could actually execute

The decisive input is how many of the twenty listings you could produce a visibly better version of within your real budget. That is a judgement about your capability against these specific competitors. No score has access to it.

The reason

A score of 6 out of 10 tells you nothing about why. You cannot check its reasoning, cannot disagree with a specific claim, and cannot tell whether the 6 came from weak demand and easy competition or strong demand and hard competition — which are entirely different situations that a single number renders identical.

When a score is the right tool

None of this means scores are useless. They are the right tool for a specific job.

  • Triage at volume. Two hundred candidates to twenty. Nothing else does this efficiently and manual assessment at that scale is not viable.
  • Relative ranking within one search. Comparing candidates surfaced by the same query, with the same method applied consistently, is a fair comparison.
  • A sanity check. If your own read is wildly different from the score, one of you has missed something. Worth investigating which.
  • Learning what to weigh. Watching how scores respond as inputs change is a reasonable way to build intuition when you are new.

What they are not right for is the final call on a product you are about to fund. Screening and deciding are different jobs, and using a screening instrument to decide is where the trouble starts.

What to look at instead for the final call

Once a score has narrowed your list, the decision needs the things the score discarded.

Instead ofLook atBecause
A competition scoreThe review distribution shapeFlat, deep, steep and barbell mean different strategies
An average review countWhere counts clusterThe average erases the entire signal
A listing-quality metricHow many you could visibly beatYour capability is the actual variable
A demand figureReview recency across the pageRecency is live demand; totals are history
A price data pointThe spread across all 20Spread tells you if positioning room exists
An overall opportunity numberReasoning you can disagree withYou cannot argue with a 6

That right-hand column is roughly thirty minutes of manual work per product — which is exactly why scores exist and exactly why they should not be the last word.

Why we ship a verdict rather than a score

SellRadar returns LAUNCH, WATCH or SKIP with written reasoning rather than a number out of ten. That is a deliberate choice and it has a real trade-off, so here is the honest version of both sides.

What a verdict gives you

  • Something you can disagree with. The reasoning names which listings look beatable and why. If you think it is wrong about position seven, you know exactly which claim to challenge.
  • Framing for how you sell. The same twenty listings produce different verdicts across eight selling models, because a page of generic brands genuinely means different things to a brand builder and a reseller.
  • The individual competitors. The listing weakness map keeps the twenty as twenty rather than averaging them into one value.

What a score gives you that a verdict does not

  • Sortability. You cannot rank forty products by LAUNCH/WATCH/SKIP as cleanly as by a number. Three buckets are coarser than a scale.
  • Speed at very high volume. For screening hundreds, a score is the better instrument and we would not pretend otherwise.

So the honest recommendation is to use both, for the jobs each is good at: score to narrow a long list, reasoning to decide on the survivors. Three SellRadar verdicts a month are free with no card, which is about the right volume for the deciding step anyway.

If you want a score, build your own

There is a middle path between trusting a vendor's number and doing thirty minutes of qualitative reading per product: score the things you have decided matter, with weights you chose.

The advantage is not accuracy. It is that you know what went into it, so you can tell when it is wrong and adjust it as you learn.

A workable starting point, scored one to five per product from the top 20 organic listings:

  • Beatable listings — how many of the twenty you could visibly out-execute within budget. Weight this heaviest; it is the most predictive single input.
  • Review distribution shape — flat and low scores high, uniformly deep scores low.
  • Generic seller share — more generic holders means more undefended positions.
  • Price spread — wide scores high, tight cluster scores low.
  • Margin headroom at the median price after fees, landed cost and advertising.
  • Unmet need — is there a recurring complaint you could fix? Present or absent.
  • Your own knowledge of the customer. Genuinely predictive and never in any vendor's model.

Sum it, and record the score alongside what actually happened when you launched. After a handful of products you will know which inputs predicted outcomes for you specifically, and you can reweight accordingly. No vendor can do that, because no vendor knows what you can execute.

The last item is the one that makes this worth doing. Every commercial score is computed identically for every user, which means it cannot represent the single largest variable in whether a market is an opportunity: who is entering it.

Frequently asked questions

Are Amazon opportunity scores accurate?

They are internally consistent rather than accurate or inaccurate — they compute what they define. The issue is not the arithmetic but the compression: a single number cannot represent how competition is distributed across twenty listings, and it cannot account for who you are or what you could execute. Useful for triage, insufficient for a funding decision.

Should I trust a high opportunity score?

Trust it enough to look closer, not enough to order inventory. A high score means the aggregate inputs look favourable. It does not mean the twenty sellers holding page one include any you could beat, which is the question that decides the outcome. Always look at the listings individually before committing.

What is a good opportunity score?

The threshold varies by vendor and is not comparable across tools, since each defines its own inputs and weighting. More usefully: rather than looking for a score above some cutoff, use scores to rank candidates within a single search and then assess the top few properly by reading their competitive pages.

Why do two products with the same score have different outcomes?

Because the score averaged away the difference. Twenty evenly-matched competitors and a strong leader plus eighteen weak listings can produce identical aggregates, but the first page is closed and the second has a contestable middle. Only looking at the distribution rather than the average distinguishes them.

Is a verdict better than a score?

For deciding, yes, because reasoning can be checked and argued with while a number cannot. For screening a long list, a score is better — it sorts, and three buckets do not. They serve different stages, and using each for its own job beats insisting on one.