Amazon reviews

What a review is made of

Helpful votes, length, photos, and duplicate text — plus the verified-purchase joint, which reverses the sign of the correlation the overview page reports at category level.

Review-level aggregates · 507.7M reviews

8 minUpdated August 2026

The reversal

Verified reviewers are the kinder ones

Compare categories and the least-verified ones look the most generous. Compare individual reviews and it flips: verified purchases average 4.191★ across 456.5M reviews, unverified ones 4.137★ across 51.1M. Verified rates higher in 29 of 33 categories. This is Simpson’s paradox, and the overview page’s category-level chart is the trap.

Rating distribution by verified-purchase flag

Counting the pairs directly. Reporting each margin separately — the rating split and the verified split — cannot produce this table.

1★2★3★4★5★

Verified minus unverified mean rating, by category

Positive means verified purchases rate higher in that category.

4 categories buck it — Beauty & Personal Care, Kindle Store, Health & Personal Care, All Beauty — and they are the media and beauty ones, where an unverified review is often a considered opinion rather than a complaint.

Why the sign flips. Media categories are simultaneously the least-verified and the best-rated, for unrelated reasons: books and music attract enthusiasts, and their reviews frequently predate or bypass an Amazon purchase. Aggregating to the category level lets that composition drive the correlation, while the within-category comparison recovers the actual behaviour. Both charts are correct; only one of them is about reviewers.

Helpfulness

Negative reviews get read

Helpful votes are the only signal in this corpus of what other shoppers valued. They are brutally skewed — the median review of any rating gets zero — so the story is in the upper tail. At the 99th percentile a one-star review collects 22.6 votes against 13.5 for a five-star one. Complaints travel 1.68× further.

Helpful votes by star rating

90th and 99th percentile vote counts. Medians are omitted because they are zero at every rating.

Helpful-vote tail by category

99th percentile votes, with the share of reviews receiving none.

Effort

The longest reviews are the ambivalent ones

Length against rating is an inverted U, not a line. A 4-star review runs 53 words on average; five stars takes 33 and one star 40. Unqualified praise is quick. Explaining a mixed verdict takes work — and so, to a lesser extent, does justifying outright condemnation.

Review length and emphasis by rating

Mean words, plus exclamation marks and ALL-CAPS words per review — the only text features carried through the extract.

Exclamation marks and shouting by rating

Per review. Emphasis is U-shaped where length is inverted-U — the extremes shout, the middle explains.

Photos & duplicates

Reviews with photos rate lower

Only 5.41% of reviews carry a photo, and they are not the happy ones: 3.997★ with an image against 4.196★ without. People reach for the camera to document a problem. Separately, exact-duplicate review text peaked at 17.8% of reviews in 2015 and sits at 9.0% by 2023.

Rating with and without a photo

Pooled across all categories and years.

Exact-duplicate review text by year

Share of reviews whose text appears verbatim on another review.

The duplicate rate is both a floor and a ceiling. It is a lower bound on coordinated review activity, because paraphrased and AI-rewritten duplicates are invisible to an exact-text hash. It is simultaneously an upper bound on misconduct, because legitimate duplicates are everywhere: the same reviewer posting one verdict across product variants, and boilerplate like “Good.” colliding by chance across millions of reviews. Treat the trend as more informative than the level.

Photo share by category

Share of reviews carrying at least one image.