What a review is made of
Helpful votes, length, photos, and duplicate text — plus the verified-purchase joint, which reverses the sign of the correlation the overview page reports at category level.
Review-level aggregates · 507.7M reviews
8 minUpdated August 2026The reversal
Verified reviewers are the kinder ones
Compare categories and the least-verified ones look the most generous. Compare individual reviews and it flips: verified purchases average 4.191★ across 456.5M reviews, unverified ones 4.137★ across 51.1M. Verified rates higher in 29 of 33 categories. This is Simpson’s paradox, and the overview page’s category-level chart is the trap.
Rating distribution by verified-purchase flag
Counting the pairs directly. Reporting each margin separately — the rating split and the verified split — cannot produce this table.
Verified minus unverified mean rating, by category
Positive means verified purchases rate higher in that category.
4 categories buck it — Beauty & Personal Care, Kindle Store, Health & Personal Care, All Beauty — and they are the media and beauty ones, where an unverified review is often a considered opinion rather than a complaint.
Why the sign flips. Media categories are simultaneously the least-verified and the best-rated, for unrelated reasons: books and music attract enthusiasts, and their reviews frequently predate or bypass an Amazon purchase. Aggregating to the category level lets that composition drive the correlation, while the within-category comparison recovers the actual behaviour. Both charts are correct; only one of them is about reviewers.
Helpfulness
Negative reviews get read
Helpful votes are the only signal in this corpus of what other shoppers valued. They are brutally skewed — the median review of any rating gets zero — so the story is in the upper tail. At the 99th percentile a one-star review collects 22.6 votes against 13.5 for a five-star one. Complaints travel 1.68× further.
Helpful votes by star rating
90th and 99th percentile vote counts. Medians are omitted because they are zero at every rating.
Helpful-vote tail by category
99th percentile votes, with the share of reviews receiving none.
Effort
The longest reviews are the ambivalent ones
Length against rating is an inverted U, not a line. A 4-star review runs 53 words on average; five stars takes 33 and one star 40. Unqualified praise is quick. Explaining a mixed verdict takes work — and so, to a lesser extent, does justifying outright condemnation.
Review length and emphasis by rating
Mean words, plus exclamation marks and ALL-CAPS words per review — the only text features carried through the extract.
Exclamation marks and shouting by rating
Per review. Emphasis is U-shaped where length is inverted-U — the extremes shout, the middle explains.
Photos & duplicates
Reviews with photos rate lower
Only 5.41% of reviews carry a photo, and they are not the happy ones: 3.997★ with an image against 4.196★ without. People reach for the camera to document a problem. Separately, exact-duplicate review text peaked at 17.8% of reviews in 2015 and sits at 9.0% by 2023.
Rating with and without a photo
Pooled across all categories and years.
Exact-duplicate review text by year
Share of reviews whose text appears verbatim on another review.
The duplicate rate is both a floor and a ceiling. It is a lower bound on coordinated review activity, because paraphrased and AI-rewritten duplicates are invisible to an exact-text hash. It is simultaneously an upper bound on misconduct, because legitimate duplicates are everywhere: the same reviewer posting one verdict across product variants, and boilerplate like “Good.” colliding by chance across millions of reviews. Treat the trend as more informative than the level.
Photo share by category
Share of reviews carrying at least one image.