Amazon reviews

How a product’s rating forms

If later reviewers copied earlier ones, it would show up as drift in a product’s rating as reviews accumulate. It does not. But what the very first review said predicts the next hundred.

Product-level aggregates · 35M items

9 minUpdated August 2026

Herding

A product's rating barely moves as reviews pile up

If later reviewers anchored on earlier ones, the mean rating at an item’s nth review would drift. It does not. Across all 236M reviews the curve runs from 4.178★ at the first review to a peak of 4.197★ at #4 and 4.183★ by #50 — a total span of 0.019stars. The first review is the outlier, and it is slightly harsher than what follows.

Mean rating at an item's nth review

Pooled across all 33 categories. Note the y-axis spans 0.04 stars — this is a flat line, drawn honestly.

Drift per review, by category

Slope of mean rating against review index over positions 2–20. Negative means later reviews run cooler.

Even the extremes are tiny — a hundredth of a star per review at most. Whatever social influence exists between Amazon reviewers, it does not show up as drift in the mean.

The finding

The first review predicts everything after it

Condition on what an item’s first reviewer said, then look only at reviews two onward. Items that opened with 1★ average 3.748★ thereafter; items that opened with 5★ average 4.261★. That is a gap of 0.513 stars that persists across every subsequent review — and unlike the flat curve above, it is enormous.

Mean of reviews 2–n, by what the first review said

Pooled across all categories. Each bar conditions on the first rating only.

This is not proof of herding. An item whose first review was one star is probably a worse item — the first review is measuring quality, not creating it. Separating the two needs something these aggregates cannot supply: variation in the first review that is unrelated to the product, such as review timing or reviewer identity.

What the data can do is show the effect is not a small-category artifact. Among the 18 categories with more than 5M conditioned reviews the gap ranges from 0.40 to 0.66 stars — consistent everywhere, never absent.

First-review gap by category

Mean of reviews 2–n after a 5★ opener, minus the same after a 1★ opener. Experience goods in amber.

Dashed lines are the unweighted means for each group: 0.493 for experience goods, 0.489 for search goods. Dot size is the number of conditioned reviews — the extremes at the top of the chart are the smallest categories, so read the ordering with that in mind.

I predicted before seeing this that herding would be strongest where quality is hard to judge before buying — books and beauty over tools and electronics. The split is in that direction but small (0.493 versus 0.489) and the within-group spread swamps it. On this evidence, Nelson’s search/experience distinction does not organise the first-review effect. Category size predicts the gap better than category type does, which is usually a sign you are looking at estimator noise rather than a behavioural difference.

Disagreement

Two ways to be a three-star product

A mean rating hides whether raters agreed. Of the 1.1M items averaging between 2.5 and 3.5 stars, 52.5% have a standard deviation above 1.8 — those are not mediocre products, they are contested ones, loved and hated in roughly equal measure. Plotting mean against spread separates the two populations that a single number merges.

Items by mean rating and rating spread

9.5M items with at least two reviews, binned. Colour is item count on a log scale.

The bright ridge along the bottom-right is the ordinary case: high mean, low disagreement. The arm reaching up and left is the contested population. The hard diagonal edge on the left is arithmetic, not behaviour — an item averaging 1.2 stars cannot have a spread of 2.

Review concentration across items

Gini of reviews-per-item within each category. Higher means a few products absorb most of the attention.

Lifecycle

A product's reviews arrive early or never

Measured in weeks since an item’s first review, 15.6% of all reviews land in week zero and 22.4% within the first month. By the end of the first year the item has collected 67.9% of the reviews it will ever get in this window. Attention decays fast, and it does not come back.

Reviews by weeks since the item's first review

First two years, all categories pooled. Log y-axis — the decay spans three orders of magnitude.

Week zero is a spike rather than a point on the curve: it contains every review posted in the same week as the item’s first, which for many items is the only week they ever get reviewed.