Amazon reviews

The rater in the rating

Summarise reviews by category and the reviewer disappears. Group them by author instead, and the star scale turns out to measure two things at once — the smaller of which is the product.

Reviewer-level aggregates · 507.7M reviews

10 minUpdated August 2026

Selection

One-time reviewers give one star 2.5× as often

Sort every review by how many its author ever wrote, and the rating climbs monotonically: 3.81★ from people who reviewed once, 4.26★ from people who reviewed ten times or more. The one-star share falls from 20.1% to 8.2% — a factor of 2.5. Nothing about the products changed; only who is holding the pen.

Mean rating by reviewer activity

Every one of the 507.7M reviews, bucketed by its author's lifetime review count.

Full rating distribution by reviewer activity

The gradient is not a shift in the mean — it is the one-star block shrinking.

1★2★3★4★5★

This is why the corpus mean is an upper bound. The Unknown category is excluded from every published figure — 11% of ratings, but 42% of users, and disproportionately one-and-done. Since sparse reviewers rate lower, dropping them pushes every mean rating up. The direction was previously an assertion; this chart is what quantifies it.

Concentration

Half the reviews come from a tenth of the reviewers

54.1M reviewers wrote 507.7M reviews — 9.39 each on average, which describes almost nobody. The top 10% wrote 52.6% of everything and the top 1% wrote 16.9%, for a Gini of 0.644. That is more unequal than most countries’ income distributions.

Lorenz curve — reviewers against reviews

Cumulative share of reviews held by the least-active reviewers, ordered by activity. The diagonal is perfect equality.

Reviewer inequality by category

Gini of reviews-per-reviewer, within each category.

Dashed line is the corpus-wide Gini (0.644). Per-category Ginis count only a reviewer’s reviews inside that category, so they are not comparable to the global figure — a generalist looks like a novice in every category they touch.

Experience

Reviewers get kinder, not harsher — probably

Conventional wisdom says critics harden with practice. The curve says the opposite: a reviewer’s first review averages 3.992★ and their tenth 4.202★. But read the caveat below before believing it — this particular chart has survivorship built into its x-axis.

Mean rating at a reviewer's nth review

Pooled across all reviewers who reached that many reviews.

The x-axis is not a clean treatment. Everyone appears at review #1; only people who kept going appear at #10. Since the previous section showed that prolific reviewers rate higher as a population, most of this rise is composition — the harsh one-and-done reviewers dropping out of the sample — not individuals mellowing.

Separating the two needs a within-reviewer comparison — the same person’s first review against their tenth. These aggregates cannot do it, because they never follow an individual over time. Read the curve as an upper bound on any real learning effect, and treat the gap between it and zero as the size of the selection problem.

Tenure

A quarter of reviewers exist for a single day

26.8% of reviewers have a first and last review on the same date — they arrived, said something, and never came back. At the other end, reviewers whose span exceeds five years are 27.6% of the population and 60.2% of the reviews.

Reviewers and reviews by tenure span

Tenure is last review minus first review. Buckets are lower bounds.

The two bars diverge sharply, and that divergence is the whole point: the reviewer population is dominated by people who barely participate, while the review corpus is dominated by people who never stopped. Any statement about “what reviewers think” has to pick which of those two populations it means.

Variance

The rater explains more than the product — with a large asterisk

Taken one at a time, knowing who wrote a review accounts for 30.2% of the variance in star ratings; knowing which product it is about accounts for 19.3%; knowing the category accounts for 0.7%. The ordering is the finding. The percentages are not a decomposition, and cannot be read as one.

Marginal variance explained, by factor

Each bar is that factor alone. They overlap heavily and sum to 50.2%, which partitions nothing.

Not a variance decomposition. User and item effects are crossed and unbalanced, so their sums of squares share variance rather than partitioning it. On complete Gift Cards data the same three marginals sum to 105.62% — impossible for an orthogonal decomposition. Do not compute a residual from these numbers, and do not read them as “30.2% is the rater, 19.3% is the product.” A further caveat travels with the item figure: 42.6% of items have exactly one review, and a group of one explains its own variance by construction.

This measure was originally specified as a per-category variance decomposition, to be read as an index of taste-driven versus quality-driven markets. It does not survive contact with the data, and the pipeline was right to publish marginals with a warning instead of a partition that does not exist. What remains is still worth knowing — the rater term is larger than the product term, consistently — but a genuine decomposition needs a crossed random-effects model, not sums of squares.