Working dataset · 2 analyses

Half a billion Amazon reviews

507,730,787 reviews across 33 product categories, from June 17, 1996 to September 14, 2023, reduced to counts, means, and distributions. No review text, no user IDs, no product IDs — which makes it a corpus you can hand to a class on day one and still learn something real from.

reviews
507.7M
categories
33
years covered
28
mean rating
4.19★
five-star
65%
verified
90%

The distribution

A mean of 4.19 describes almost nothing

Star ratings are not bell-shaped. Across all 507.7M reviews, 75.6% of the mass sits at the two ends of the scale and only 24.4% sits in the middle three. The distribution is J-shaped: 5★ is the mode, 1★ is the runner-up, and the arithmetic mean lands where almost nobody actually rates.

Share of all reviews by star rating

507,730,787 reviews, 33 categories, June 17, 1996 – September 14, 2023.

  1. 565.5%
  2. 412.6%
  3. 37.0%
  4. 24.9%
  5. 110.1%
This is the single most consequential fact about review data. Any model, dashboard, or pricing rule treating “average rating” as a location parameter is summarising a bimodal distribution with a number that describes neither mode. The useful summaries are share-at-5★, share-at-1★, and the ratio between them.

The categories

Three categories are a third of the corpus

Home & Kitchen, Clothing Shoes & Jewelry, Electronics together hold 34.9% of every review ever written. Subscription Boxes, the smallest slice, holds 16,216 4,157× fewer than Home & Kitchen. Compare categories on rates, never on raw counts.

All 33 categories

Sort by any column.

  1. 1Home & Kitchen67.4M
  2. 2Clothing Shoes & Jewelry66M
  3. 3Electronics43.9M
  4. 4Books29.5M
  5. 5Tools & Home Improvement27M
  6. 6Health & Household25.6M
  7. 7Kindle Store25.6M
  8. 8Beauty & Personal Care23.9M
  9. 9Cell Phones & Accessories20.8M
  10. 10Automotive20M
  11. 11Sports & Outdoors19.6M
  12. 12Movies & TV17.3M
  13. 13Pet Supplies16.8M
  14. 14Patio Lawn & Garden16.5M
  15. 15Toys & Games16.3M
  16. 16Grocery & Gourmet Food14.3M
  17. 17Office Products12.8M
  18. 18Arts Crafts & Sewing9M
  19. 19Baby Products6M
  20. 20Industrial & Scientific5.2M
  21. 21Software4.9M
  22. 22CDs & Vinyl4.8M
  23. 23Video Games4.6M
  24. 24Musical Instruments3M
  25. 25Amazon Fashion2.5M
  26. 26Appliances2.1M
  27. 27All Beauty702K
  28. 28Handmade Products664K
  29. 29Health & Personal Care494K
  30. 30Gift Cards152K
  31. 31Digital Music130K
  32. 32Magazine Subscriptions71K
  33. 33Subscription Boxes16K

Verification

The least-verified categories are the best-rated

89.9% of reviews carry a verified-purchase flag — but that share collapses in media. Books, Kindle, CDs & Vinyl, Digital Music, and Movies & TV average 71.2% verified against 93.3% everywhere else, and they rate 4.39★ against 4.15★. People review books they did not buy on Amazon, and they are kinder when they do.

Verified-purchase share against mean rating

One dot per category, sized by review volume. Media categories in amber.

Media (book, music, video)Everything else
Treat the verified flag as a sampling variable, not a quality filter. Restricting to verified reviews does not just remove noise — it removes Books, Kindle, and CDs from your sample far more aggressively than it removes Automotive, and it shifts the rating distribution while it does so.

Analyses

Questions asked of this corpus

Each analysis states what slice it is computed over. The aggregate pages cover all 507.7M reviews; others work from smaller samples where the review text itself is needed.

The data

Five CSVs, no text, no identifiers

The published aggregates are counts, means, and distributions only — no review text, no user ID, no product ID. That is what makes them safe to hand out and what makes them useless for per-product or NLP work; for that you need the HuggingFace source.

FileRowsWhat it holds
category_stats_all.csv33One row per category — volume, mean rating, mean length, verified share, the 1★–5★ split, first and last review date.
ts_yearly_all.csv798Category × year, 1996–2023. The only chronological file.
ts_monthly_all.csv396Category × calendar month. Seasonality, all years pooled.
ts_dayofweek_all.csv231Category × weekday (0 = Monday), all years pooled.
ts_hourofday_all.csv792Category × hour (0–23), all years pooled.

Plain HTTPS — no credentials

import pandas as pd

BASE = "https://ontopic-public-data.t3.storage.dev/amazon-reviews/merged_results/"
cats = pd.read_csv(BASE + "category_stats_all.csv")
yrs  = pd.read_csv(BASE + "ts_yearly_all.csv")

S3 protocol

import boto3, pandas as pd

s3 = boto3.client("s3", endpoint_url="https://t3.storage.dev",
                  region_name="auto")
obj = s3.get_object(Bucket="ontopic-public-data",
                    Key="amazon-reviews/merged_results/"
                        "category_stats_all.csv")
cats = pd.read_csv(obj["Body"])

The bucket answers anonymous GETs on virtual-host style URLs (bucket.t3.storage.dev/key); the path-style form t3.storage.dev/bucket/key returns 403.

Four ways to get this wrong

  1. Only ts_yearly is a timeline. The monthly, weekday, and hour files pool every year together. Plotting them left to right as a time axis produces a chart that means nothing.

  2. Filter on count before trusting a rate. A category-year holding one review reports rating_5_pct = 100.0. Every rate chart here drops cells under 500 reviews.

  3. Volumes span four orders of magnitude. 67.4M reviews in Home & Kitchen against 16,216 in Subscription Boxes. Normalise before you compare.

  4. Percent columns are 0–100. Not 0–1. Dividing twice, or not at all, is the most common bug against these files.

Derived from McAuley-Lab/Amazon-Reviews-2023 and inherits its terms. Aggregation ran on Google Cloud Run, one job per category, streaming each raw_review_* split; the merged CSVs were migrated to Tigris in August 2026 with every object verified by MD5. Charts on these pages read a 33-category JSON built from those CSVs by scripts/fetch-amazon-aggregates.mjs.