What’s on the shelf
Prices, brand concentration, and the attribute vocabulary of 35 million products — and the three measures this corpus is structurally unable to answer.
Product metadata · 35M items
7 minUpdated August 2026Prices
Price is missing more often than it is present
Before any price chart: the median category is 60.3% null. Coverage runs from 14.6% missing in Kindle Store to 100.0% in Subscription Boxes. Every decile below is computed on the minority of items that carry a price, and there is no reason to think that minority is a random sample of the shelf.
Price missingness by category
Share of items with no price in the metadata.
Price distribution where price exists
10th to 90th percentile, with the median marked. Log scale.
Concentration
Measured by reviews, most categories are unconcentrated
HHI over brands, computed two ways: by share of items listed and by share of reviews received. Gift Cards is the outlier at 5,205.3 on reviews — a single seller dominating what is effectively Amazon’s own product. Kindle Store sits at 4.1. For scale, US antitrust guidelines call a market “highly concentrated” above 2,500, so all but one of these markets are formally competitive.
Brand HHI by category, items versus reviews
Log scale — the range spans three orders of magnitude. Higher means more concentrated.
Everything sits above the diagonal, meaning attention is more concentrated than the catalogue is: the brands with the most listings are not merely proportionally reviewed, they absorb more than their share. That gap is the interesting quantity, not either HHI alone.
Structure
Two thousand ways to describe a product
The details field is a free-form dictionary, and across 32 categories it contains a vocabulary of over two thousand distinct keys. The top thirty carry most of the coverage; the tail is where any attempt at structured product comparison goes to die.
Most common product-detail keys
Items carrying each key, summed across categories.
Category-tree depth
Mean breadcrumb depth, for the 29 of 33 categories that populate the field at all.
4 categories report 100% empty breadcrumbs and are omitted rather than drawn as zero — an absent field and a depth of zero are different facts.
Negative results
Three things this dataset cannot tell you
Publishing only the measures that worked would misrepresent the corpus. These three were specified, attempted, and abandoned — each for a different and instructive reason.
The co-purchase graph does not exist
bought_togetheris null for every item in all 33 categories — 0 have any coverage at all. The field is documented in the dataset schema, which makes it look available; it is empty in practice. This was the one interaction signal that needed no reviewer-level shuffle, and it is simply not there.“Items that never got reviewed” is unanswerable
Exactly zero of the 35,003,183 items have a rating count of zero. That is not a fact about Amazon — the metadata split only contains items that appear in the reviews split, so the denominator the measure needs, every listed product whether reviewed or not, is absent by construction. Publishing “0%” would have stated an artifact as a finding.
Page rating versus computed rating — deferred, not abandoned
Comparing the rating Amazon displays against the mean of the reviews actually present would measure how much of a star rating comes from ratings-without-reviews. It needed a second full pass over the 245 GB review corpus during the metadata scan, roughly doubling that tier’s cost for one number. The by-item extract now exists, so it is cheap the next time round.
Why this section exists. A dataset page that lists only its successes teaches the wrong lesson about data work. Two of these three are properties of how the corpus was assembled rather than of Amazon, and telling them apart is most of the skill.