Mountain Dew outsold Coca-Cola in 41% of the 2,542 ZIP codes in this case, and by at least 1.9 to 1 in one ZIP in ten. Newport was more than a third of cigarette units in 7.5% of ZIPs (about one in thirteen), and ZYN nicotine pouches were a quarter or more of smokeless units in 12.3% (about one in eight). Every store sells the same six categories: beer, soft drinks, energy drinks, cigarettes, smokeless tobacco and cigars. Which brands sell within each category depends on the ZIP code.
A distributor planning assortments or a retailer scouting a site faces 40 brand shares per ZIP, too many to read, map or compare across 2,542 ZIPs. This case asks whether a few underlying dimensions summarize them, whether ZIP codes fall into segments, and what the shares reveal about the people who live there.
Forty brand shares for 2,542 ZIP codes
A brand's share is its units divided by all units in its category, so the shares within a category add to 100. The file covers calendar year 2022 and pools at least two convenience stores in each ZIP code.Data. Unit sales from PDI point-of-sale systems, obtained through Dewey. Demographics are ACS-based ZIP tables; adult smoking rates are CDC PLACES model estimates; the 2020 presidential vote is allocated to ZIP codes by Fekrazad.1 The Northeast is thin, and Pennsylvania and New Jersey contribute little beer data because convenience stores there rarely sell beer.
| Item | Contents |
|---|---|
| Unit of analysis | ZIP code, 2,542 ZIPs in 44 states |
| Stores | 7,626 convenience stores, 2 to 21 per ZIP (median 2) |
| Brand shares | 40 brands: 16 beer, 7 soft drinks, 1 energy drink (Monster), 6 cigarette, 4 smokeless, 6 cigar |
| Demographics | Race and ethnicity, income, education, age, density, poverty, adult smoking rate, location |
| Outcome used later | Republican share of the 2020 two-party presidential vote, by ZIP |
Brand shares follow regional lines
Figure 1 maps any of the 40 shares. Mountain Dew outsold Coca-Cola in 80% of the ZIPs in the Midwest heartland segment (defined in Figure 5) and in 1% of the ZIPs in the Hispanic metro segment. Brands that sell in the same places carry overlapping information, and the number of independent dimensions behind the 40 columns measures how much.
Source: PDI convenience-store sales (Dewey), calendar 2022; 2,542 ZIP codes, 7,626 stores. Shares are percentages of units within each category; every ZIP counts once (no weights). ZIP positions are Census 2020 ZCTA internal points on an Albers projection.
Standardizing puts every brand on one scale
Marlboro takes 43.1% of cigarette units in the average ZIP and Kool 2.0%, so the two live on different scales. Standardizing subtracts a brand's average across ZIPs and divides by its standard deviation:
Here xij is brand j's share in ZIP i, xj its average across ZIPs and sj its standard deviation. A z of +2 means two standard deviations above the typical ZIP, whatever the brand. Without this step the largest brands (Marlboro, Bud Light, Coca-Cola) have the largest variances and would dominate every component. The 40 columns leave out each category's "other" column: it is an exact function of the listed shares, because the shares in a category add to 100.
Three principal components summarize the 40 shares
Brands move together. Where Busch Light sells well, so do Mountain Dew and Monster; where Corona sells well, so do Modelo and Fanta. Principal component analysis2 replaces the 40 correlated standardized shares with uncorrelated weighted sums, ordered by how much of the variation each keeps. The first is
with the weights aj chosen to make PC1 vary as much as possible across ZIPs. The second component does the same among weighted sums that are uncorrelated with the first, and so on. The variance of component k is called its eigenvalue, λk. The 40 standardized shares have a total variance of 40, so component k keeps
Figure 2 plots the share each component keeps. The first three keep 15.6%, 14.1% and 10.2% of the variance, 39.9% together, and every later component keeps 4.9% or less. The bars flatten after the third, the elbow that the scree test looks for.3 A common alternative, the eigenvalue-above-1 rule, keeps any component that explains more variance than one standardized brand does (above 2.5% of the total).4 It keeps ten components here, which together explain 66.2%. Ten components are more than an analyst can name, and three can be named, as the next section shows. The choice of three is a judgment about interpretability, not the output of a test.
Three components stand out in the scree plot, while the eigenvalue-above-1 rule would keep ten
Source: PDI convenience-store sales (Dewey), 2022; eigendecomposition of the correlation matrix of the 40 standardized brand shares across 2,542 ZIP codes; unweighted.
The components line up with race, region and education
A component's loadings are the correlations between each brand's share and the component score. Brands with large positive loadings define one end of a component, and brands with large negative loadings define the other. The signs are arbitrary: a component and its mirror image carry the same information. The last column below correlates each ZIP's component score with its census characteristics; the names are mine.
| Component | High end | Low end | Correlated with | A name |
|---|---|---|---|---|
| 1 (15.6%) | Fanta, Corona, Sprite, Heineken, Coca-Cola, Newport, Modelo | Mountain Dew, Busch Light, Monster, Twisted Tea, Pepsi | % white −0.69, latitude −0.64, % Black +0.61, % foreign-born +0.53, Republican share −0.53 | Diverse, Southern and metro vs. white rural heartland |
| 2 (14.1%) | White Claw, American Spirit, ZYN, Camel, Backwoods, Mike's Hard | Bud Light, Michelob Ultra, Natural Light, Dr Pepper, Monster | % adults smoking −0.59, % some college +0.53, income +0.46, latitude +0.41, Republican share −0.38 | Educated, affluent West and North vs. Anheuser-Busch country |
| 3 (10.2%) | Dr Pepper, Marlboro, Copenhagen, Swisher Sweets, Coors Light | Newport, Grizzly, Pepsi, Mountain Dew, Game | longitude −0.61 (western ZIPs score high), % Hispanic +0.36, % Black −0.34, Republican share +0.30 | Texas and the West vs. the Southeast |
The components were computed from brand shares alone, and their scores still correlate with race, region, education and politics. Figure 3 plots every brand's loadings on two components at once, so brands that sell in the same ZIPs sit close together.
Brands that sell in the same ZIPs sit close together on the brand map
Source: PDI convenience-store sales (Dewey), 2022; 40 standardized brand shares across 2,542 ZIP codes; demographic correlations are Pearson correlations across ZIPs, unweighted.
On components 1 and 2 the Coca-Cola system (Fanta 0.77, Sprite 0.69, Coca-Cola 0.56) sits with the Mexican imports (Corona 0.73, Modelo 0.54), Heineken (0.62) and the menthol cigarettes Newport (0.55) and Kool (0.50). PepsiCo's Mountain Dew (−0.75) and Pepsi (−0.45) sit with Busch Light (−0.60), Twisted Tea (−0.53) and Monster (−0.56). Bud Light (−0.70), Michelob Ultra (−0.67) and Natural Light (−0.59) anchor the bottom of component 2, opposite White Claw (0.76), American Spirit (0.75) and ZYN (0.73). Two brands that are close sell well in the same ZIPs. That is co-location in store sales, and says nothing about whether consumers see the brands as similar: this is a revealed-preference map, not a perceptual map built from survey ratings.
Figure 4 shows, for each component, the brands with the largest loadings and the correlations that support the names. Component 1 has correlations of −0.69 with percent white and +0.61 with percent Black, component 2 of +0.53 with the share of adults with some college and −0.59 with the adult smoking rate, and component 3 of −0.61 with longitude.
Brand loadings
Correlation with ZIP characteristics
Source: PDI convenience-store sales (Dewey), 2022; ACS-based ZIP demographics; CDC PLACES adult smoking estimates; 2020 vote from Fekrazad (2025). Correlations across 2,542 ZIPs, unweighted. "Latitude" is positive in the north and "longitude" is positive in the east.
K-means groups the ZIPs into six segments
PCA summarizes the variables; clustering groups the observations. The k-means algorithm56 splits the ZIPs into K segments so that each ZIP is as close as possible to its segment's average profile:
Here si holds ZIP i's scores on the first six components (53.5% of the variance), c(i) is its segment and μc is the segment's average score. The algorithm alternates two steps until nothing changes: assign each ZIP to the nearest segment center, then move each center to the average of its ZIPs. With K = 6 the segments below result. The numbering the algorithm assigns is arbitrary, and the names are mine. A segment over-indexes on a brand when its ZIPs' average standardized share is above the all-ZIP average.
| Segment | ZIPs | Over-indexes on | Top states by ZIPs | Profile (averages across ZIPs) |
|---|---|---|---|---|
| Hispanic metro | 308 | Modelo, Mexican sodas, Corona, Coca-Cola, Dutch Masters, Heineken | FL, TX, IL, GA, CA | 25% Hispanic, $82k median household income, 45% Republican |
| Midwest heartland | 736 | Mountain Dew, Busch Light, Monster, Twisted Tea, Pepsi, Pall Mall | OH, IL, MI, WI, MO | 90% white, 5% of ZIPs urban, 62% Republican |
| Carolinas and the Atlantic coast | 412 | Game, Newport, Heineken, Grizzly, Pepsi, Dutch Masters | NC, FL, SC, VA | 25% Black, 49% Republican |
| The West, "new nicotine" | 220 | ZYN, Camel, American Spirit, White Claw, Coors Light, Mike's Hard | CA, WA, ID, OR, UT | 66% some college, $79k, 12% of adults smoke, 49% Republican |
| Deep South, large Black population | 262 | Kool, Sprite, Fanta, Newport, Bud Light, Miller High Life | GA, TX, MS, LA, AL | 43% Black, $55k, 42% Republican |
| Rural red South and Texas | 604 | Dr Pepper, Michelob Ultra, Bud Light, Natural Light, Copenhagen, Marlboro | TX, AL, GA, TN, OK | 82% white, 3% of ZIPs urban, 71% Republican |
The names label the brands a segment over-indexes on and do not describe every ZIP in it. In the Hispanic metro segment 33% of ZIPs are urban and the average ZIP is 25% Hispanic. Figure 5 maps the segments and profiles each one.
Source: PDI convenience-store sales (Dewey), 2022; ACS-based ZIP demographics; CDC PLACES; 2020 vote from Fekrazad (2025). K-means with K = 6 on the first six component scores of 2,542 ZIPs. Segment averages are unweighted averages across ZIPs, not population-weighted.
The segments sit on a continuum, so six is a judgment call
The silhouette score7 measures how well separated segments are. For each ZIP it compares ai, the average distance to the other ZIPs in its own segment, with bi, the average distance to the ZIPs in the nearest other segment:
The score is near 1 for a ZIP deep inside a tight segment, near 0 on a border and negative when the ZIP is closer to another segment. Averaged over all ZIPs it is 0.200 for the six segments here and between 0.187 and 0.215 for every K from 3 to 8 (Figure 6). A common rule of thumb reads averages below 0.25 as no substantial cluster structure.8 ZIPs lie on a continuum from the Midwest heartland basket to the Hispanic metro basket, and k-means draws borders through it. Four or five segments are as defensible as six (K = 4 has the highest score, 0.215); six gives more useful names. A manager should pick the number of segments that supports a decision and check that they survive a different random start or a different method, such as hierarchical clustering with Ward's linkage.
Average silhouette scores stay between 0.19 and 0.22 for every number of segments from 3 to 8
Source: PDI convenience-store sales (Dewey), 2022; silhouette computed with Euclidean distance on the first six component scores, unweighted. Thresholds follow Kaufman and Rousseeuw (1990).
Brand shares predict the 2020 vote about as well as census variables do
A sharper test of what the basket contains is to predict each ZIP's Republican share of the 2020 two-party presidential vote from its brand shares with a linear regression,
and to judge it by out-of-sample R2. Fit the model on 80% of the ZIPs, predict the other 20%, and repeat five times so that every ZIP is predicted once by a model that never saw it (five-fold cross-validation).9 Out-of-sample R2 is the share of the variation in Republican vote that these held-out predictions explain:
where yi comes from a model fit without ZIP i's fold. In-sample R2 rises with every variable added; out-of-sample R2 rewards only variables that carry information about ZIPs the model has not seen.
Eight census variables (percent white, Black, Hispanic and Asian, median household income, percent with some college, median age and log density) give 0.68. The 40 brand shares give 0.70, the first six components 0.61, and demographics and brands together 0.85 (Figure 7). Across 20 random fold assignments each of these values moves by less than 0.01, while the five folds of a single split range from 0.64 to 0.72 for the brand shares. The average Republican share is 56.8% with a standard deviation of 18.3 points across ZIPs; the typical miss falls from 10.4 points with demographics alone to 7.1 points with demographics and brands.
Brand shares explain 70% of the variation in 2020 Republican vote out of sample, and brands plus demographics explain 85%
Out-of-sample R², by predictors
Source: PDI convenience-store sales (Dewey), 2022; ACS-based ZIP demographics; 2020 vote from Fekrazad (2025). Ordinary least squares with five-fold cross-validation on 2,542 ZIPs, unweighted. R² values are means over 20 random fold assignments.
The brand shares carry information the census variables do not, and the likeliest source is region and culture: the eight demographics contain no location. The table below lists the brands whose shares correlate most with the Republican vote, first raw and then net of the eight demographics (a partial correlation). The raw correlations mostly reflect race, since Newport, Fanta and Sprite sell where more Black and Hispanic residents live. Net of demographics the Republican side is led by Michelob Ultra, Dr Pepper and Marlboro, and the Democratic side by American Spirit, Pepsi and White Claw, which is mostly a Southern against Northern and Western split. State fixed effects would be the natural next step. Bud Light is third on the net-Republican list (0.28), so even allowing for demographics it sold better in Republican ZIPs in 2022, which is where case 3 begins. The correlations describe ZIPs and predict votes. Brands do not move votes, and votes do not move brands; both follow who lives where.
| Higher in Republican ZIPs | Higher in Democratic ZIPs | |
|---|---|---|
| Raw correlation | Monster 0.47, Dr Pepper 0.41, Michelob Ultra 0.38, Marlboro 0.38, Mountain Dew 0.36 | Newport −0.57, Heineken −0.44, Fanta −0.42, Dutch Masters −0.39, Sprite −0.34 |
| Net of the 8 demographics | Michelob Ultra 0.54, Dr Pepper 0.42, Bud Light 0.28, Marlboro 0.26, Natural Light 0.25 | American Spirit −0.44, Pepsi −0.37, White Claw −0.34, Twisted Tea −0.34, ZYN −0.29 |
Bud Light's two strongholds fell by different amounts in 2023
Bud Light's average share of beer units in 2022 was highest in the Deep South segment (18.7%) and the Rural red South and Texas segment (17.6%), against 8.6% in the West and 14.7% across all ZIPs. Case 3 follows the brand through the 2023 boycott. Of that case's 1,156 ZIPs, 1,143 are also in this file. In the 52 weeks after April 3, 2023, compared with the 52 weeks before, Bud Light's share fell 4.5 points (27%) in the rural red South, from 16.9%, and 3.0 points (17%) in the Deep South segment, from 17.8%. The West shows why points and percent can tell different stories: its share was small, 8.5%, and still fell by 30% (2.5 points). A brand concentrated in a few segments depends on what happens in those segments, and the 2022 data show where Bud Light was concentrated the year before the boycott.
Bud Light's share was highest in the Deep South and rural red South segments in 2022, and fell by 17% and 27% a year after the boycott began
Bud Light share of beer units, 2022 (%)
Change, 52 weeks after vs. 52 weeks before April 3, 2023 (points; % in brackets)
Source: PDI convenience-store sales (Dewey); left panel is calendar 2022 for 2,542 ZIPs, right panel is the case 3 ZIP-week panel for 1,143 ZIPs, weeks −52 to −1 against weeks 0 to 51 relative to April 3, 2023; unweighted averages across ZIPs.
| Segment | ZIPs in this file | Bud Light share 2022, % | ZIPs in both cases | Share, 52 weeks before, % | Change, points | Change, % |
|---|---|---|---|---|---|---|
| Deep South, large Black population | 262 | 18.7 | 127 | 17.8 | −3.0 | −17 |
| Rural red South and Texas | 604 | 17.6 | 290 | 16.9 | −4.5 | −27 |
| Midwest heartland | 736 | 14.3 | 307 | 13.8 | −3.9 | −28 |
| Carolinas and the Atlantic coast | 412 | 13.9 | 202 | 13.4 | −2.8 | −21 |
| Hispanic metro | 308 | 12.1 | 137 | 12.2 | −2.6 | −21 |
| The West, "new nicotine" | 220 | 8.6 | 80 | 8.5 | −2.5 | −30 |
The share before the boycott averages the 52 weeks to March 2023, so it differs slightly from the calendar-2022 share.
What these data cannot show
Every number here describes ZIP codes, not people. "Newport is a Democratic brand" means that ZIPs where Newport's share is high also tend to vote Democratic; it says nothing about how Newport smokers vote, and a correlation across areas can differ in size and even in sign from the same correlation across individuals.10
The shares come from what stores sold, which reflects shelf space, distributor contracts, promotions and prices as well as what shoppers prefer. Co-location on the brand map can come from distribution. The sample is convenience stores only: chains and independents in the PDI panel, mostly in the South and Midwest, with a thin Northeast, and no grocery, supermarket or bar sales. Shares also hide volume. A ZIP with two stores gives a noisier share than a ZIP with twenty, the median ZIP has two, and every ZIP counts once in every average above.
Segment numbers and component signs are arbitrary and change with the random start. The segments are an approximation to a continuum, with a silhouette score of 0.200, and the names are descriptions I chose. The components and segments were fit on all 2,542 ZIPs, and the "first six components" prediction uses components computed before the folds were cut. They use no information about the vote, so the leak is small, but a stricter test would refit them within each fold. The 2020 vote is allocated to ZIP codes from other geographies, and the smoking rates are modeled estimates, so both carry measurement error that the correlations do not show.
Questions for discussion
- A regional distributor is launching a flavored malt beverage. Using the segment profiles in Figure 5, which segments would you target first, and what would you want to know that these data cannot tell you?
- A retailer is scouting a site in a ZIP with no sales history. Which census variables would you use to predict where the ZIP sits on the brand map, and how would you check that prediction before committing to an assortment?
- Bud Light's two strongest segments lost 17% and 27% of their share after April 2023. What does it mean for a brand to be exposed to a segment, and what would you want to see besides segment shares before calling a brand concentrated?
- The eigenvalue-above-1 rule keeps ten components and the scree plot suggests three. Argue for each from the decision the components would support. Then use Figure 6 to say whether six segments are too many.
- Demographics give an out-of-sample R² of 0.68, brands 0.70 and both 0.85. What does the gain from combining them say about what the brand shares contain, and what additional variable would you add to test your reading?
- In Figure 7 the typical miss (root mean squared error) of the brand-share model is 10.1 points of Republican vote. Write one sentence about Newport and the vote that these data support and one that they do not.
Replicate this analysis
The folder 02_brand_map has the 40 brand shares and the demographics for every ZIP, and a README with step-by-step exercises. R (prcomp, kmeans) and Python (numpy or scikit-learn) do each step in a few lines. Excel has no built-in PCA; the free Real Statistics add-in does it. The steps are to standardize the 40 columns, take the eigenvectors of their correlation matrix, scale them by the square roots of the eigenvalues to get loadings, cluster the first six score columns with K = 6, and regress rep_share_2020 on the predictors with five-fold cross-validation. Your segment numbers and component signs will differ from the ones here.
Data sources
Brand shares: PDI Technologies point-of-sale data from convenience stores, obtained through Dewey under a research licence. Demographics: American Community Survey ZIP tables. Adult smoking, binge drinking and obesity rates: CDC PLACES model-based estimates. 2020 presidential vote by ZIP code area: Fekrazad (2025), doi:10.1038/s41597-025-05140-3. ZIP positions: Census 2020 ZCTA internal points; state outlines: Census cartographic boundaries. The files are derived from licensed data; check the Dewey and PDI licence terms before posting them outside the course.
Methods note
Sample: 2,542 ZIP codes in 44 states, each pooling at least two convenience stores (7,626 stores in all), calendar year 2022. Each brand share is the brand's units as a percentage of its category's units. The analysis uses 40 shares (16 beer, 7 soft drinks, Monster, 6 cigarette, 4 smokeless, 6 cigar) and leaves out the "other" columns and the category-mix columns. Shares are standardized across ZIPs (mean 0, standard deviation 1), and every ZIP counts once; there are no population or store weights. PCA is the eigendecomposition of the 40 by 40 correlation matrix; a loading is the correlation between a brand's share and the component score, and signs are set so that Fanta, White Claw and Dr Pepper load positively on components 1, 2 and 3. Segments are k-means with K = 6 on the first six component scores (53.5% of the variance); the assignments come from one k-means run, and the silhouette scores use Euclidean distance on the same scores. Prediction is ordinary least squares with an intercept and random five-fold cross-validation; out-of-sample R2 pools the held-out predictions, is averaged over 20 random fold assignments, and the whiskers in Figure 7 are the fold-level R2 range from the first split. The demographic list is percent white, Black, Hispanic and Asian, median household income, percent with some college, median age and log population density. The Bud Light link uses case 3's ZIP-week panel: each ZIP's average share over the 52 weeks before and after the week of April 3, 2023, then averaged within segment. The percent change is the ratio of the segment averages. No standard errors are reported.
How to cite
@misc{singh2026brand,
author = {Singh, Vishal},
title = {What the checkout counter says about a neighborhood},
year = {2026},
note = {Teaching case, NYU Stern School of Business},
url = {https://vishalsingh.org}
}
References
- Fekrazad (2025). Scientific Data, doi:10.1038/s41597-025-05140-3 (2020 presidential vote allocated to ZIP code areas).
- Jolliffe, I. T. (2002). Principal Component Analysis, 2nd edition. Springer.
- Cattell, R. B. (1966). The scree test for the number of factors. Multivariate Behavioral Research 1(2), 245–276.
- Kaiser, H. F. (1960). The application of electronic computers to factor analysis. Educational and Psychological Measurement 20(1), 141–151.
- MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, 281–297.
- Lloyd, S. P. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory 28(2), 129–137.
- Rousseeuw, P. J. (1987). Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics 20, 53–65.
- Kaufman, L. and Rousseeuw, P. J. (1990). Finding Groups in Data: An Introduction to Cluster Analysis. Wiley.
- Hastie, T., Tibshirani, R. and Friedman, J. (2009). The Elements of Statistical Learning, 2nd edition. Springer.
- Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review 15(3), 351–357.