§3.4

Uncertainty and Statistical Charts

Executives do not need confidence intervals because they enjoy statistical ritual. They need them because decisions happen before certainty arrives, and an interval is the cheapest honest answer to the only question that matters at that moment: is this difference big enough, precise enough, and stable enough to act on? The charts in this section are the last ones in Part II, and they are the ones that carry statistical content — distributions that explain why a transform is needed, scatterplots that show a slope before any equation names it, and intervals that mark the boundary of what the data can support. If a manager can read these four visuals, the regression in Part III arrives as compact notation for something already understood rather than as a new subject.

The executive question: what does the chart show, and where does it stop?

The Progresso soup pattern is visually strong: share falls in warm months and rises in the soup season, and price moves against demand. But the data is a panel of store-months, not a census of every possible market condition. Store coverage changes over time. Store-level behavior varies enormously. A chart that shows the finding without the limits around the finding is not being confident, it is being incomplete.

Statistical charts do that limit-marking work in three moves, and they build on each other:

  1. Shape before summary. Look at the distribution before trusting any average of it.
  2. Slope before equation. Read the relationship visually before estimating it.
  3. Interval before verdict. Show the range of plausible values, and say what the range is about.

Shape before summary

Raw Progresso store-month volume has a long right tail: many ordinary store-months and a few very large ones. That shape has an immediate managerial consequence. The mean sits well above where the bulk of stores actually are, so any staffing or inventory plan built around "the typical store-month" from the raw average over-resources the ordinary store and under-resources the rare giant.

Figure 1 shows the raw histogram — fixed-width bins clipped near the 99th percentile so the shape is visible rather than dominated by the largest stores — beside the log transform that compresses the tail and makes typical store-months comparable.

Raw volume is a long-tail distribution

Most store-months are modest; a few stores move enormous volume. The count axis is shared with the log view so the shape change is the only difference.

x-axis

Slopes change once seasonality enters the picture

A forest plot: each estimate is a center, a 95% interval, and a comparison to zero. The dashed line is the no-relationship baseline; later pricing chapters handle identification.

Figure 1. Fixed-width histograms show the long-tail soup volume problem; the log transformation compresses that tail and makes typical store-months easier to compare. The coefficient intervals in the same figure preview what an estimate looks like as a visual object.

The log is not a cosmetic move, and treating it as one is the most common misunderstanding in this whole area. Taking logs changes the question from differences in units to differences in percentages — which, for a pricing manager, is usually the question that was being asked all along. Use it because the business logic is multiplicative, not because the chart looks tidier afterward.

Figure 1 also introduces coefficient intervals, and the point there is not regression mechanics. The point is that a statistical estimate is a visual object with four parts: a center, an interval, a comparison group, and a limit statement. Every estimate in the rest of this book has all four, whether or not the chart shows them.

Slope before equation

The most intuitive statistical chart in the soup case is log(volume) against log(price). In a log-log chart the slope carries elasticity intuition directly: a one percent price difference is associated with some percent difference in volume. That is why the pricing chapters use log regression, and it is worth seeing before the equation arrives.

-2.46
The national month-adjusted log-log slope: a 1% higher Progresso price is associated with about 2.46% lower volume in this descriptive preview.

That number is not the final price elasticity. It is an intuition-building estimate, and the soup case also shows exactly why it cannot be the last word. Winter months are high-demand months. Non-winter months are lower-demand months. If Progresso's pricing moves across that same seasonal cycle — and the indexed chart showed that it does — then the scatterplot is mixing price behavior with seasonal demand, and the slope is measuring both at once.

Figure 2 separates the scatter by winter and non-winter, with winter defined here as October through February.

Log price–volume slope by season

A downward log-log slope previews price elasticity. The fitted line and slope are descriptive, not yet causal.

Figure 2. Winter and non-winter scatterplots both slope downward, but the comparison is still descriptive: seasonality and pricing strategy move together, so the slope cannot separate them.

Figure 2 is useful precisely because it is not enough. It hands the reader the visual intuition for elasticity while leaving the identification question standing in plain view: what price variation is independent of demand shocks, promotions, inventory, store mix, and seasonality? That question is the whole subject of Chapter 6, and Part II's contribution is to make it feel necessary rather than pedantic.

Interval before verdict

Figure 3 shows monthly mean Progresso share across store-months with intervals attached. The intervals are narrow, because there are tens of thousands of store-month observations. Narrow is useful. Narrow is not the same as certain.

Seasonal share intervals are narrow — but narrow is not causal

Dot area encodes coverage: active store-months range from 6,643 to 7,984. A precise mean built on thin coverage still deserves a second look.

Region × season on one shared axis

A real dot-and-whisker on a single share axis. Winter share sits above non-winter in every region, but the East operates at a different level entirely.

Figure 3. Progresso share is reliably lower in the summer months, but the intervals describe store-month variation — not the causal effect of price.

Figure 3 should be read in two passes. First read the pattern: share is lower in summer and higher across the broader soup season. Then read the caveat: the interval describes observed store-month variation around a monthly summary. It says nothing at all about what would happen if Progresso changed price tomorrow.

That distinction is central to the rest of the book, and it is worth stating as a rule: statistical precision does not repair a weak comparison. A very narrow interval around a biased estimate is still a biased estimate — it is just a confidently wrong one. Precision and identification are separate properties, earned by separate means, and a chart that shows only the first invites readers to assume the second.

There is a second, less glamorous reason intervals differ in width here. The soup panel is unbalanced: stores enter and exit the observed panel, so monthly active store counts vary. That does not make the data unusable. It means the coverage note belongs next to the figure, not buried in an appendix. If a December point rests on many more active stores than a June point, the reader needs that before reading the chart as pure seasonality — and the same principle will matter again in model monitoring, survey analysis, A/B tests, and AI evaluation.

Figure 4. Uncertainty language is only useful when the chart explains what kind of uncertainty the interval represents.
RuleInterpretation
Rule 1Narrow intervals can still be biased if the comparison is wrong.
Rule 2Coverage changes over time, so active store counts belong near the chart.
Rule 3Store-month intervals describe observed variation, not the causal effect of price.

Figure 4 sits deliberately close to the chart rather than in a technical appendix, because it encodes a reporting habit: every interval should answer what is varying, what is being averaged, and which decision would change if the interval were wider.

Concept check

Four questions spanning what a narrow interval certifies, why the log transform is a change of question, how to read overlapping and unequal-width intervals, and what a distribution reveals that a mean cannot.

  1. 1.
    A dashboard shows last quarter's regional sales gap with very tight error bars, so a VP concludes the gap is a reliable basis to shift the ad budget. What is the strongest objection?
  2. 2.
    Store-month soup volume is sharply right-skewed, and this section takes the log before plotting price against volume. What is the real reason to use logs here?
  3. 3.
    Two months' mean-share intervals overlap slightly; a teammate says "overlap means no real difference, so ignore it." A second month's interval is much wider than the first. How should you reason?
  4. 4.
    Across all Progresso store-months, average monthly unit volume is about 900 units, and the finance team proposes staffing and inventory plans assuming a typical store-month sells around 900. Before signing off, you plot the full distribution. Which finding would most change how you act on that 900 figure?