In Week 1 you saw that competitor prices are a cloud, not a point. Now we learn to describe that cloud precisely — picture it, find its centre, measure its spread, locate any single price inside it — and turn the description into an actual answer to "what should we charge?" This is the textbook's chapter 2, and every exam has a question from it.
The session is three hours in two parts. Part 1 (~1 h) is the lesson: your instructor presents the deck and you follow the same numbered sections here — every slide names its section. Part 2 (~2 h) is eight exercises worked together, given on the slides and as forms below; solutions open as your instructor reveals them.
Thirty Margherita prices are on your notepad. Before you look at a single one, an instinct check: what share of Geneva pizzerias charge more than CHF 20 for a Margherita?
Thirty is a sample; Geneva's ~300 pizzerias are the population. Every number this week is a statistic — computed from the thirty, and it would come out a little different from a different thirty. That "a little different" is Week 5's whole subject; for now we simply describe what we have.
Description first. You cannot reason about a cloud you have not looked at.
Before any average, plot the data. A frequency histogram groups prices into classes and shows how many fall in each; a relative frequency histogram shows the same bars as a share of the whole. Drag the class width and watch the story change — too wide hides detail, too narrow turns it into noise.
The tail stretches to the right — a few pricey tourist spots. That is a right-skewed distribution, and it will matter the moment we pick an average.
A histogram shows shape but loses the actual values. A stem-and-leaf diagram shows the shape and every price: the stem is the tens digit, each leaf is a ones digit.
Read it sideways: it is a histogram you can still read the raw data off.
For symmetric data they roughly agree. For skewed data they split — and the gap is the lesson.
The mean chases extremes; the median resists them. Report the median when a few big values would otherwise mislead.
Below, Pâquis and Rive both average CHF 18. But one is tightly clustered and one is all over the map. The standard deviation (the typical distance from the mean, from Week 1) and the interquartile range put a number on that difference.
Three rulers for spread, weakest to strongest: range (max − min, fooled by one outlier), IQR (the middle 50%, resistant), and standard deviation (uses every value).
Week 1 built the standard deviation of a population: average the squared distances, divide by N. For a sample you divide by n − 1. The reason: the sample's own mean x̄ sits a little closer to the sample's values than the true µ would, so the squared distances come out slightly too small. Dividing by one less pushes them back up. On the exam, the word "sample" is your cue to use n − 1 — most lost points in this chapter are exactly here.
The p-th percentile is the value below which p% of the data sit. The quartiles Q₁, Q₂ (the median) and Q₃ are the 25th, 50th and 75th. Sort the data and you get five landmarks: minimum, Q₁, median, Q₃, maximum. The box plot draws them: a box from Q₁ to Q₃ (the middle 50%), a line at the median, whiskers to the data, and dots for outliers beyond the fences Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR.
A z-score says how many standard deviations a value sits from the mean: z = (x − x̄) ÷ s. Type a competitor's price and see.
A z near 0 is typical; |z| above ~2 is genuinely unusual. z-scores let you compare positions across different datasets on one ruler.
Suppose your pizzeria's daily covers (customers served) are bell-shaped with mean 120 and standard deviation 15. The Empirical Rule says ≈68% of days fall within ±1 SD, ≈95% within ±2, ≈99.7% within ±3. Slide to see the band.
Chebyshev makes no shape assumption: at least (1 − 1/k²) of any data lies within k standard deviations. Weaker than 68–95–99.7, but always true. Compare the two columns above as you slide.
The Empirical Rule is a generous estimate that needs a bell; Chebyshev is a cautious guarantee that needs nothing.
Your instructor puts each exercise on the slide; you work it here, by hand, with a calculator — the exam is by hand, so this is the rehearsal. Check marks your number immediately, Hint nudges without giving it away, and the Solution button opens when the room does that exercise together.
Around 100 minutes for all eight, on one dataset of fifteen Eaux-Vives pizzerias. Your answers are saved and reach your instructor's dashboard. Nothing here is graded.
One attempt per question, every answer explained. Your score syncs to your instructor's dashboard as a self-check but does not count toward your grade. Have a calculator to hand — two of these ask for a number, exactly as the exam will.
In Week 3 we leave description for probability: random variables, and in particular the binomial distribution — the tool for "what is the chance we sell out of dough on a Friday?" Before that, the Week 2 homework: a fresh dataset, everything from this week by hand.
The competitor data, described: a typical price (the median), a middle-of-the-market band (Q₁ to Q₃, where half of all pizzerias sit), and a clear outlier you should ignore. Pick your Margherita price and see exactly where it lands among the competition.
"Around the median, inside the middle 50%" is a position you can justify to your partner and your bank. That is the whole job of descriptive statistics: turning a cloud into a defensible decision.
📎 Everything saves automatically and reaches your instructor's dashboard. This is the second instalment of the individual project.