HEG GenΓ¨ve APPLIED STATISTICS Β· WEEK 1 ← Course
This week
0%
Course
0%
Week 1 · Applied Statistics · HEG Genève

Why statistics is hard β€” and what it actually is

Before a single calculation, we deal with the real reason statistics feels impossible: it is a model of the world, not the truth. We meet that idea β€” and the traps that follow from it β€” by trying to do one ordinary thing: open a pizzeria in Geneva and decide what to charge. Along the way you get the vocabulary of the textbook's first chapter.

The session is three hours in two parts. Part 1 (~1 h) is the lesson: your instructor presents the deck and you follow the same numbered sections here β€” every slide names its section. Part 2 (~2 h) is eight exercises worked together, given on the slides and as forms below; solutions open as your instructor reveals them.

🏫 In class: 3 h Β· lesson ~1 h, then exercises ~2 h Β· bring this page open πŸ“– Textbook: Saylor, ch. 1 β€” Introduction β–Ά The Lecture button (top-right) shows the deck we use in class

By the end of Week 1 you will be able to

  • Say why a single dataset has no single "true" answer hiding inside it β€” and why the spread is half the story.
  • Tell a population from a sample, a parameter from a statistic, and descriptive from inferential statistics.
  • Classify any variable as qualitative or quantitative, discrete or continuous β€” and know why the type decides what you may compute.
  • Explain the core idea of the whole course: a method can be true while its conclusion is only conditional on a model you chose.
  • Name the traps β€” noise, sampling, confounding, the invisible population β€” before they ambush you.
  • Recognise the everyday words that mean something dangerously different in statistics.

Resources β€” tick them off as you go

βœ“Lecture deck β€” the β–Ά Lecture button above; presented in class in Week 1.
βœ“Saylor Β§1.1 β€” Basic definitions and concepts β†— β€” population, sample, parameter, statistic, kinds of variables. Ten minutes; read it after section 1.3.
βœ“Saylor Β§1.2 β€” Overview β†— and Β§1.3 β€” Presentation of data β†— β€” short; the second one is the seed of Week 2.
βœ“Part 2 Β· the exercise set β€” eight exercises, ~95 min, done with your instructor in the second half of the session. Section 1.7 below.
βœ“Week 1 homework β€” ~2 h of calculation on a ten-pizzeria dataset, with hints and worked solutions. Not graded; it is how you find out whether the lesson stuck.
Section 1.8 links the four intuition-building tools that already live in your toolkit.
← Course page
A reset link lives at the bottom of section 1.9.
Section 1.1 Β· The hookslide 3

The pizza problem

Your decision

You're opening a pizzeria in Geneva. What do you charge for a Margherita?

You and a partner have signed the lease on a small place near Plainpalais. The oven is in. The very first number you must commit to is the price of the Margherita β€” the dish everyone uses to judge whether you're cheap, fair, or a rip-off.

No data yet. No formulas. Just your gut. Pick a price and lock it in. We come back to this exact number at the end of the week β€” and again at the end of the semester.

Set your opening price
CHF 18.0
Drag to choose. There is no right answer yet β€” that's the point.
Locked. You just made a decision under uncertainty with almost no information β€” which is exactly what statistics is for. Everything we build from here is about replacing that gut number with one you can defend.
Why this week exists

Most statistics courses open with a formula. This one opens with a decision, because the formulas only make sense once you have felt the problem they solve: you must act, the world is noisy, and you can never see all of it.

Statistics is the discipline of deciding well when you cannot see everything.

Section 1.2 Β· Variationslide 5

There is no number in the data

The first surprise

So you go and look at what others charge.

Sensible. You spend an afternoon walking Geneva, ordering a Margherita at every pizzeria you pass, and writing down the price. Click to visit them one at a time and watch the prices land.

Walk Geneva β€” collect the prices
0 of 16 visited
Each pizzeria you visit drops one price onto the line.
The words the textbook uses

Population, sample, parameter, statistic.

You care about all Geneva pizzerias β€” roughly three hundred of them. That is the population. You measured sixteen: that is your sample. Any number you compute from the sixteen β€” the average, say β€” is a statistic. The unknown true value across all three hundred is a parameter. This distinction runs through the whole course, and the exam will ask it in a dozen disguises.

sample β†’ statistic (you can compute it)  Β·  population β†’ parameter (usually unknown)

Two kinds of statistics follow from it. Descriptive statistics summarise the data you actually have β€” the sixteen prices. Inferential statistics use the sample to say something about the population β€” the three hundred β€” and to say how sure you can be. Weeks 1–2 are descriptive; from Week 5 onward the course is inference.

Try it

Which is it?

Population Β· Sample Β· Parameter Β· Statistic
Section 1.3 Β· The data itselfslide 8

What kind of thing did you write down?

Two kinds of variable

Numbers you can do arithmetic on β€” and labels you cannot.

Everything you record about a pizzeria is a variable: it varies from one pizzeria to the next. The textbook splits variables in two. Qualitative variables are categories β€” the neighbourhood, the type of oven, whether it delivers. Quantitative variables are numbers with arithmetic meaning β€” the price, the number of tables. An average price makes sense; an "average neighbourhood" does not.

Quantitative variables split once more. Discrete ones are counts with gaps between possible values β€” 12 tables, 13 tables, never 12.4. Continuous ones can take any value on a scale, limited only by your measuring instrument β€” a price can be CHF 16.50, a dough ball can weigh 247.3 g.

The type of the variable decides which displays and which calculations are even legal. Get this reflex automatic.

Try it

Classify each thing you could record about a pizzeria.

Qualitative Β· Quantitative discrete Β· Quantitative continuous
Presenting the data

A list, a table, or a picture.

The sixteen prices as you wrote them down are a data list. The moment you start counting how many pizzerias charge each price you have built a data frequency table β€” the same information, organised. Flip between the two below. Week 2 adds the third form, the picture, and that is where a shape first becomes visible.

Same sixteen prices, two presentations
The list, in the order you collected it.
Section 1.4 Β· Measuring the spreadslide 10

How wide is the cloud?

From "there's a spread" to one number

The mean says where the cloud sits. The standard deviation says how wide it is.

You already trust the mean β€” the balance point of the prices. The standard deviation just answers the next question: how far is a typical pizzeria from that mean? That is the whole idea β€” a typical distance from the mean. The formula only looks frightening because of one twist, and below you can watch that twist happen.

Build the standard deviation β€” step by step
β€”
mean
β€”
naΓ―ve avg distance
β€”
standard deviation
Five pizzerias on the line. The red line is the mean β€” the balance point.
Why squares? (the twist)

Watch step 2 carefully. If you simply average the raw distances, the ones on the left are negative and the ones on the right are positive, and they cancel to exactly zero β€” every single time, for any dataset. The "average distance from the mean" is always 0. Useless. That is the dead end your intuition hits.

The fix is mechanical, not mystical: square each distance (a negative squared turns positive), average the squares β€” that average is the variance β€” then take the square root to get back into francs. That square root is the standard deviation.

Standard deviation = the side of the average square you build on the distances from the mean.

One thing to notice as you drag: the standard deviation always comes out a little larger than the naΓ―ve average distance. Squaring gives extra weight to the pizzerias that sit far out β€” so the standard deviation is especially sensitive to the big departures, which is usually exactly what you care about.

Section 1.5 Β· The keystone ideaslide 12

The map is not the territory

The most important idea in the course

To get a recommendation, you have to impose a model. And the model is your choice.

The same 16 prices can be read two completely reasonable ways. Flip between them and watch the recommended price change β€” on identical data.

Same data Β· two models

Neither model is a lie. Each is a defensible model of "the Geneva pizza market," and each gives a different, confident-sounding answer. The data didn't decide β€” you did, when you chose the model.

"All models are wrong, but some are useful."β€” George Box

The method is true. The conclusion is only conditional β€” on a model you chose and are responsible for.

Hold onto this. Every test, interval and p-value in this course is a deduction that is valid given its assumptions. Whether those assumptions fit your pizzeria is a judgement the mathematics can never make for you. Most of the pain in learning statistics is mistaking the certainty of the maths for certainty about the world.

Section 1.6 Β· The trapsslide 14

Four traps, before they ambush you

Trap 1 Β· Signal vs noise

"Sales were up 12% on Friday β€” something's working!"

Maybe. But random variation alone throws big swings around all the time. Below, every week has the same true average β€” only noise differs. Run a few months and count how often pure chance hands you a "+10% or more" week.

Is a 12% jump real?
big swings from pure noise: 0
True weekly average never changes. The bars only move because of randomness.

Statistics is, at heart, the discipline of asking: is this signal, or is it noise?

Trap 2 Β· Sampling

Who exactly did you ask?

Suppose you only surveyed the cheap takeaway counters by the train station. Is that "Geneva"? Pick a sampling approach and see how the answer lurches.

Where do you sample?
β€”
your estimate
β€”
pizzerias

The sample you can reach is rarely the population you care about. This gap is why sampling distributions β€” the keystone of the whole course, Week 5 β€” exist.

Trap 3 Β· Correlation β‰  causation

"Pricier pizzerias get better reviews β€” so let's just raise the price!"

The dots really do trend upward. Then reveal what you couldn't see β€” the neighbourhood each pizzeria sits in.

Price vs rating
Trap 4 Β· The invisible population

What is "all Geneva pizzerias", exactly?

The ones open today? Including the kebab shop that does two pizzas? The one opening next month β€” yours? The population you reason about is partly invented; the "true average price" is a feature of a model, not a stone tablet. You will spend this course making confident statements about a thing you can never fully see.

Comfort with reasoning about things you cannot observe is the quiet skill statistics demands.

Run the noise simulation, pick a sample, and reveal the lurking variable to finish this section.
Section 1.7 Β· Part 2 Β· Exercisesslides 17–24

Eight exercises, worked together

The second half of the session

Same exercises on the screen and on this page

Your instructor puts each exercise on the slide; you work it here. Check marks a numeric answer immediately, Hint nudges without giving it away, and the Solution button opens when the room does that exercise together β€” so nobody races ahead and everybody has the worked answer afterwards.

Around 95 minutes for all eight. Your answers are saved and reach your instructor's dashboard, which is how the pace of the room gets set. Nothing here is graded.

Section 1.8 Β· Checkpoint

Prove it to yourself

Warm-up Β· false friends

Statistics hijacks ordinary words and gives them sharp, different meanings.

Each card shows a word and its everyday meaning β€” the trap. Tap to flip it to what it actually means in statistics. This is doubly hard if English isn't your first language: you're decoding the language and the jargon at once. Not scored β€” it is here to loosen you up before the quiz.

Tap each word to flip it
0 of 6 flipped
Self-check Β· not graded

Ten questions on this week

One attempt per question, every answer explained. Your score syncs to your instructor's dashboard as a self-check but does not count toward your grade β€” the final exam tests exactly this style of thinking, on paper, with numbers.

β€”

Where this goes next

In Week 2 you learn to describe the cloud precisely β€” picture it, find its centre, measure its spread, locate any single price inside it β€” and turn the description into a price you can defend. Do the Week 1 homework first: it makes you compute by hand what the widgets computed for you.

Section 1.9 Β· Workshop

Open the pizzeria file

The course exercise

One pizzeria, the whole semester

Every week ends here: a few questions that make you apply the week's ideas to your pizzeria, in your own words. Everything saves automatically and reaches your instructor's dashboard β€” no "submit" button, just type. By Week 16 these pages are the file you would take to a bank.

πŸ“Ž This is the seed of the individual project (15%). Nothing here is graded on its own; it is the material the project grows out of.

Week 1 Β· four questions
1.The price you locked in on instinct in section 1.1 β€” copied here so it is on the record. Change it only if you have changed your mind.
2.Define the population you are actually reasoning about when you set a price. Be precise: which pizzerias count, and which don't?
3.How would you sample it in one afternoon β€” and what bias do you fear most in the sample you can actually reach?
4.One market or two? Which model from section 1.5 fits the business you intend to run β€” and what would make you change your mind?
The map of the semester

Every week is a real decision for your Geneva pizzeria.

We keep one running business and let its decisions pull in exactly the statistics each one needs. A handful of questions carry all the classical topics.

Unlike some subjects, the topics are tightly stacked. A shaky grasp of distributions silently breaks confidence intervals three weeks later, with no obvious symptom. This is the real reason students feel lost in Week 8 β€” the crack was in Week 3. It's hard for structural reasons β€” not because you're "not a maths person."

Warm up your intuition now

These standalone playgrounds already live in your toolkit β€” each one drills a trap you just met:

Fill in all four fields to complete this section.