Intro to probability · The normal distribution

From sum
to mean

One question, two parts, and one difference that trips up most of the class: when you add observations up, the spread grows; when you average them, the spread shrinks. This page builds that from scratch — until you can solve a question like this with your eyes closed.

Practice this
The given

The problem we're solving

A population is normally distributed with mean 80 and variance 16. A researcher draws 180000 samples of size 16 each. Given:

μ = 80 // population mean σ² = 16 // population variance σ = 4 // standard deviation = square root of the variance n = 16 // size of each sample a = P( 3 · (X₁ + X₂ + ... + X₁₆) > 3852 ) b = P( X̄ < 78.5 ) // we need the product a·b

The four options on the exam were:

a) 0.148 b) 0.7851 c) 0.0268 d) 0.5106
Watch out for the first trap — the figure "180000 samples" doesn't enter any calculation. It only tells us that sampling is repeated many times. Questions like this always carry one extra piece of data; don't go looking for a place to plug it in.
Practice this
01 · The ABCs

What every symbol means

Before any calculation — let's make sure every character in the question is clear. These are all the symbols you'll meet, and nothing more.

μ
Mu — the expectationThe mean of the whole population. The number everything clusters around. Here: 80.
σ
Sigma — standard deviationHow spread out the values are around the expectation, in the same units as the data. Here: 4.
σ²
VarianceThe standard deviation squared. Variance is handy for calculating; standard deviation is handy for understanding. Here: 16.
n
Sample sizeHow many observations are in one sample. Here: 16.
Xi
The i−th observationMeasurement number i in the sample. X1 is the first, X16 the last.
Σ
Capital sigma — a sumAn instruction to add. The sign in the question simply means X1 + X2 + ... + X16.
X̄
"X bar" — the sample meanThe sum divided by n. A random variable in its own right: every sample gives it a different value.
Z
Standard normalA special normal variable with expectation 0 and standard deviation 1. It's the only one that has a table.
Φ
Phi — the cumulative distribution functionΦ(z) is exactly the probability P(Z < z). It's the number you read off the table.
N( , )
Distribution notationN(a,b) = normal with expectation a and variance b. The second number is the variance, not the standard deviation!
Practice this
02 · Intuition

What a normal distribution actually is

Imagine measuring the height of every student in your year. Most values will land close to the average, and only a few will be extreme. Draw it — and you get a bell. That's the normal distribution.

Two numbers define it completely: where the bell is centered (that's μ) and how wide it is (that's σ). Play with the sliders and see exactly what each one does.

The area under the whole bell is always 1, that is, 100%. So when the bell gets wider it must get lower — there's no "more probability", it's just spread over a wider range. Probability = area. That's the one idea to keep from this part.

The 68−95−99.7 rule of thumb

In every normal distribution, whatever μ and σ are:

P( μ - 1·σ < X < μ + 1·σ ) ≈ 68% P( μ - 2·σ < X < μ + 2·σ ) ≈ 95% P( μ - 3·σ < X < μ + 3·σ ) ≈ 99.7%

It's a great sanity check in the exam: if you end up with a probability of 0.4 for something three standard deviations from the mean — you went wrong somewhere.

Practice this
03 · The tools

Two rules for expectation and variance

The whole question rests on four lines. They're worth memorizing exactly as written:

Rule A — a sum of independent variables

E( X₁ + X₂ + ... + Xₙ ) = n · μ Var( X₁ + X₂ + ... + Xₙ ) = n · σ²

Both the expectation and the variance simply add up. Each extra observation adds another μ to the expectation and another σ² to the variance.

Rule B — multiplying by a constant

E( c · X ) = c · E(X) Var( c · X ) = c² · Var(X)

This is the biggest trap in the question: the expectation is multiplied by c, but the variance is multiplied by c squared. Why? Because variance measures squared distances. Stretch every value by a factor of 3, and every distance from the mean grows 3 times — so every squared distance grows 9 times.

The safe way to remember it: the standard deviation behaves "normally" —

SD( c · X ) = |c| · SD(X)

and you can always get the variance from it by squaring. If you get confused, work with the standard deviation and square at the end.

Two variables: the original, and the same variable multiplied by 3. The center moves 3 times as far — and the width stretches 3 times too, which means the variance grows 9 times.
Practice this
04 · The heart of the question

Sum vs. mean

The two parts of the problem look alike, but they pull in opposite directions. Here's the line that has to stick:

Sum — the spread grows

Var(S) = n · σ²

Every new observation adds noise of its own. The pile gets bigger and wilder.

Mean — the spread shrinks

Var(X̄) = σ² / n

The deviations cancel each other out. The bigger the sample, the steadier the mean.

Where it comes from — one line

The mean is just the sum divided by n, that is, the sum times the constant 1/n. Apply Rule B:

Var(X̄) = Var( S / n ) = (1/n)² · Var(S) // Rule B: the constant, squared = (1/n²) · n · σ² // Rule A = σ² / n

That's it. The whole difference between the two parts comes from that square on the n.

Watch it happen

Below we simulate real samples from the problem's population. Single observations and means of 16 observations are both centered around 80, but look at the width.

samples = 0 SD(X) → 4 SD(Xbar) → 1
With n = 16 the standard deviation of the mean is 4 times smaller than that of a single observation, because the square root of 16 is 4. Notice that it shrinks with the square root of n, not with n — to cut the width by a factor of 2, you need a sample 4 times as large.
Practice this
05 · The common language

Standardizing: how everything becomes Z

There are infinitely many normal distributions — one for every pair of μ and σ. Nobody can print a table for each of them. The solution: translate every question into the language of one standard bell, centered at 0 with a standard deviation of 1.

Z = ( X - E(X) ) / SD(X)

The two steps this formula performs:

1. Subtract the expectation

Shifts the bell so its center sits on 0. Now the value says "how far above or below the mean I am".

2. Divide by the standard deviation

Squeezes or stretches the width to 1. Now the value says "how many standard deviations above the mean I am".

What a z-score means: the number 0.25 we'll get later simply says "this value is a quarter of a standard deviation above the mean". That's all Z is — a uniform ruler for distance.

Drag slowly: the bell of N(80, 16) travels left until it reaches 0, then shrinks to width 1. The shaded area doesn't change during the whole journey — and that's exactly why we're allowed to compute the probability on the standard bell instead of the original one.
Practice this
06 · The tool

From Z to a probability

The Z table always gives the area to the left, that is

Φ(z) = P( Z < z )

Every other case follows from it. These three lines cover any question:

P( Z < z ) = Φ(z) // straight from the table P( Z > z ) = 1 - Φ(z) // the complement P( Z < -z ) = 1 - Φ(z) // symmetry around 0

Why is the third line true? Because the standard bell is perfectly symmetric around 0, so the left tail to the left of (−z) has the same area as the right tail to the right of z. Most tables print only positive values, so you need this flip.

Phi(z) = 0.5987 1 - Phi(z) = 0.4013
The shaded area is P(Z < z). The two values we'll need in the problem are 0.25 and −1.5 — press the buttons to jump to them.

The exact table values we'll use:

Φ(0.25) = 0.5987 Φ(1.50) = 0.9332
Practice this
07 · The solution

Full solution, step by step

You now have all the tools. The order is fixed and doesn't change from question to question: identify the expression → find the expectation → find the variance → take the square root → standardize → table.

Part A — computing a

1

What's the expression? Call the sum S; then the expression inside the probability is 3S.

S = X₁ + X₂ + ... + X₁₆
2

The distribution of S — by Rule A:

E(S) = n · μ = 16 · 80 = 1280 Var(S) = n · σ² = 16 · 16 = 256 S ~ N(1280 , 256)
3

Multiply by 3 — by Rule B. The expectation times 3, the variance times 9:

E(3S) = 3 · 1280 = 3840 Var(3S) = 3² · 256 = 9 · 256 = 2304 3S ~ N(3840 , 2304) SD(3S) = √(2304) = 48
4

Standardize and look it up in the table:

z = (3852 - 3840) / 48 = 12 / 48 = 0.25 a = P(3S > 3852) = P(Z > 0.25) = 1 - Φ(0.25) = 1 - 0.5987
a = 0.4013
The distribution of 3S, and the right tail beyond 3852. Notice the tail is a little less than half — which makes sense, because 3852 is only slightly above the center.
Practice this
07 · The solution · continued

Part B — computing b

1

The distribution of the mean — this time the variance gets divided:

E(X̄) = μ = 80 Var(X̄) = σ²/n = 16/16 = 1 X̄ ~ N(80 , 1) SD(X̄) = √(1) = 1
2

Standardize. This time z comes out negative, so we use symmetry:

z = (78.5 - 80) / 1 = -1.5 b = P(X̄ < 78.5) = P(Z < -1.5) = 1 - Φ(1.5) = 1 - 0.9332
b = 0.0668
The distribution of X̄, and the left tail below 78.5. This bell is much narrower than the one for a single observation — so the probability is small: 78.5 is one and a half standard deviations from the center.
The answer
a · b = 0.4013 · 0.0668 = 0.02680...

ab ≈ 0.0268 — that is, option c.

Practice this
08 · Before the exam

Traps and the shortcut

Trap 1 — mixing up variance and standard deviation. In the notation N(a,b) the second number is the variance. If it says 16, the standard deviation is 4. Before every standardization ask yourself: "is the number I'm about to divide by a square root?"
Trap 2 — multiplying the variance by the constant instead of its square. If you see 3S, the variance gets multiplied by 9. Safest: work with the standard deviation, where you just multiply by the constant.
Trap 3 — using σ instead of σ/√n. The moment you see X̄, the sample size has to enter the calculation. If you got the same answer as for a single observation — you forgot it.
Trap 4 — forgetting to flip the direction. The table gives the area to the left. If the question asks "greater than", you need 1 - Φ(z).

The shortcut — for the confident

Notice that the sum and the mean are really the same variable in disguise. Since S = 16 · X̄, we have:

3S = 3 · 16 · X̄ = 48 · X̄ P(3S > 3852) = P( X̄ > 3852/48 ) = P( X̄ > 80.25 ) z = (80.25 - 80) / 1 = 0.25 // exactly the same z

Same z, same answer, in a third of the work. It's also a great check: if the two methods disagree — one of them has a mistake.

The memory card

Var(X₁+...+Xₙ) = n · σ² // sum: grows n times Var(X̄) = σ² / n // mean: shrinks n times Var(cX) = c² · Var(X) // constant: squared SD = √(Var) // always a square root before standardizing Z = (x - E) / SD // then to the table P(Z < -z) = 1 - Φ(z) // symmetry
Practice this
09 · Self-check

Where you stand

The four check questions are spread across the relevant slides. Here's where you are — click a row to jump to its question.

Three questions to ask yourself before the exam

If you can answer these without scrolling up — you're ready for this question:

1. When is the variance divided by n, and when is it multiplied by n ? 2. Where does the square come in when you multiply by a constant ? 3. When do you need 1 - Phi(z) instead of Phi(z) ?

And if something still doesn't sit right — mark it on the slide with "Mark what I didn't get", or ask for a different explanation. Don't leave a gap.

Practice this