Design studies, collect real data, and argue conclusions with actual statistical reasoning, from describing a single variable through confidence intervals and significance tests, taught unit by unit to the College Board framework. A graphing calculator is allowed on the entire exam, and every lesson shows the calculator route alongside the by-hand one.
3H EXAM40 MCQ6 FRQ48 LESSONS90 PRACTICE PROBLEMSPREREQ: ALGEBRA II
Course overview
What this course covers, and how the exam weights it.
AP Statistics follows the nine units of the College Board course framework, which fall into four themes: exploring data (Units 1–2), collecting data (Unit 3), probability and sampling distributions (Units 4–5), and statistical inference (Units 6–9). Unit 1 and the inference units carry the most weight, and the free-response section always ends with an investigative task that pulls from several units at once.
U1Exploring One-Variable Data15–23%
U2Exploring Two-Variable Data5–7%
U3Collecting Data12–15%
U4Probability, Random Variables, and Probability Distributions10–20%
U5Sampling Distributions7–12%
U6Inference for Categorical Data: Proportions12–15%
U7Inference for Quantitative Data: Means10–18%
U8Inference for Categorical Data: Chi-Square2–5%
U9Inference for Quantitative Data: Slopes2–5%
All nine units are open, 48 lessons in all. Every
lesson pairs a short explanation with worked examples and a problem
to try yourself, the same problem types that show up on the exam.
Each unit closes with a short video walk-through and a ten-problem
practice set with hidden answers.
Free preview: open any 5 lessons, or watch one unit video, without
an account. The counter on the left keeps track.
Lesson 1.1 · Unit 1 · CED topics 1.1–1.2
Individuals, variables, and types of data
Every statistics problem starts with a spreadsheet, even if you never see it.
Rows are the things being measured; columns are what's recorded about each one.
Before you pick a graph or a summary number, you have to know which kind of column
you're looking at: a mean of bus routes is nonsense.
Definitions
An individual is one object described by the data (a person, a car, a
county). A variable is a characteristic that can take different values
for different individuals. A categorical variable places each individual
into a group or label (transportation, blood type, zip code). A quantitative
variable takes numerical values for which arithmetic makes sense (height, GPA, number
of siblings). Quantitative variables are discrete if the values come from
a countable list (0, 1, 2, … siblings) and continuous if any value in an
interval is possible (hours of sleep, 7.25 or 7.251).
Worked example · Reading a data table
Here are five rows from a survey of Ridgeview High School students.
Student
Grade
Transportation
GPA
Sleep (h)
Siblings
Amara
10
Bus
3.6
7.5
2
Ben
12
Car
3.1
6.0
0
Chloe
9
Walk
3.9
8.25
1
Diego
11
Car
2.8
6.5
3
Elena
12
Bus
3.4
7.0
1
The individuals are the students. Transportation is categorical.
Grade looks numerical, but you'd never average it meaningfully: it labels a
group, so treat it as categorical. GPA and Sleep are quantitative and
continuous (any value in a range is possible). Siblings is quantitative and
discrete: you can count the possible values.
Worked example · Numbers that aren't quantitative
A dataset records each runner's bib number, finishing time, and age group (under 20,
20–39, 40+). Which variables are quantitative?
Only finishing time. Bib number is an identifier: adding two bib numbers means
nothing. Age group is categorical even though it's built from a number; once ages
are binned into labels, arithmetic is gone. Ask whether the average of the
values would mean anything. If not, it's categorical.
Exam tip: when asked to "identify the variable," name it with its units and its type:
"the quantitative variable is hours of sleep per night." Half-answers lose the point.
Try it
A used-car website lists, for each car: make, model year, mileage, color, price, and
number of previous owners. Classify each variable as categorical, quantitative
discrete, or quantitative continuous.
Show answer
Make and color are categorical. Model year is best treated as categorical (a label
for a group of cars). Mileage and price are quantitative and effectively continuous.
Number of previous owners is quantitative and discrete (0, 1, 2, …).
Lesson 1.2 · Unit 1 · CED topics 1.3–1.4
Representing and comparing categorical data
A categorical variable has no shape, center, or spread. All you can do is count how
many individuals fall into each category, and then convert those counts into
proportions so that groups of different sizes can be compared fairly.
Definitions
A frequency table lists each category with its count. A relative
frequency table lists each category with its proportion (or percent) of the
total: \(\text{relative frequency} = \dfrac{\text{count}}{\text{total}}\). The
relative frequencies of one variable always sum to 1 (allowing for rounding). A
bar chart draws one bar per category with height equal to the frequency
or relative frequency; bars don't touch. A pie chart shows the same
relative frequencies as slices and only works when the categories make up a whole.
Worked example · From counts to proportions
Fifty Ridgeview students were asked how they usually get to school.
Transportation
Car
Bus
Walk or bike
Other
Total
Frequency
20
15
10
5
50
Relative frequency
0.40
0.30
0.20
0.10
1.00
Each relative frequency is count ÷ 50: for example \(20/50 = 0.40\). Describing the
distribution: "Car was the most common way to get to school (40% of students), followed
by bus (30%); only 10% used some other method." There is no "shape" here: just which
categories are large and which are small.
Worked example · Comparing two groups
The same question was asked of 40 ninth graders and 25 seniors separately.
Group
Car
Bus
Walk or bike
Other
Total
9th grade (count)
10
20
8
2
40
9th grade (rel. freq.)
0.25
0.50
0.20
0.05
1.00
12th grade (count)
15
3
5
2
25
12th grade (rel. freq.)
0.60
0.12
0.20
0.08
1.00
Raw counts mislead here because the groups are different sizes. Compare relative
frequencies instead: 60% of seniors arrive by car compared with only 25% of
ninth graders, while half of ninth graders ride the bus versus 12% of seniors.
The proportion walking or biking is the same (20%) in both groups. A side-by-side bar
chart with relative frequency on the vertical axis shows this at a glance.
Exam tip: with groups of different sizes, graders expect percentages, not counts, and
comparative words; "higher than," "the same as," "roughly twice."
Try it
In a survey of 80 students, 32 said they prefer texting, 28 prefer voice calls, 12
prefer video calls, and the rest chose "other." Build the relative frequency table and
write one sentence describing the distribution.
Show answer
Other has \(80 - 72 = 8\) students. Relative frequencies: texting \(32/80 = 0.40\),
voice \(28/80 = 0.35\), video \(12/80 = 0.15\), other \(8/80 = 0.10\) (sum 1.00).
Texting was the most popular choice (40%), voice calls were close behind (35%), and
video calls and other methods together accounted for only a quarter of students.
Lesson 1.3 · Unit 1 · CED topics 1.5–1.6
Graphs for quantitative data and describing shape
A quantitative variable does have a shape, and you can only see it in a picture. Three
graphs do the job: a dotplot (one dot per value), a stemplot (digits sorted into rows),
and a histogram (values grouped into equal-width bins).
Shape vocabulary
A distribution is symmetric if its left and right halves are roughly
mirror images, skewed right if the longer tail points toward larger
values, and skewed left if the longer tail points toward smaller values.
It is unimodal with one clear peak, bimodal with two, and
roughly uniform if all values are about equally common. Also report
gaps (empty stretches), clusters (groups of values), and
any apparent outliers: values far from the rest.
Worked example · Stemplot and histogram
Resting pulse rates (beats per minute) of 20 students, sorted:
58
62
64
66
68
68
70
70
72
72
72
74
74
76
78
78
80
84
88
96
A stemplot uses the tens digit as the stem and the ones digit as the leaf:
Stem
Leaves
5
8
6
2 4 6 8 8
7
0 0 2 2 2 4 4 6 8 8
8
0 4 8
9
6
A histogram with bins of width 10 has heights 1, 5, 10, 3, 1 for the bins 50–59,
60–69, 70–79, 80–89, 90–99. Both pictures show the same thing: a unimodal
distribution peaking in the 70s, roughly symmetric with a slightly longer right tail,
centered near 72–74 bpm, and one possible high outlier at 96. A key that
says "7 | 2 means 72 bpm" is required on a stemplot.
Worked example · Choosing bin width
With bins of width 2 the pulse data would spread over 20 bins, most holding zero or
one value: pure noise. One bin of width 50 gives a single block, no information. Aim
for 5 to 10 equal-width bins, and be consistent about which bin a boundary value joins.
Try it
Number of books 20 students read last year: 0, 1, 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 5,
5, 6, 7, 8, 10, 12, 20. Sketch a dotplot mentally and describe the shape, including
any gaps or outliers.
Show answer
The distribution is skewed right and unimodal, with a peak at 3 books. Most students
read 0–8 books; there is a gap between 12 and 20, and the value 20 is an apparent
outlier. (Center is around 3–4 books; the long right tail pulls the mean, 5.05, above
the median, 3.5: a preview of the next lesson.)
Lesson 1.4 · Unit 1 · CED topic 1.7
Measures of center and spread
Once you can see a distribution, you want two numbers for it: where it sits (center)
and how wide it is (spread). Each comes in a resistant version that ignores extreme
values and a non-resistant version that uses every data point.
Formulas
Mean: \(\bar{x} = \dfrac{\sum x_i}{n}\). Median: the middle
value when the data are ordered (average the two middle values if \(n\) is even).
Range: max − min. Interquartile range:
\(\text{IQR} = Q_3 - Q_1\), where \(Q_1\) and \(Q_3\) are the medians of the lower and
upper halves. Sample standard deviation:
\[s_x = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n - 1}},\]
roughly the typical distance of a value from the mean; its square is the
variance \(s_x^2\). Divide by \(n - 1\), not \(n\).
Worked example · Computing everything by hand
Hours of sleep last night for eight students: 5, 6, 6.5, 7, 7, 7.5, 8, 9.
Mean: \(\bar{x} = 56/8 = 7\) hours. Median: the middle two are 7 and 7, so 7 hours.
Lower half 5, 6, 6.5, 7 gives \(Q_1 = 6.25\); upper half 7, 7.5, 8, 9 gives
\(Q_3 = 7.75\); IQR = 1.5 hours. Range = 4 hours. For the SD, the squared deviations
\((x_i - 7)^2\) are 4, 1, 0.25, 0, 0, 0.25, 1, 4, which sum to 10.5:
\[s_x = \sqrt{\frac{10.5}{7}} = \sqrt{1.5} \approx 1.22 \text{ hours}.\]
A typical student's sleep was about 1.22 hours from the mean.
Calculator
Enter the data in L1, then STAT → CALC → 1-Var Stats L1. Read \(\bar{x}\),
\(Sx\) (the sample SD with \(n - 1\): use this one, not \(\sigma x\)), and scroll down for
the five-number summary.
Worked example · Resistance and the outlier rule
Suppose the 9 was really 14 hours. The mean jumps to \(61/8 = 7.625\) hours and the SD
to about 2.74, but the median stays 7 and the IQR stays 1.5. Median and IQR are
resistant; mean and SD are not. Is 14 an outlier? The fences are
\(Q_1 - 1.5\cdot\text{IQR} = 6.25 - 2.25 = 4.0\) and
\(Q_3 + 1.5\cdot\text{IQR} = 7.75 + 2.25 = 10.0\); since 14 > 10, it is an outlier
by the 1.5×IQR rule. For skewed data or data with outliers, report median and IQR.
Transformations. Adding a constant shifts center and position by that
constant and leaves spread unchanged (everyone sleeps 30 minutes more: mean 7.5, SD
still 1.22). Multiplying by \(k\) multiplies center and spread by \(k\): in minutes,
the mean is 420 and the SD is \(60 \times 1.2247 \approx 73.5\).
Try it
Compute the mean and sample standard deviation of 4, 8, 6, 5, 12 by hand.
A boxplot compresses a distribution into five numbers, which makes it the best tool for
putting two groups side by side. The exam's favorite Unit 1 question is exactly that:
"Compare the distributions." Your answer needs shape, outliers, center, and spread,
each stated with comparative language and in context.
Method
The five-number summary is minimum, \(Q_1\), median, \(Q_3\), maximum. A
boxplot draws a box from \(Q_1\) to \(Q_3\) with a line at the median.
Whiskers extend to the smallest and largest values that are not outliers; any
point beyond the fences \(Q_1 - 1.5\,\text{IQR}\) or \(Q_3 + 1.5\,\text{IQR}\) is plotted
separately as an outlier. A boxplot hides modes and gaps, so never call a distribution
"unimodal" from a boxplot alone.
Worked example · Two groups, four sentences
Commute times (minutes) for 12 students who drive and 12 who ride the bus:
Drive
5
8
10
10
12
14
15
15
18
20
22
45
Bus
15
18
20
22
25
25
28
30
30
32
35
40
Group
Min
Q1
Median
Q3
Max
IQR
Upper fence
Drive
5
10
14.5
19
45
9
32.5
Bus
15
21
26.5
31
40
10
46
For the drivers, \(Q_1\) is the median of the lower six values (10) and \(Q_3\) the
median of the upper six (19). The upper fence is \(19 + 1.5(9) = 32.5\), so the
45-minute commute is an outlier and the whisker stops at 22. The bus group has none.
Comparison:Shape: the drive distribution is skewed right
(a long upper tail ending in a high outlier, and a mean of 16.2 min above the median of
14.5), while the bus distribution is roughly symmetric. Outliers: one driver at 45 minutes; the
bus group has none. Center: the median commute for bus riders (26.5 min) is
higher than for drivers (14.5 min): bus riders typically take about 12 minutes longer.
Spread: the IQRs are similar (10 min for bus, 9 min for drive), so the middle
halves are about equally variable, though the drivers' range is larger because of the
outlier.
Exam tip: every sentence must name both groups and use a comparison word ("higher
than," "similar to," "more variable than"). Listing each group's statistics without
comparing them earns partial credit at best. Use median and IQR when either group is
skewed or has outliers.
Try it
Nine quiz scores: 3, 7, 8, 9, 10, 11, 12, 15, 26. Find the five-number summary and
determine whether any value is an outlier.
Show answer
Median = 10 (the fifth value). Lower half 3, 7, 8, 9 gives \(Q_1 = 7.5\); upper half
11, 12, 15, 26 gives \(Q_3 = 13.5\). Five-number summary: 3, 7.5, 10, 13.5, 26.
IQR = 6; fences \(7.5 - 9 = -1.5\) and \(13.5 + 9 = 22.5\). Since 26 > 22.5, the
score of 26 is an outlier; 3 is not.
Lesson 1.6 · Unit 1 · CED topic 1.10
The normal distribution: z-scores, percentiles, and the empirical rule
Many measurements (heights, test scores, measurement errors) pile up in a symmetric,
bell-shaped pattern. When a distribution is approximately normal, the mean and standard
deviation tell you everything, and one calculation converts any value into a percentile.
Definitions
The z-score of a value is \(z = \dfrac{x - \mu}{\sigma}\): how many standard
deviations \(x\) lies above (positive) or below (negative) the mean. A value's
percentile is the percent of the distribution at or below it. A normal
distribution is written \(N(\mu, \sigma)\) and obeys the 68–95–99.7 rule:
about 68% of values lie within \(1\sigma\) of the mean, 95% within \(2\sigma\), 99.7%
within \(3\sigma\).
Worked example · Finding a proportion
Heights of adult women are approximately \(N(64.5, 2.5)\) inches. What proportion are
shorter than 60 inches? Between 62 and 68 inches?
\(z = \dfrac{60 - 64.5}{2.5} = -1.8\). The area to the left of \(z = -1.8\) is
0.0359, so about 3.6% of women are shorter than 60 inches. For the
interval, the z-scores are \(\dfrac{62 - 64.5}{2.5} = -1.0\) and \(\dfrac{68 - 64.5}{2.5} = 1.4\),
and the area between them is \(0.9192 - 0.1587 = \) 0.7605, about 76%
(the calculator, carrying more decimals, gives 0.7606).
Sanity check: the empirical rule puts 68% between 62 and 67 inches, so 76% for a
slightly wider interval is reasonable.
Worked example · Finding a cutoff
How tall must a woman be to be in the tallest 10%?
We need the value with area 0.90 to its left; that z-score is \(z = 1.2816\).
Unstandardize: \(x = \mu + z\sigma = 64.5 + 1.2816(2.5) \approx\)
67.7 inches. Say which area you're using, "area to the left", so
the grader knows the sign wasn't luck.
Calculator
normalcdf(lower, upper, μ, σ) gives the area between two values; use
−1E99 or 1E99 for an open tail. Here: normalcdf(−1E99, 60, 64.5, 2.5) = 0.0359 and
normalcdf(62, 68, 64.5, 2.5) = 0.7606. invNorm(area, μ, σ) returns the
value with that area to its left: invNorm(0.90, 64.5, 2.5) = 67.70. On the exam,
show the z-score and a shaded sketch too: a bare "normalcdf" earns no credit.
Assessing normality. Before using the normal model, check that a
histogram is roughly symmetric and bell-shaped with no outliers and that about 68% of
the data fall within one SD of the mean. Skewed data (income, commute times) don't
belong in normalcdf.
Try it
Scores on a chemistry test are approximately \(N(72, 8)\). (a) What percent of
students scored above 85? (b) What score marks the 25th percentile?
Show answer
(a) \(z = (85 - 72)/8 = 1.625\); area to the right is \(1 - 0.9479 = 0.0521\), about
5.2% (normalcdf(85, 1E99, 72, 8)). (b) invNorm(0.25) gives \(z = -0.6745\), so
\(x = 72 - 0.6745(8) \approx 66.6\) points.
Unit 1 practice · 10 problems
Unit 1 practice: Exploring One-Variable Data
Ten problems covering the whole unit, in roughly exam order. Work each one on
paper before revealing the answer: the reveal shows the key steps, not just the
number. A calculator is fine for problems 4, 8, and 9; show the z-score and a
sketch anyway.
A fitness app records, for each user: user ID number, age (years), average daily steps, hours of sleep per night, favorite activity (run, bike, swim, or walk), and number of workouts last week. Identify the individuals and classify each variable as categorical, quantitative discrete, or quantitative continuous.
Show answer
The individuals are the app's users. User ID is categorical (an identifier: averaging IDs means nothing). Favorite activity is categorical. Age and hours of sleep are quantitative and continuous (any value in a range is possible). Average daily steps and number of workouts last week are quantitative and discrete: they are counts. The test: if the average of the values would mean something, the variable is quantitative.
Sixty freshmen and ninety seniors were asked which type of club they belong to.
Group
Sports
Arts
STEM
None
Total
Freshmen
27
15
12
6
60
Seniors
30
18
27
15
90
A student says, "More seniors than freshmen are in sports clubs, so sports are more popular with seniors." Compute the relative frequencies for each group and write two sentences comparing the distributions.
Show answer
Freshmen: sports \(27/60 = 0.45\), arts \(0.25\), STEM \(0.20\), none \(0.10\). Seniors: sports \(30/90 = 0.333\), arts \(0.20\), STEM \(0.30\), none \(0.167\). The student is wrong: the groups differ in size, so compare proportions, not counts. Sports is the most common club type in both grades, but a higher proportion of freshmen (45%) than seniors (33%) belong to one. STEM clubs are more popular among seniors (30% vs. 20%), and seniors are more likely to belong to no club at all (17% vs. 10%).
A histogram of the sale prices (thousands of dollars) of 40 homes in a town has bars of height 14, 12, 7, 4, 2, and 1 over the intervals 100–200, 200–300, 300–400, 400–500, 500–600, and 600–700. Describe the shape of the distribution, state which interval contains the median, and explain whether the mean is greater than, less than, or about equal to the median.
Show answer
The distribution is unimodal and skewed right: the peak is in the 100–200 interval and the frequencies trail off in a long tail toward the expensive homes, with no gaps. The first bar holds 14 homes and the first two hold 26, so the 20th and 21st values, the median for \(n = 40\), fall in the 200–300 interval. The mean is greater than the median: the long right tail pulls the mean toward the high values (using bar midpoints, the mean is roughly \(278\) thousand dollars), while the median is resistant.
The numbers of emails five employees received yesterday were 8, 11, 12, 15, and 19. By hand, find the mean, median, range, IQR, and sample standard deviation, and interpret the standard deviation in context.
Show answer
Mean \(\bar x = 65/5 = 13\) emails; median = 12 (the third value). Range \(= 19 - 8 = 11\). Lower half 8, 11 gives \(Q_1 = 9.5\); upper half 15, 19 gives \(Q_3 = 17\); IQR \(= 7.5\) emails. Deviations from 13 are \(-5, -2, -1, 2, 6\); squares 25, 4, 1, 4, 36 sum to 70, so \[s_x = \sqrt{\frac{70}{4}} = \sqrt{17.5} \approx 4.18 \text{ emails}.\] An employee's email count typically differs from the mean of 13 by about 4.18 emails.
Wait times (minutes) for 11 patients at a clinic, in order: 14, 18, 20, 21, 22, 23, 25, 26, 28, 31, 45. Find the five-number summary, use the 1.5×IQR rule to check for outliers, and describe the boxplot you would draw.
Show answer
Median = 23 (the 6th value). Lower half 14, 18, 20, 21, 22 gives \(Q_1 = 20\); upper half 25, 26, 28, 31, 45 gives \(Q_3 = 28\). Five-number summary: 14, 20, 23, 28, 45. IQR \(= 8\); fences \(20 - 1.5(8) = 8\) and \(28 + 1.5(8) = 40\). Since \(45 \gt 40\), the 45-minute wait is an outlier; 14 is not. The boxplot has a box from 20 to 28 with a line at 23, a lower whisker to 14, an upper whisker to 31 (the largest non-outlier), and a separate point at 45.
Daily high temperatures in a city over a month have mean 20 °C, median 19 °C, standard deviation 4 °C, and IQR 6 °C. (a) Find the mean, median, standard deviation, and IQR in degrees Fahrenheit, using \(F = 1.8C + 32\). (b) A thermometer was found to read 2 °C too low; every reading is corrected by adding 2 °C. What are the corrected mean and standard deviation in °C?
Show answer
(a) Multiplying by 1.8 scales both center and spread; adding 32 shifts center only. Mean \(= 1.8(20) + 32 = 68\) °F; median \(= 1.8(19) + 32 = 66.2\) °F; SD \(= 1.8(4) = 7.2\) °F; IQR \(= 1.8(6) = 10.8\) °F. (b) Adding a constant shifts every value and the center by 2 but leaves spread unchanged: mean 22 °C, SD still 4 °C.
Final exam scores for two sections of the same course are summarized by five-number summaries. Section A (in person): 52, 64, 72, 80, 96. Section B (online): 44, 58, 66, 84, 100. Neither section has outliers. Compare the two distributions.
Show answer
Shape: Section A is roughly symmetric (the median sits in the middle of the box and the whiskers are similar in length), while Section B is skewed right: the upper half of its box (66 to 84) and upper whisker are longer than the lower ones. Outliers: neither section has any. Center: the median score in Section A (72) is higher than in Section B (66). Spread: Section B's scores are more variable than Section A's: an IQR of 26 points versus 16, and a range of 56 versus 44. Every sentence names both groups and uses a comparison word; that is what the rubric rewards.
Composite ACT scores are approximately \(N(21, 5.4)\) and SAT total scores are approximately \(N(1050, 200)\). Priya scored 30 on the ACT; Marcus scored 1350 on the SAT. Who did relatively better? Use z-scores and percentiles to justify.
Show answer
Priya: \(z = \dfrac{30 - 21}{5.4} \approx 1.67\); Marcus: \(z = \dfrac{1350 - 1050}{200} = 1.50\). Priya's score is 1.67 standard deviations above the mean versus 1.50 for Marcus, so Priya did relatively better. Her score is at about the 95th percentile (area to the left of 1.67 is 0.952); his is at about the 93rd (area 0.933).
Finishing times in a large 5K race are approximately normal with mean 28 minutes and standard deviation 5 minutes. (a) What proportion of runners finish in under 20 minutes? (b) What proportion finish between 25 and 35 minutes? (c) How fast must a runner be to finish in the fastest 5%?
Show answer
(a) \(z = \dfrac{20 - 28}{5} = -1.6\); area to the left is \(0.0548\), so about 5.5% finish under 20 minutes (normalcdf(−1E99, 20, 28, 5)). (b) \(z = -0.6\) and \(z = 1.4\); area between is \(0.9192 - 0.2743 = 0.6449\), about 64%. (c) The fastest 5% have area 0.05 to the left: \(z = -1.645\), so \(x = 28 - 1.645(5) \approx 19.8\) minutes (invNorm(0.05, 28, 5) = 19.78). Runners must finish in under about 19.8 minutes.
IQ scores are approximately \(N(100, 15)\). Using the 68–95–99.7 rule, approximately what percent of people have IQ scores between 85 and 130?
47.5% is only the piece from 100 to 130 (half of the 95% within two standard deviations). It leaves out the 34% between 85 and the mean, so it counts just one side of the interval.
68% is the percent within one standard deviation of the mean, 85 to 115, but this interval reaches 130, two standard deviations above the mean. Because the upper half is wider than the lower half, the answer must fall between 68% and 95%.
From 85 to 100 is one standard deviation below the mean, holding half of 68%, or 34%; from 100 to 130 is two standard deviations above the mean, holding half of 95%, or 47.5%. Together, \(34\% + 47.5\% = 81.5\%\) (the exact normal area is 0.8186).
95% is the percent within two standard deviations on both sides, 70 to 130. The interval here starts at 85, not 70, so the 13.5% between 70 and 85 has to be removed.
Lesson 2.1 · Unit 2 · CED topics 2.1–2.3
Two categorical variables: two-way tables and conditional distributions
Unit 1 handled one variable at a time. Now the question is whether knowing one thing
about an individual tells you anything about another. For two categorical variables the
tool is a two-way table, and the trick is to read it by rows or columns, not by raw
cell counts.
Definitions
A two-way table counts individuals by two categorical variables at once.
The marginal distribution of one variable uses the totals in the margins:
\(\dfrac{\text{row or column total}}{\text{grand total}}\). A conditional
distribution restricts attention to one row (or column) and divides each cell by
that row's total. Two variables have an association if the conditional
distributions differ from group to group: knowing the value of one variable changes
the probabilities for the other. If the conditional distributions are (nearly) the
same, there is no association.
Worked example · Reading the table three ways
Two hundred students were asked whether they hold a part-time job.
Grade
Has a job
No job
Total
10th
20
60
80
12th
66
54
120
Total
86
114
200
Marginal: overall, \(86/200 = 0.43\) of students have a job. Conditional on
grade: among 10th graders, \(20/80 = 0.25\) have a job; among 12th graders,
\(66/120 = 0.55\) do. Conditional the other way: among students with jobs,
\(66/86 \approx 0.767\) are seniors. Which conditional distribution you want depends on
the question: "of the seniors, what proportion…" means divide by the senior row
total, 120.
Is there an association? Yes. The proportion with a job is much higher
for 12th graders (55%) than for 10th graders (25%), so grade level and job status are
associated in this sample. A segmented bar chart shows this: one bar per
grade, each split into "job" and "no job" segments by percent; here the segments would
differ sharply. Identical bars would mean no association.
Exam tip: to argue for an association, quote two conditional proportions and say they
differ. To argue against it, show they are about equal. Never argue from raw counts when
the groups have different sizes.
Try it
Of 60 boys surveyed, 18 prefer cats and 42 prefer dogs. Of 40 girls, 12 prefer cats and
28 prefer dogs. Is there an association between gender and pet preference? Justify
with conditional relative frequencies.
Show answer
The proportion preferring cats is \(18/60 = 0.30\) for boys and \(12/40 = 0.30\) for
girls. The conditional distributions are identical, so there is no association
between gender and pet preference in this sample. (Raw counts, 18 vs. 12, would have
misled you.)
Lesson 2.2 · Unit 2 · CED topics 2.4–2.5
Scatterplots and correlation
For two quantitative variables, the picture is a scatterplot: the explanatory variable
on the horizontal axis, the response on the vertical, one point per individual. You
describe it in four words (form, direction, strength, unusual features), and then, if
the form is linear, one number summarizes the strength.
Definition
The correlation \(r\) measures the direction and strength of a
linear relationship:
\[r = \frac{1}{n-1}\sum \left(\frac{x_i - \bar{x}}{s_x}\right)\left(\frac{y_i - \bar{y}}{s_y}\right).\]
Properties: \(-1 \le r \le 1\); the sign gives direction; \(|r|\) near 1 means the
points hug a line. \(r\) has no units and doesn't change if you swap \(x\) and \(y\) or
change units (hours to minutes). It measures only linear association, a perfect
curve can have \(r\) near 0, and it is not resistant: one outlier can move it a lot.
Worked example · Describing a scatterplot
Hours studied and score on a unit test for eight students:
Hours (x)
1
2
2
3
4
5
6
7
Score (y)
62
68
65
74
78
82
85
91
Plotting these shows a linear form, positive direction,
and strong association, students who studied longer tended to score
higher, with no outliers or clusters. Summary statistics: \(\bar{x} = 3.75\),
\(s_x = 2.121\), \(\bar{y} = 75.625\), \(s_y = 10.211\). The sum of the products of
z-scores is 6.936, so \(r = 6.936/7 \approx 0.991\): a very strong positive linear
association, matching the picture.
Calculator
Enter x in L1 and y in L2. Turn on diagnostics once (2nd → CATALOG → DiagnosticOn), then
STAT → CALC → LinReg(a+bx) L1, L2. The screen shows \(r\) and \(r^2\) along
with the line you'll meet in the next lesson.
Worked example · What r doesn't say
Across U.S. cities, monthly ice cream sales and drowning deaths have a strong positive
correlation. Does ice cream cause drowning? No: hot weather drives both. Temperature is
a lurking variable, and this is the rule you'll write on every exam:
correlation does not imply causation. Even \(r = 0.99\) for hours and
scores doesn't prove that studying raised the scores; only a randomized experiment can.
Try it
(a) If the study times above are converted to minutes, what happens to \(r\)? (b) A
scatterplot of \(y = x^2\) for \(x = -3, -2, \ldots, 3\) is a perfect parabola. Is \(r\)
near 1?
Show answer
(a) Nothing: \(r\) is unchanged by a change of units, still 0.991. (b) No. The
relationship is perfectly curved but not linear; by symmetry \(r = 0\) exactly. A
correlation near 0 means no linear association, not no association.
Lesson 2.3 · Unit 2 · CED topics 2.6–2.7
The least-squares regression line
When a scatterplot is linear, you want the line that fits it best so you can predict
\(y\) from \(x\). "Best" has a precise meaning: the line that makes the sum of the
squared vertical distances from the points as small as possible.
Formulas
The least-squares regression line (LSRL) is \(\hat{y} = a + bx\), where
\(\hat{y}\) is the predicted response. Its slope and intercept are
\[b = r\,\frac{s_y}{s_x}, \qquad a = \bar{y} - b\bar{x},\]
so the line always passes through \((\bar{x}, \bar{y})\). Slope: for each
additional unit of \(x\), the predicted \(y\) changes by \(b\) units. Intercept:
the predicted \(y\) when \(x = 0\) (meaningful only if \(x = 0\) makes sense and is
near the data).
Worked example · Building and interpreting the line
For the hours-and-scores data (\(\bar{x} = 3.75\), \(s_x = 2.121\), \(\bar{y} = 75.625\),
\(s_y = 10.211\), \(r = 0.991\)):
\[b = 0.9909\cdot\frac{10.211}{2.121} \approx 4.77, \qquad a = 75.625 - 4.77(3.75) \approx 57.74.\]
So \(\widehat{\text{score}} = 57.74 + 4.77(\text{hours})\).
Slope in context: for each additional hour studied, the predicted test
score increases by about 4.77 points. Intercept in context: a student who
studied 0 hours is predicted to score about 57.7 points: plausible here, since 0 is
just below the smallest observed value of 1 hour. Always say "predicted": the line
describes the average pattern, not any one student.
Worked example · Prediction and extrapolation
Predict the score for 3.5 hours: \(\hat{y} = 57.74 + 4.77(3.5) \approx 74.4\) points.
Now try 20 hours: \(\hat{y} = 57.74 + 4.77(20) \approx 153\): above the maximum possible
score. Predicting far outside the range of the data (1 to 7 hours) is
extrapolation, and the exam wants you to name it as unreliable.
Calculator
With x in L1 and y in L2: STAT → CALC → LinReg(a+bx) L1, L2, Y1 (the Y1
stores the equation for graphing and predicting). Report the equation with variable
names, not \(x\) and \(y\): \(\widehat{\text{score}} = 57.74 + 4.77(\text{hours})\).
Try it
A regression of weekly sales (dollars) on advertising spend (dollars) has
\(\bar{x} = 10\), \(s_x = 2\), \(\bar{y} = 50\), \(s_y = 6\), \(r = 0.80\). Find the LSRL and
interpret its slope.
Show answer
\(b = 0.80(6/2) = 2.4\); \(a = 50 - 2.4(10) = 26\). \(\widehat{\text{sales}} = 26 + 2.4(\text{ad spend})\).
For each additional dollar spent on advertising, predicted weekly sales increase by
about $2.40.
Lesson 2.4 · Unit 2 · CED topic 2.8
Residuals, residual plots, r², and s
A regression line is a model, and every model needs a report card. Residuals tell you
how far each point misses; a residual plot tells you whether a line was the right shape
at all; and two numbers, \(r^2\) and \(s\), grade the fit overall.
Definitions
Residual = actual − predicted = \(y - \hat{y}\). A positive residual
means the line underpredicted that point. A residual plot graphs residuals
against \(x\) (or against \(\hat{y}\)); a linear model is appropriate when the plot shows
no pattern: random scatter around 0. A curve in the residual plot means
the relationship isn't linear. The coefficient of determination
\(r^2\) is the fraction of the variation in \(y\) that is explained by the LSRL on \(x\).
The standard deviation of the residuals
\[s = \sqrt{\frac{\sum (y_i - \hat{y}_i)^2}{n - 2}}\]
is the typical size of a prediction error, in the units of \(y\).
Worked example · Residuals for the study data
Using \(\hat{y} = 57.74 + 4.77x\):
Hours
1
2
2
3
4
5
6
7
Score
62
68
65
74
78
82
85
91
Predicted
62.51
67.28
67.28
72.05
76.82
81.59
86.36
91.13
Residual
−0.51
0.72
−2.28
1.95
1.18
0.41
−1.36
−0.13
The student who studied 3 hours scored 74, but the line predicted 72.05; the residual
of \(+1.95\) means the model underpredicted by about 2 points. The residuals bounce
above and below zero with no curve, so a linear model is appropriate. Their squares sum
to 13.21, giving \(s = \sqrt{13.21/6} \approx 1.48\) points.
Worked example · Reading computer output
Predictor
Coef
SE Coef
T
P
Constant
57.738
1.121
51.48
0.000
Hours
4.770
0.264
18.04
0.000
S = 1.484 R-Sq = 98.2%
Read the equation from the Coef column: \(\widehat{\text{score}} = 57.738 + 4.770(\text{hours})\).
\(r^2\) in context: about 98.2% of the variation in test scores is explained
by the linear relationship with hours studied. \(s\) in context: actual
scores are typically about 1.48 points away from the scores predicted by the line. And
\(r = +\sqrt{0.982} \approx 0.991\): take the sign from the slope.
Try it
Output for predicting a lemonade stand's daily cups sold from the high temperature
(°F): Constant −38.20, Temp 1.450, S = 6.10, R-Sq = 72.3%. Write the equation and
interpret \(r^2\) and \(s\) in context.
Show answer
\(\widehat{\text{cups}} = -38.20 + 1.450(\text{temp})\). About 72.3% of the variation in
daily cups sold is explained by the linear relationship with temperature. Actual sales
are typically about 6.1 cups from the predicted sales. (Also \(r = +\sqrt{0.723} \approx 0.85\).)
Lesson 2.5 · Unit 2 · CED topic 2.9
Influential points and transforming to achieve linearity
Two things break a regression: a single point that drags the line around, and a
relationship that was never straight to begin with. You need to recognize both from a
plot.
Definitions
An outlier in regression is a point with an unusually large residual: far
from the line in the \(y\) direction. A high-leverage point has an \(x\)
value far from \(\bar{x}\). A point is influential if removing it would
substantially change the slope, intercept, or \(r\). High-leverage points are usually
influential; a \(y\)-outlier near \(\bar{x}\) mostly changes \(r\) and \(s\).
Worked example · One point, three fates
Start with the study data (\(\hat{y} = 57.74 + 4.77x\), \(r = 0.991\)) and add one
ninth student.
(4 h, 60 points): slope barely moves, to 4.65, but \(r\) drops to 0.85. An outlier
with little leverage: it hurts the fit, not the direction.
(12 h, 110 points), on the trend: slope 4.37, \(r = 0.994\). High leverage but not
very influential, because it agrees with the pattern.
(12 h, 60 points): slope collapses to 0.39 and \(r\) to 0.12. High leverage
and off the trend: strongly influential.
Describe influence by naming what changes: "removing this point would make the slope
steeper and \(r\) closer to 1."
Method · Transforming
If the residual plot is curved, transform a variable and refit. Exponential growth
(\(y = ab^x\)) becomes linear as \(\log y\) versus \(x\); a power relationship
(\(y = ax^p\)) becomes linear as \(\log y\) versus \(\log x\). Choose the model with a
patternless residual plot and higher \(r^2\), then undo the log to predict.
Worked example · Log transformation
A bacteria culture is counted every hour:
Hours
0
1
2
3
4
5
Count
100
210
390
820
1600
3300
log(Count)
2.000
2.322
2.591
2.914
3.204
3.519
A line on the raw data has \(r^2 = 0.81\) and residuals 501, 23, −386, −544, −353,
759: a clear U shape, so linear is wrong. Each count is roughly double the last:
exponential growth. Regressing \(\log(\text{count})\) on hours gives
\[\widehat{\log(\text{count})} = 2.004 + 0.302(\text{hours}), \qquad r^2 = 0.9996,\]
with patternless residuals. At 7 hours: \(\log \hat{y} = 2.004 + 0.302(7) = 4.118\), so
\(\hat{y} = 10^{4.118} \approx 13{,}100\) bacteria. (Since \(10^{0.302} \approx 2.0\), the
model says the population doubles each hour.)
Try it
A scatterplot of car value against age curves downward and flattens, and the linear
residual plot is U-shaped. A fit of \(\log(\text{value})\) on age gives
\(\widehat{\log(\text{value})} = 4.42 - 0.071(\text{age})\). Predict the value of a
6-year-old car, and explain why the log model was chosen.
Show answer
\(\log \hat{y} = 4.42 - 0.071(6) = 3.994\), so \(\hat{y} = 10^{3.994} \approx \$9{,}860\).
The log model was chosen because the linear model's residual plot was curved, while
the transformed data are linear with a patternless residual plot: the
exponential-decay shape of depreciation.
Unit 2 practice · 10 problems
Unit 2 practice: Exploring Two-Variable Data
Ten problems covering the whole unit, in roughly exam order. Work each one on
paper before revealing the answer: the reveal shows the key steps, not just the
number. Use LinReg(a+bx) for problem 5; everything else is by hand.
A counselor asked 100 juniors and 120 seniors about their plans after graduation.
Grade
Four-year college
Work or trade school
Gap year
Total
Juniors
55
25
20
100
Seniors
84
24
12
120
Total
139
49
32
220
(a) What proportion of all students surveyed plan to attend a four-year college? (b) Give the conditional distribution of plans for each grade. (c) Is there an association between grade and plans? Justify.
Show answer
(a) Marginal: \(139/220 \approx 0.632\). (b) Juniors: college \(55/100 = 0.55\), work \(0.25\), gap year \(0.20\). Seniors: college \(84/120 = 0.70\), work \(0.20\), gap year \(0.10\). (c) Yes. The conditional distributions differ: 70% of seniors plan on a four-year college versus 55% of juniors, and juniors are twice as likely to plan a gap year (20% vs. 10%). Knowing a student's grade changes the likely plan, so the variables are associated in this sample. A segmented bar chart would show the two bars split differently.
A scatterplot of asking price (dollars) against age (years) for 30 used cars shows points running from upper left to lower right in a fairly narrow, straight band, with one 3-year-old car priced far above the band; \(r = -0.87\). (a) Describe the association. (b) If price were recorded in thousands of dollars instead, what would \(r\) be? (c) A classmate says \(r = -0.87\) proves the relationship is strong and linear. What's wrong with that?
Show answer
(a) There is a strong, negative, linear association between age and asking price: older cars tend to have lower prices. One unusual point, a 3-year-old car priced well above the trend, is an outlier in the \(y\) direction. (b) Still \(-0.87\); correlation has no units and is unchanged by a change of units in either variable. (c) \(r\) measures the strength of a linear association only: it cannot tell you the form. A strongly curved pattern can also produce \(|r|\) near 0.87, so you must look at the scatterplot (or a residual plot) to judge linearity.
For 40 test plots, the amount of fertilizer applied (kg per hectare) has \(\bar x = 40\) and \(s_x = 12\); crop yield (tonnes per hectare) has \(\bar y = 6.5\) and \(s_y = 0.9\); the correlation is \(r = 0.85\). Find the least-squares regression line, interpret its slope in context, and predict the yield for a plot that received 55 kg/ha.
Show answer
\(b = r\dfrac{s_y}{s_x} = 0.85\cdot\dfrac{0.9}{12} = 0.06375\) and \(a = \bar y - b\bar x = 6.5 - 0.06375(40) = 3.95\), so \(\widehat{\text{yield}} = 3.95 + 0.0638(\text{fertilizer})\). Slope: for each additional kilogram of fertilizer per hectare, the predicted yield increases by about 0.064 tonnes per hectare. At 55 kg/ha: \(\hat y = 3.95 + 0.06375(55) \approx 7.46\) tonnes per hectare. (55 is within about 1.25 SDs of \(\bar x\), so this is not an extrapolation.)
The regression line for predicting a music student's audition score from weekly practice hours is \(\widehat{\text{score}} = 12.5 + 3.2(\text{hours})\). One student practiced 6 hours per week and scored 30. Find and interpret the residual.
Show answer
Predicted score \(= 12.5 + 3.2(6) = 31.7\). Residual \(= \text{actual} - \text{predicted} = 30 - 31.7 = -1.7\). The student scored 1.7 points lower than the line predicted for someone practicing 6 hours per week; the model overpredicted this student's score. A negative residual means the point lies below the line.
The ages (years) and resale values (hundreds of dollars) of six laptops of the same model:
Age (x)
2
4
5
7
8
10
Value (y)
31
27
26
20
19
14
Use your calculator to find the LSRL, \(r\), and \(r^2\). Then predict the value of a 6-year-old laptop and find the residual for the 5-year-old laptop.
Show answer
LinReg(a+bx) with age in L1 and value in L2: \(\widehat{\text{value}} = 35.69 - 2.143(\text{age})\), \(r = -0.995\), \(r^2 = 0.990\). Each additional year of age is associated with a predicted drop of about $214 in resale value. At age 6: \(\hat y = 35.69 - 2.143(6) \approx 22.83\), or about $2,280. For the 5-year-old laptop: predicted \(35.69 - 2.143(5) = 24.98\), so the residual is \(26 - 24.98 = +1.02\) hundred dollars: it sold for about $100 more than the line predicted.
Software output for predicting a home's monthly heating cost (dollars) from the month's average outside temperature (°F), based on 18 months:
Predictor
Coef
SE Coef
T
P
Constant
210.40
9.80
21.47
0.000
Temp
−3.850
0.420
−9.17
0.000
S = 14.2 R-Sq = 84.6%
Write the equation, find \(r\), and interpret the slope, \(r^2\), and \(s\) in context. Predict the heating cost for a month averaging 30 °F.
Show answer
\(\widehat{\text{cost}} = 210.40 - 3.85(\text{temp})\). Since the slope is negative, \(r = -\sqrt{0.846} \approx -0.92\). Slope: for each 1 °F increase in average outside temperature, the predicted monthly heating cost decreases by about $3.85. \(r^2\): about 84.6% of the variation in monthly heating cost is explained by the linear relationship with outside temperature. \(s\): actual heating costs are typically about $14.20 away from the cost predicted by the line. At 30 °F: \(210.40 - 3.85(30) = 94.90\), so about $95.
Two residual plots are described. Plot I: residuals are negative for the smallest and largest \(x\)-values and positive in the middle, forming an arch. Plot II: residuals scatter randomly above and below zero with no visible pattern. For each, state whether a linear model is appropriate and explain what the plot tells you.
Show answer
Plot I: a linear model is not appropriate. The arch means the line systematically overpredicts at the ends and underpredicts in the middle: the true relationship is curved, and a transformation or nonlinear model should be considered. Plot II: a linear model is appropriate. Random scatter about zero means the line has captured the pattern and what remains is unexplained noise. A residual plot magnifies departures from linearity that are hard to see on the scatterplot itself.
A scatterplot of weekly earnings against hours worked for 20 part-time employees shows a positive linear pattern for those working 10 to 30 hours. One employee worked 60 hours but reported unusually low earnings, landing well below the trend. Is this point an outlier, high-leverage, influential, or some combination? Describe how removing it would change the slope, \(r\), and \(s\).
Show answer
The point is high-leverage, its \(x\)-value (60 hours) is far from \(\bar x\), and because it also lies off the trend, it is influential. Removing it would make the slope steeper (the low point at far right drags the right end of the line down), move \(r\) closer to 1 (the remaining points hug the line more tightly), and decrease \(s\) (the residuals shrink). Note that the point's own residual may not look huge, precisely because it pulls the line toward itself; influence is judged by what changes when the point is removed.
A town's population (thousands) was recorded each year since 2000. A linear model of population on years-since-2000 has a U-shaped residual plot. Fitting \(\log(\text{population})\) instead gives \(\widehat{\log(\text{pop})} = 1.70 + 0.045(\text{years})\) with \(r^2 = 0.99\) and a patternless residual plot. Predict the population in 2015, explain why the log model was chosen, and interpret the slope 0.045.
Show answer
For 2015, years \(= 15\): \(\log\hat y = 1.70 + 0.045(15) = 2.375\), so \(\hat y = 10^{2.375} \approx 237\) thousand people. The log model was chosen because the linear model's residual plot was curved (a systematic misfit) while the transformed data are linear with random residuals and a very high \(r^2\). A straight line in \(\log y\) versus \(x\) means exponential growth: each year multiplies the predicted population by \(10^{0.045} \approx 1.109\), about 10.9% growth per year.
Across a large number of fires, the correlation between the number of firefighters sent and the dollar amount of damage is \(r = 0.92\). Which statement is correct?
This confuses correlation with causation. A strong association in observational data never establishes that one variable causes the other, and here the causal story is backwards anyway: bigger, more damaging fires attract more firefighters.
The fraction of variation explained is \(r^2\), not \(r\): \(0.92^2 = 0.846\), so about 85% of the variation in damage is explained, not 92%. Reading \(r\) itself as a percent of variation is one of the most common Unit 2 errors.
Bigger fires call for more firefighters and cause more damage, so fire size is a lurking variable that explains the association without either measured variable causing the other.
Correlation has no units and is unchanged by any linear change of units: dividing every damage figure by 1,000 rescales the variable but leaves \(r\) exactly 0.92. Only the slope of a regression line would change.
Lesson 3.1 · Unit 3 · CED topics 3.1–3.2
Populations, samples, and types of studies
Units 1 and 2 described data you already had. Unit 3 asks where the data came from,
no amount of clever analysis rescues a badly collected sample.
Definitions
The population is the entire group you want to know about; a
sample is the part of it you actually measure. A parameter is
a number describing the population (\(\mu\), \(\sigma\), \(p\)); a statistic
is a number computed from the sample (\(\bar{x}\), \(s\), \(\hat{p}\)). Statistics
estimate parameters. A census measures every individual in the
population; a sample survey measures a sample. An observational
study records variables without intervening; an experiment
deliberately imposes treatments and observes the response.
Worked example · Naming the pieces
A school district wants to know the mean number of hours its 4,200 high school students
sleep on school nights. It surveys 150 randomly chosen students and finds a mean of
7.1 hours.
Population: all 4,200 high school students in the district. Sample: the 150 students
surveyed. Parameter: \(\mu\), the true mean sleep of all 4,200 students (unknown).
Statistic: \(\bar{x} = 7.1\) hours. Write the parameter with a symbol and in
words, "the true mean hours of sleep of all district students," not just "\(\mu\)."
7.1 is what we observed; \(\mu\) is what we want.
Worked example · Classifying studies
The government attempts to count every resident.: Census.
Researchers track 5,000 adults for 20 years, recording diet and heart disease.:
Observational study: they measured what people chose to eat.
Researchers randomly assign 200 adults to a Mediterranean or standard diet and
measure blood pressure.: Experiment: the treatment was imposed.
A pollster calls 1,000 randomly chosen voters.: Sample survey.
The test question that separates the two big categories: did the researchers
decide who got which condition? If yes, it's an experiment.
Why not always take a census? Cost, time, and sometimes impossibility (testing every
light bulb to failure leaves none to sell). A well-chosen sample of a few thousand
estimates a national parameter very accurately.
Try it
A quality inspector selects 40 phone batteries from a day's production of 12,000 and
finds that 3 fail a charge test. Identify the population, sample, parameter, and
statistic, and say whether this is a census.
Show answer
Population: all 12,000 batteries produced that day. Sample: the 40 tested. Parameter:
\(p\), the true proportion of the day's batteries that would fail. Statistic:
\(\hat{p} = 3/40 = 0.075\). Not a census: only a sample was tested (and a census would
be impractical if the test damages the battery).
Lesson 3.2 · Unit 3 · CED topic 3.3
Random sampling methods
Letting people choose whether to be in a sample, or letting the researcher pick
convenient ones, produces samples that systematically differ from the population. The
cure is chance: when a random mechanism picks the sample, nothing can tilt the selection.
Definitions
A simple random sample (SRS) of size \(n\) is chosen so that every set of
\(n\) individuals has the same chance of being selected. A stratified random
sample splits the population into homogeneous groups (strata) and takes an SRS
from every stratum. A cluster sample splits the population into
groups (clusters), randomly selects some clusters, and includes everyone in
them. A systematic sample picks a random start and then every \(k\)th
individual. Convenience samples (whoever is nearby) and voluntary
response samples (whoever chooses to answer) are not random and are biased.
Worked example · Selecting an SRS with a random digit table
Choose 5 of 40 students. Label the students 01 to 40, then read two-digit groups from a
line of random digits, skipping numbers over 40 and any repeats, until you have five.
Line 112
19223
95034
05756
28713
96409
12531
Pairs: 19 ✓, 22 ✓, 39 ✓, 50 (skip), 34 ✓, 05 ✓. The sample is students 05, 19, 22, 34,
and 39. On the exam, write out every step: label, read, skip rule, stop rule.
Calculator
MATH → PROB → randInt(1, 40, 5) generates five integers from 1 to 40
(ignore repeats and generate more); randIntNoRep(1, 40, 5) gives five
distinct ones directly. State the labels and the command in your answer.
Worked example · Which method and why
A school of 1,600 wants student opinion on a new schedule. Freshmen and seniors likely
feel differently, so a stratified sample, an SRS of 50 from each grade, guarantees every grade is represented and is more precise than an SRS of 200. Always
justify the strata: "opinions likely differ across grades." If the
school instead randomly picks 8 of its 64 homerooms and surveys everyone in them, that's
a cluster sample: cheaper and faster, but less precise. A stratum should
be internally alike; a cluster should look like a mini-population.
Try it
A college has 300 first-year students and 200 seniors. Describe how to select a
stratified random sample of 60 that keeps the class proportions, and explain why
stratifying beats an SRS here if the question is about campus housing satisfaction.
Show answer
Select an SRS of \(60 \times 300/500 = 36\) first-years (label 001–300, use
randIntNoRep(1, 300, 36)) and, separately, an SRS of 24 seniors. Housing satisfaction
probably differs by class year, so sampling within each stratum guarantees both are
represented in proportion and reduces the variability of the estimate.
Lesson 3.3 · Unit 3 · CED topics 3.4–3.5
Bias, confounding, and why observational studies can't prove cause
Even a random sample goes wrong if some people can't be reached, refuse, or don't answer
truthfully. And even a perfect survey can't answer "does X cause Y?" if the people who
did X differ from the rest in other ways.
Sources of bias
Bias is a systematic tendency to overestimate or underestimate the truth.
Undercoverage: some groups have no chance to be selected (a landline
survey misses cell-only households). Nonresponse: selected individuals
can't be reached or refuse, and they differ from responders. Response
bias: answers are inaccurate; people lie about sensitive behavior or
misremember. Wording: leading or confusing questions push answers one
way. Bias is about the method; a bigger sample does not fix it.
Worked example · Naming the bias and its direction
A city mails a survey about a proposed tax increase to 5,000 households; 900 return it,
and 71% oppose the tax. Name the likely bias and its direction.
Nonresponse bias. Only 18% responded, and people angry about a tax are more
motivated to mail back a survey than the indifferent or mildly in favor. So 71% likely
overestimates the true proportion who oppose. The graded answer has three parts:
the name, why responders differ from nonresponders, and the direction of the error.
Definition · Confounding
Two variables are confounded when their effects on the response can't be
separated: a third variable is associated with the explanatory variable and also
affects the response. In an observational study the groups chose themselves, so they
can differ in countless other ways. That's why observational studies cannot
establish cause and effect; only random assignment (next lesson) breaks the
link between the treatment and everything else.
Worked example · Spotting the confounder
Adults who take a daily multivitamin have lower rates of heart disease in an
observational study. Can we conclude vitamins protect the heart?
No. People who take vitamins are also more likely to exercise, eat well, and see a
doctor. Exercise, say, is associated with taking vitamins and lowers heart
disease risk, so the two effects are confounded. To earn the point, name a specific
variable and link it to both the explanatory variable and the response.
Try it
A survey asks, "Given that reckless drivers cause thousands of deaths each year, do you
support tougher penalties for speeding?", and 89% say yes. Name the problem and its
likely direction.
Show answer
Wording (question) bias: the preamble about deaths pushes respondents toward "yes,"
so 89% likely overestimates the proportion who would support tougher penalties if
asked neutrally ("Do you support or oppose tougher penalties for speeding?").
Lesson 3.4 · Unit 3 · CED topics 3.5–3.6
Designing experiments
An experiment is the only design that can support a cause-and-effect claim, and it earns
that power from one move: letting chance decide who gets which treatment. Everything else
is about making the comparison sharper.
Definitions and principles
Experimental units (called subjects if human) receive
treatments, which are the specific conditions imposed; the
response variable is what's measured afterward. Three principles:
comparison/control: include at least two treatments (often one is a
placebo, a fake treatment, or the current standard) so outside variables
affect all groups equally; random assignment: chance puts units into
groups, balancing unknown variables and defeating confounding; replication: enough units per group that chance variation doesn't hide a real effect.
Blinding: subjects don't know their treatment (single-blind), and neither
do those measuring the response (double-blind).
Worked example · Completely randomized design
A grower has 40 tomato seedlings and wants to test whether a new fertilizer increases
yield compared with the current one. Describe a completely randomized design.
Label the seedlings 01–40. Use randIntNoRep(1, 40); the first 20 numbers drawn get the
new fertilizer, the other 20 the current one. Grow all plants under identical conditions
(water, light, spacing). At season's end, record yield in kilograms per plant and
compare the two groups' mean yields. Say what "compare" means: if the new-fertilizer
mean is higher by more than random-assignment variation would explain, conclude the
fertilizer increased yield.
Worked example · Blocking and matched pairs
Suppose 20 seedlings sit in a sunny plot and 20 in a shady one. Sunlight affects yield
regardless of fertilizer, so use a randomized block design: sun and shade
are blocks, and within each block randomly assign 10 plants to each fertilizer. Blocking
removes sunlight variation from the comparison, making a fertilizer effect easier to
detect. Matched pairs is blocking at its finest: each pair of similar units
(or each subject, measured twice) gets both treatments, with the assignment within the
pair randomized.
Exam tip: you stratify when sampling; you block when assigning treatments.
Blocks are formed on a variable known to affect the response, never randomly.
Try it
Researchers want to know whether caffeine improves reaction time. They have 30
volunteers. Describe a matched pairs design, and state what makes it double-blind.
Show answer
Each volunteer is their own pair: on one day they take a caffeine pill, on another an
identical-looking placebo, order decided by a coin flip for each subject. Measure
reaction time after each and analyze the 30 differences. It is double-blind if neither
the volunteer nor the person recording reaction times knows which pill was taken when.
Lesson 3.5 · Unit 3 · CED topic 3.7
Scope of inference: what each design lets you conclude
Every study allows some conclusions and forbids others. The rules come down to two
questions: were the individuals chosen at random, and were the treatments assigned at
random? Each "yes" unlocks one kind of claim.
Rule
Random selection of individuals from a population permits
generalizing the results to that population. Random assignment
of treatments permits concluding that the treatment caused the difference
in response. The two are independent, giving four cases:
Random assignment
No random assignment
Random selection
Cause and effect, generalizable to the population
Association only, generalizable to the population
No random selection
Cause and effect, but only for individuals like those in the study
Association only, only for those studied
Worked example · Classifying four studies
An SRS of 500 adults nationwide is surveyed; those who sleep less than 6 hours
report more headaches.: Random selection, no random assignment: we can generalize
the association to all U.S. adults, but can't say short sleep causes headaches
(stress could confound).
Sixty volunteers are randomly assigned to 6 or 8 hours of sleep for a week; the
8-hour group has fewer headaches.: Random assignment, no random selection: sleep
duration caused the difference, but only for people similar to these
volunteers.
An SRS of 200 employees at a company is randomly assigned to standing or sitting
desks; standers report less back pain.: Both: the desks caused the difference, and
the result generalizes to all employees at that company.
A teacher compares test scores of students who chose to attend review sessions
with those who didn't.: Neither: an association among these students only.
Worked example · Writing the conclusion
For the standing-desk study, a full-credit scope statement reads: "Because desks were
randomly assigned, we can conclude that standing desks caused the reduction in reported
back pain. Because the employees were randomly selected from this company, the result
can be generalized to all employees of the company." Two sentences, each naming the
random step and the conclusion it justifies.
Exam tip: medical experiments use volunteers, so "cause and effect, but not generalizable
beyond similar volunteers" is the most common correct answer. A large sample doesn't
justify generalizing: size isn't randomness.
Try it
A university randomly selects 400 of its students and asks about hours of part-time
work and GPA; students working more than 20 hours have lower GPAs on average. What can
be concluded, and what can't?
Show answer
Random selection means the association between heavy part-time work and lower GPA can
be generalized to all students at this university. But work hours were not randomly
assigned, students chose them, so we cannot conclude that working more causes lower
GPAs; a confounding variable such as financial stress could be responsible.
Unit 3 practice · 10 problems
Unit 3 practice: Collecting Data
Ten problems covering the whole unit, in roughly exam order. Most answers here are
sentences, not numbers: write them out in full before revealing the answer, and
check that every design you describe includes the random step. No calculator needed.
A city's transportation department wants to know what proportion of the 84,000 vehicles registered in the city would fail an emissions test. Inspectors randomly select 500 registered vehicles and test them; 35 fail. Identify the population, the sample, the parameter, and the statistic.
Show answer
Population: all 84,000 vehicles registered in the city. Sample: the 500 vehicles that were tested. Parameter: \(p\), the true proportion of all registered vehicles in the city that would fail the emissions test (unknown). Statistic: \(\hat p = 35/500 = 0.07\), the proportion of the sampled vehicles that failed. Always write the parameter in words as well as in symbols.
Which of the following is an experiment?
This is an observational study: the people chose whether to smoke, so the researchers did not impose the explanatory variable. Comparing two existing groups looks like treatment versus control, but smoking is confounded with other lifestyle differences.
A survey is a sampling study: it measures the opinions of a random sample but imposes no treatment on anyone. Random selection of subjects is not the same thing as random assignment of treatments.
In an experiment the researchers impose the treatment: they decided who took vitamin D and who took the placebo, using random assignment. The one question to ask is whether the researchers assigned the conditions: only here is the answer yes.
This is a census, an attempt to measure every individual in the population. It collects data but imposes no treatment and compares no groups, so it is neither an experiment nor even a sample.
Name the sampling method in each case. (a) A principal numbers all 1,200 students and uses randIntNoRep(1, 1200, 100) to choose 100. (b) She selects an SRS of 25 students from each of the four grades. (c) She randomly selects 5 of the 40 homerooms and surveys every student in them. (d) She posts a survey link on the school website for anyone to answer. (e) She stands at the front door and surveys every 12th student who enters, starting with a randomly chosen student among the first 12.
Show answer
(a) Simple random sample: every group of 100 students is equally likely. (b) Stratified random sample: the grades are strata, and an SRS is taken from every stratum. (c) Cluster sample: homerooms are clusters, some are chosen at random, and everyone in the chosen clusters is included. (d) Voluntary response sample, not random; people with strong opinions are overrepresented. (e) Systematic random sample: a random start, then every 12th individual.
A manager wants to choose 4 of her 30 employees for a focus group. Describe how to select an SRS using this line from a table of random digits, and give the sample it produces: 41072 50733 18960 27411 96085
Show answer
Label the employees 01 through 30. Read two-digit groups from left to right, skipping any number above 30 (and 00) and any repeat, until four distinct labels are chosen. The pairs are 41 (skip), 07 ✓, 25 ✓, 07 (repeat, skip), 33 (skip), 18 ✓, 96 (skip), 02 ✓. The sample is employees 02, 07, 18, and 25. On the exam, all four pieces are graded: the labels, how you read the digits, the skip rule, and the stopping rule.
A university with 6,000 undergraduates and 2,000 graduate students wants to survey 200 students about extending library hours. (a) Explain why a stratified random sample by student level is likely to give a more precise estimate than an SRS of 200, and describe how to select it. (b) Describe a cluster sample the university could use instead, and give one advantage and one disadvantage.
Show answer
(a) Undergraduates and graduate students probably use the library differently (graduate students more often at night), so their opinions are likely to differ. Stratifying guarantees both groups are represented in the right proportion, an SRS of \(200 \times 6000/8000 = 150\) undergraduates and a separate SRS of 50 graduate students, and removes the between-group variation from the estimate, reducing its variability. An SRS might, by chance, badly over- or underrepresent graduate students. (b) Randomly select, say, 10 of the university's course sections and survey every student in those sections. Advantage: cheaper and faster; you visit 10 classrooms instead of tracking down 200 individuals. Disadvantage: students in the same section tend to be similar, so the estimate is less precise than an SRS of the same size.
(a) A TV station invites viewers to text "YES" or "NO" on a proposed stadium tax; of 3,000 texts, 78% say NO. Name the type of bias and its likely direction. (b) A polling firm calls landline numbers between 10 a.m. and 2 p.m. on weekdays to estimate the proportion of adults who work full time. Name two sources of bias and the direction of each.
Show answer
(a) Voluntary response bias. People who feel strongly, especially those angry about a new tax, are far more likely to bother texting, so 78% probably overestimates the proportion of all viewers (let alone all residents) who oppose the tax. (b) Undercoverage: adults without landlines (younger and often working adults) have no chance of being selected. Nonresponse: full-time workers are unlikely to be home to answer between 10 and 2, so they are selected but not reached. Both push the estimate the same way: underestimating the proportion who work full time. Neither is fixed by calling more numbers.
An observational study of 2,000 elementary students finds that those who eat breakfast every day score higher on standardized tests than those who don't. A headline reads, "Breakfast boosts test scores." Name a specific confounding variable, explain how it is confounded with eating breakfast, and explain why the headline is not justified.
Show answer
Household income is a plausible confounder. Higher-income families are more likely to provide breakfast every day (associated with the explanatory variable) and more likely to provide books, tutoring, and quiet study space that raise test scores (affects the response). Because the breakfast eaters and non-eaters differ in income, the effect of breakfast cannot be separated from the effect of income. The students chose (or their families chose) whether they ate breakfast, nothing was randomly assigned, so this observational study can show an association but cannot establish that breakfast causes higher scores. Full credit requires naming the variable and linking it to both the explanatory variable and the response.
Sixty volunteers agree to take part in a study of whether a 10-minute guided meditation before an exam reduces test anxiety, measured by a standard 0–40 anxiety questionnaire completed immediately after the exam. Describe a completely randomized design, including how you would carry out the random assignment and what comparison you would make.
Show answer
Number the volunteers 01 to 60. Use randIntNoRep(1, 60, 30): the 30 volunteers whose numbers appear do the 10-minute guided meditation immediately before the exam; the other 30 sit quietly for 10 minutes instead (the control, so both groups get the same 10-minute pause and only the meditation differs). All 60 take the same exam under identical conditions and then complete the anxiety questionnaire. Compare the mean anxiety scores of the two groups; if the meditation group's mean is lower by more than random assignment alone would plausibly produce, conclude that the meditation reduced anxiety. Random assignment balances hidden variables (baseline anxiety, preparation) across the groups; replication (30 per group) keeps chance variation from hiding a real effect.
Suppose 24 of the 60 volunteers in the previous problem have a diagnosed anxiety disorder and 36 do not. Explain why a randomized block design would be better than the completely randomized design, describe how to carry it out, and state how blocking differs from stratifying.
Show answer
A diagnosed anxiety disorder strongly affects the response (anxiety score) regardless of treatment. If the completely randomized design happened to put more of those 24 volunteers in one group, the treatment effect would be blurred. Block on diagnosis: within the 24 diagnosed volunteers, randomly assign 12 to meditation and 12 to control (label 01–24, randIntNoRep(1, 24, 12)); within the 36 without a diagnosis, randomly assign 18 and 18 the same way. Compare meditation with control within each block, then combine. Blocking removes the diagnosis-related variation from the comparison, making a real meditation effect easier to detect. Stratifying is the analogous idea when sampling from a population; blocking is done when assigning treatments in an experiment. Blocks are always formed on a variable known to affect the response, never at random.
(a) A random sample of 300 adults in one state finds that those who drink two or more cups of coffee daily report lower rates of depression. (b) In a separate study, 80 volunteers are randomly assigned to drink two cups of regular or decaffeinated coffee daily for eight weeks; the regular-coffee group reports lower depression scores. For each study, state whether the results can be generalized to a larger population and whether a cause-and-effect conclusion is justified. Justify each answer.
Show answer
(a) Random selection from the state's adults permits generalizing the association between heavy coffee drinking and lower depression rates to all adults in the state. But coffee consumption was not randomly assigned, people chose it, so a confounding variable (sleep habits, social activity, income) could explain the link, and we cannot conclude that coffee reduces depression. (b) Random assignment permits a cause-and-effect conclusion: for people like these volunteers, drinking regular coffee reduced depression scores relative to decaf. But the volunteers were not randomly selected from any population, so the result generalizes only to people similar to them. Two random steps, two separate conclusions.
Lesson 4.1 · Unit 4 · CED topics 4.1–4.2
Randomness, simulation, and the law of large numbers
A random process is unpredictable in the short run but has a stable pattern in the long
run, and that long-run pattern is what probability measures. When the math is hard, you
can find the pattern by imitating the process many times: a simulation.
Definition
The probability of an outcome is the proportion of times it would occur
in a very long series of repetitions of the chance process. The law of large
numbers says that as the number of repetitions grows, the observed proportion
gets closer to the true probability. It says nothing about the next few trials.
Method · Designing a simulation
Describe how one trial is modeled with a chance device (digits, randInt),
what each outcome represents, and what to record. Run many trials.
Count how often the event happened; that proportion is the estimate.
Worked example · Cereal box prizes
One in five cereal boxes contains a prize. Estimate the probability that you need
more than 5 boxes to get your first prize.
Describe: one random digit represents one box; 0–1 means "prize" (probability
0.2) and 2–9 means "no prize." Read digits until a prize appears and record the number
of boxes opened. Run: 200 trials. Count: the proportion of trials with 6
or more boxes estimates the probability. In one run, 67 of 200 trials needed more than
5 boxes: estimate \(67/200 = 0.335\). (The exact answer, \(0.8^5 \approx 0.328\), comes
in Lesson 4.7.)
Calculator
randInt(0, 9, 10) gives ten random digits at once. For the box example,
randInt(1, 5) with 1 = prize is a cleaner model. Whatever you use, say what each number
stands for.
Worked example · The "law of averages" myth
A fair coin has landed heads six times in a row. Is tails now more likely? No: flips
are independent, so the chance of tails is still 0.5. The law of large numbers says the
proportion of heads drifts toward 0.5 over thousands of flips; it does not say
the coin compensates for past results. Believing tails is "due" is the gambler's fallacy.
Try it
A basketball player makes 70% of her free throws. Describe how to simulate one set of
10 shots to estimate the probability that she makes at least 8.
Show answer
Use randInt(1, 10, 10); each 1–7 is a made shot (probability 0.7), each 8–10 a miss.
Count the makes and record whether the count is 8 or more. Repeat many times (say
100 sets); the proportion of sets with 8 or more makes estimates the probability.
(Assumes shots are independent.)
Lesson 4.2 · Unit 4 · CED topics 4.3–4.4
Probability rules: sample spaces, complements, and the addition rule
Once you can list what might happen, probability becomes bookkeeping. A few rules
connect the probabilities of related events, and a Venn diagram keeps the accounts
straight.
Rules
The sample space \(S\) is the set of all possible outcomes. For any event
\(A\), \(0 \le P(A) \le 1\) and \(P(S) = 1\). Complement rule:
\(P(A^c) = 1 - P(A)\). General addition rule:
\[P(A \cup B) = P(A) + P(B) - P(A \cap B).\]
Events are mutually exclusive (disjoint) if they can't happen together,
\(P(A \cap B) = 0\), in which case \(P(A \cup B) = P(A) + P(B)\). Subtracting the
overlap once is the whole point: outcomes in both events would otherwise be counted
twice.
Worked example · Venn diagram
At one school, 45% of seniors take AP Statistics, 30% take AP Calculus, and 15% take
both. Find the probability that a randomly chosen senior takes at least one of the
two, and the probability they take neither.
Let \(A\) = takes Statistics, \(B\) = takes Calculus.
\(P(A \cup B) = 0.45 + 0.30 - 0.15 = 0.60\). By the complement rule,
\(P(\text{neither}) = 1 - 0.60 = 0.40\). In the Venn diagram, the overlap holds 0.15,
the Statistics-only region holds \(0.45 - 0.15 = 0.30\), the Calculus-only region
holds \(0.30 - 0.15 = 0.15\), and the outside holds 0.40. Check: the four regions sum
to 1.
Worked example · Two dice
Roll two fair dice; the sample space has 36 equally likely ordered pairs.
\(P(\text{sum} = 7) = 6/36\) (six pairs: 1-6, 2-5, 3-4, 4-3, 5-2, 6-1). "Sum is 7" and
"doubles" are mutually exclusive (a doubled number gives an even sum), so
\(P(\text{sum} = 7 \text{ or doubles}) = 6/36 + 6/36 = 1/3\). For "at least one 6," go
through the complement: 25 of the 36 pairs have no 6, so
\(P(\text{at least one 6}) = 1 - 25/36 = 11/36 \approx 0.306\).
Exam tip: the phrases "at least one" and "or" are signals: complement rule for the
first, addition rule for the second. Define your events in symbols before computing,
and show the rule you're applying.
Try it
Among adults in a town, 62% drink coffee daily, 24% drink tea daily, and 10% drink
both. What is the probability a randomly chosen adult drinks coffee or tea daily? Neither?
Are "drinks coffee" and "drinks tea" mutually exclusive?
Conditional probability, tree diagrams, and independence
Knowing that one event happened often changes the odds of another. Conditional
probability makes that precise, gives you a rule for "and," and provides the test for
whether two events are truly independent.
Rules
The conditional probability of \(A\) given \(B\) is
\[P(A \mid B) = \frac{P(A \cap B)}{P(B)}.\]
Rearranged, this is the general multiplication rule:
\(P(A \cap B) = P(A)\,P(B \mid A)\). Events \(A\) and \(B\) are independent
if \(P(A \mid B) = P(A)\), knowing \(B\) doesn't change the chance of \(A\), and then
\(P(A \cap B) = P(A)\,P(B)\). Independent and mutually exclusive are different: disjoint
events with positive probability are never independent.
Worked example · From a two-way table
Grade
Has a job
No job
Total
10th
20
60
80
12th
66
54
120
Total
86
114
200
Choose one student at random. \(P(\text{job} \mid \text{12th}) = 66/120 = 0.55\): the
condition shrinks the sample space to the 120 seniors. \(P(\text{12th} \mid \text{job}) = 66/86 \approx 0.767\).
Are "has a job" and "12th grade" independent? \(P(\text{job}) = 86/200 = 0.43\), which is
not equal to \(P(\text{job} \mid \text{12th}) = 0.55\), so no: knowing a
student is a senior raises the chance they have a job.
Worked example · Tree diagram
About 2% of people have a certain condition. A screening test detects it in 95% of
people who have it, but also gives a positive result to 4% of people who don't. If a
randomly chosen person tests positive, what is the probability they have the condition?
The tree's first branches are \(P(D) = 0.02\) and \(P(D^c) = 0.98\); the second set of
branches are the test results given each. Multiply along branches:
Branch
Calculation
Probability
D and positive
0.02 × 0.95
0.0190
D and negative
0.02 × 0.05
0.0010
No D and positive
0.98 × 0.04
0.0392
No D and negative
0.98 × 0.96
0.9408
\(P(\text{positive}) = 0.0190 + 0.0392 = 0.0582\), so
\[P(D \mid \text{positive}) = \frac{0.0190}{0.0582} \approx 0.326.\]
Only about a third of positives actually have the condition: because the condition
is rare, the 4% false positives from the large healthy group outnumber the true
positives.
Try it
Two cards are dealt without replacement from a standard deck. Find the probability
that both are hearts, and explain which rule you used.
Show answer
General multiplication rule: \(P(H_1 \cap H_2) = P(H_1)\,P(H_2 \mid H_1) = \dfrac{13}{52}\cdot\dfrac{12}{51} \approx 0.0588\).
The draws are not independent, after one heart is removed, only 12 of 51 remain.
Lesson 4.4 · Unit 4 · CED topics 4.7–4.8
Discrete random variables: expected value and standard deviation
A random variable attaches a number to each outcome of a chance process: the number
of toppings on the next pizza ordered, your winnings on a raffle ticket. Its probability
distribution is a table, and from the table you compute a mean and a standard deviation
just as you did for data, with probabilities playing the role of relative frequencies.
Formulas
A discrete random variable \(X\) takes a countable set of values \(x_i\)
with probabilities \(p_i\) that sum to 1. Its mean (expected value) and
variance are
\[\mu_X = E(X) = \sum x_i\,p_i, \qquad \sigma_X^2 = \sum (x_i - \mu_X)^2\,p_i,\]
and \(\sigma_X = \sqrt{\sigma_X^2}\). Interpret \(\mu_X\) as the long-run average value
over many repetitions, and \(\sigma_X\) as the typical distance of a value from that
average.
Worked example · Toppings on a pizza
Let \(X\) = number of toppings on a randomly selected pizza order.
x
0
1
2
3
4
P(X = x)
0.15
0.35
0.30
0.15
0.05
x · P(x)
0
0.35
0.60
0.45
0.20
(x − 1.6)² · P(x)
0.384
0.126
0.048
0.294
0.288
\(\mu_X = 0 + 0.35 + 0.60 + 0.45 + 0.20 = 1.6\) toppings. Interpretation: over many
orders, the average number of toppings would be about 1.6. The variance is the sum of
the last row, \(\sigma_X^2 = 1.14\), so \(\sigma_X = \sqrt{1.14} \approx 1.07\) toppings:
the number of toppings typically differs from the mean by about 1.07. Also,
\(P(X \ge 2) = 0.30 + 0.15 + 0.05 = 0.50\).
Worked example · Expected value of a raffle
A club sells 500 raffle tickets at $5 each. One ticket wins $1,000, two win $100, and
five win $20. Let \(W\) = your net winnings from one ticket.
The values are \(995, 95, 15, -5\) with probabilities \(1/500, 2/500, 5/500, 492/500\):
\[E(W) = \frac{995(1) + 95(2) + 15(5) + (-5)(492)}{500} = \frac{-1200}{500} = -\$2.40.\]
If you bought many tickets, you would lose an average of $2.40 per ticket. Expected
value is a long-run average, not a value you can get on a single ticket.
Calculator
Enter the values in L1 and the probabilities in L2, then 1-Var Stats L1, L2.
The screen's \(\bar{x}\) is \(\mu_X\) and \(\sigma x\) is \(\sigma_X\) (use \(\sigma x\),
not \(Sx\), for a probability distribution). Still write the formula with at least the
first terms shown.
Try it
\(X\) = number of pets in a randomly chosen household, with \(P(0) = 0.40\),
\(P(1) = 0.35\), \(P(2) = 0.15\), \(P(3) = 0.10\). Find and interpret \(\mu_X\); find \(\sigma_X\).
Show answer
\(\mu_X = 0(0.40) + 1(0.35) + 2(0.15) + 3(0.10) = 0.95\) pets: the average over many
households would be about 0.95 pets.
\(\sigma_X^2 = (0.95)^2(0.40) + (0.05)^2(0.35) + (1.05)^2(0.15) + (2.05)^2(0.10) = 0.9475\),
so \(\sigma_X \approx 0.97\) pets.
Lesson 4.5 · Unit 4 · CED topic 4.9
Transforming and combining random variables
Real quantities are built from other random quantities: a bill is a fixed fee plus a
per-item charge, a delivery time is prep time plus driving time. You can find the mean
and standard deviation of the combination without listing every outcome, as long as
you remember the one rule everyone gets wrong.
Rules
For constants \(a\) and \(b\): \(\mu_{aX+b} = a\mu_X + b\) and \(\sigma_{aX+b} = |a|\,\sigma_X\)
(adding \(b\) shifts the center but not the spread). For any two random variables,
\(\mu_{X \pm Y} = \mu_X \pm \mu_Y\). If \(X\) and \(Y\) are independent,
\[\sigma^2_{X \pm Y} = \sigma_X^2 + \sigma_Y^2.\]
Variances add: for a sum and for a difference. Standard deviations never
add. Subtracting a variable still adds uncertainty.
Worked example · A linear transformation
From Lesson 4.4, the number of toppings \(X\) has \(\mu_X = 1.6\) and \(\sigma_X = 1.068\).
A pizza costs $12 plus $1.50 per topping, so \(C = 12 + 1.5X\).
\(\mu_C = 12 + 1.5(1.6) = \$14.40\) and \(\sigma_C = 1.5(1.068) \approx \$1.60\). The $12
base fee moved the mean but did nothing to the spread.
Worked example · Sum of independent variables
Prep time \(P\) at a pizza shop has mean 12 minutes and SD 2 minutes; driving time \(D\)
has mean 15 minutes and SD 4 minutes, independent of prep. Total delivery time
\(T = P + D\):
\[\mu_T = 12 + 15 = 27 \text{ min}, \qquad \sigma_T = \sqrt{2^2 + 4^2} = \sqrt{20} \approx 4.47 \text{ min}.\]
The tempting answer \(2 + 4 = 6\) minutes is wrong. Note that you must state (or be
told) that \(P\) and \(D\) are independent before using the variance rule.
Worked example · A difference
Machine A fills cans with mean 355 mL and SD 3 mL; independent Machine B has mean
350 mL and SD 4 mL. For a random can from each, \(A - B\) has mean \(355 - 350 = 5\) mL
and SD \(\sqrt{3^2 + 4^2} = 5\) mL: the variances still add even though we subtracted.
Exam tip: write the variance step explicitly; "\(\sigma_T^2 = 4 + 16 = 20\), so
\(\sigma_T = 4.47\)". Skipping to the square root hides whether you added variances or SDs.
Try it
\(X\) has mean 50 and SD 6; \(Y\) has mean 30 and SD 8; they are independent. Find the
mean and SD of (a) \(X + Y\), (b) \(X - Y\), (c) \(2X + 10\).
Show answer
(a) Mean 80, SD \(\sqrt{36 + 64} = 10\). (b) Mean 20, SD also 10. (c) Mean
\(2(50) + 10 = 110\), SD \(2(6) = 12\).
Lesson 4.6 · Unit 4 · CED topics 4.10–4.11
The binomial distribution
Count the successes in a fixed number of identical yes/no trials (free throws made out
of 10, defective parts in a batch of 20, correct guesses on 12 questions), and you have
a binomial random variable. Its distribution has a formula, a mean, and a standard
deviation you can write down without a table.
Conditions (BINS)
Binary: each trial is a success or a failure.
Independent: the outcome of one trial doesn't affect another.
Number: the number of trials \(n\) is fixed in advance.
Same probability: \(P(\text{success}) = p\) on every trial.
Then \(X\) = number of successes is binomial, \(X \sim B(n, p)\), with
\[P(X = k) = \binom{n}{k} p^k (1-p)^{n-k}, \qquad \mu_X = np, \qquad \sigma_X = \sqrt{np(1-p)}.\]
10% condition: when sampling without replacement from a population of
size \(N\), the trials are approximately independent as long as \(n \le 0.10N\).
Worked example · Guessing on a quiz
A 12-question multiple-choice quiz has four options per question. A student guesses on
every question. Let \(X\) = number correct.
BINS: each guess is right or wrong; guesses are independent; \(n = 12\) is fixed;
\(p = 0.25\) each time. So \(X \sim B(12, 0.25)\).
\[P(X = 3) = \binom{12}{3}(0.25)^3(0.75)^9 = 220(0.015625)(0.07508) \approx 0.258.\]
\(P(X \le 3) = P(0) + P(1) + P(2) + P(3) \approx 0.649\). For "at least 6 correct," use the
complement: \(P(X \ge 6) = 1 - P(X \le 5) \approx 1 - 0.946 = 0.054\). Mean and SD:
\(\mu_X = 12(0.25) = 3\) correct, \(\sigma_X = \sqrt{12(0.25)(0.75)} = 1.5\) correct.
Interpretation: a guessing student gets 3 right on average, typically within about
1.5 of that.
Calculator
2nd → DISTR: binompdf(n, p, k) gives \(P(X = k)\) and
binomcdf(n, p, k) gives \(P(X \le k)\). Here binompdf(12, 0.25, 3) = 0.258,
binomcdf(12, 0.25, 3) = 0.649, and \(1 - \) binomcdf(12, 0.25, 5) = 0.054. On the exam,
name the distribution and its parameters, "\(X\) is binomial with \(n = 12\), \(p = 0.25\)", and state the probability you're computing before quoting the command.
Common error: "at least 6" is \(X \ge 6\), so subtract binomcdf up to 5, not 6. Draw a
quick number line if the boundary is confusing.
Try it
A player makes 70% of her free throws, independently. In 10 attempts, find \(P(X = 8)\),
\(P(X \ge 8)\), and the mean and SD of \(X\).
The binomial counts successes in a fixed number of trials. Flip the question: how
many trials until the first success?, and the number of trials is no longer
fixed; it's the random variable. That's the geometric setting, and it has the simplest
formulas in the unit.
Conditions and formulas
The trials must be binary, independent, and have the same success probability \(p\):
the same as binomial except that instead of a fixed \(n\), you count trials
until the first success. Then \(X\) = number of the trial on which the first
success occurs, and for \(k = 1, 2, 3, \ldots\)
\[P(X = k) = (1-p)^{k-1}\,p, \qquad \mu_X = \frac{1}{p}, \qquad \sigma_X = \frac{\sqrt{1-p}}{p}.\]
Two shortcuts: \(P(X \gt k) = (1-p)^k\) (the first \(k\) trials all fail), so
\(P(X \le k) = 1 - (1-p)^k\). Geometric distributions are always skewed right.
Worked example · Cereal boxes, exactly
One in five cereal boxes holds a prize. Let \(X\) = number of boxes opened to find
the first prize.
Each box is a prize or not, boxes are independent, \(p = 0.2\) every time, and we count
boxes until the first prize: geometric with \(p = 0.2\).
\[P(X = 3) = (0.8)^2(0.2) = 0.128,\]
the chance the first two boxes miss and the third hits.
\(P(X \le 3) = 1 - 0.8^3 = 0.488\), and \(P(X \gt 5) = 0.8^5 \approx 0.328\): the exact
value the Lesson 4.1 simulation estimated as 0.335. On average you'd open
\(\mu_X = 1/0.2 = 5\) boxes, with \(\sigma_X = \sqrt{0.8}/0.2 \approx 4.47\) boxes; the large
SD reflects the long right tail (sometimes you open 15 boxes).
Calculator
2nd → DISTR: geometpdf(p, k) gives \(P(X = k)\) and
geometcdf(p, k) gives \(P(X \le k)\). Here geometpdf(0.2, 3) = 0.128 and
geometcdf(0.2, 3) = 0.488; \(P(X \gt 5) = 1 - \) geometcdf(0.2, 5) = 0.328. Note the
argument order is \(p\) first, the reverse of the binomial commands' \(n\)-first habit.
Worked example · Binomial or geometric?
"A quality inspector tests parts until finding a defective one" is geometric (trials
until first success: here "success" is a defect). "A quality inspector tests 30 parts
and counts the defective ones" is binomial with \(n = 30\). The giveaway is whether
the number of trials is fixed (binomial) or is itself the thing being counted (geometric).
Try it
You roll a fair die until a 6 appears. Find the probability that the first 6 comes on
the fourth roll, the probability it takes at most four rolls, and the expected number
of rolls.
Unit 4 practice: Probability, Random Variables, and Probability Distributions
Ten problems covering the whole unit, in roughly exam order. Work each one on paper
before revealing the answer. Use binompdf, binomcdf, geometpdf, and geometcdf for
problems 8 and 9, but write the formula with numbers substituted as well.
Thirty percent of customers at a coffee shop pay with the shop's loyalty app. Describe how to use a random number generator to simulate one trial that estimates the probability that at least 4 of the next 10 customers pay with the app, and explain how you would use many trials to estimate the probability.
Show answer
Describe: use randInt(1, 10, 10) to generate ten integers, one per customer. Let 1, 2, or 3 represent a customer who pays with the app (probability 0.3) and 4 through 10 a customer who doesn't. Count how many of the ten integers are 1–3 and record whether the count is at least 4. Run: repeat for many trials, say 200. Count: the proportion of trials in which at least 4 of the 10 used the app is the estimate. (This assumes customers' payment choices are independent.)
An American roulette wheel has 38 equally likely slots, 18 of them red. After the ball lands on black seven spins in a row, a gambler bets heavily on red, saying red is "due." Explain what is wrong with the gambler's reasoning and what the law of large numbers actually says.
Show answer
Each spin is independent, so the probability of red on the next spin is still \(18/38 \approx 0.474\): the wheel has no memory and does not compensate for past results. The law of large numbers says that over a very large number of spins the proportion of reds will get close to 0.474; it makes no promise about the next few spins and does not require a short-run streak to be balanced out. Believing that a streak makes the opposite outcome more likely is the gambler's fallacy.
At a high school, 55% of students have a part-time job, 35% play a school sport, and 20% do both. A student is chosen at random. (a) Find the probability that the student has a job or plays a sport. (b) Find the probability the student does neither. (c) Are "has a job" and "plays a sport" mutually exclusive? Explain.
Show answer
Let \(J\) = has a job, \(S\) = plays a sport. (a) General addition rule: \(P(J \cup S) = 0.55 + 0.35 - 0.20 = 0.70\). (b) Complement: \(P(\text{neither}) = 1 - 0.70 = 0.30\). (c) No: mutually exclusive events cannot both occur, but \(P(J \cap S) = 0.20 \ne 0\). In a Venn diagram: job only 0.35, both 0.20, sport only 0.15, neither 0.30, summing to 1.
A random sample of 150 students was asked whether they had used a phone during class that day.
Grade
Used phone
Did not
Total
Freshmen
36
24
60
Seniors
54
36
90
Total
90
60
150
One student is selected at random. Find \(P(\text{used phone})\), \(P(\text{used phone} \mid \text{freshman})\), and \(P(\text{freshman} \mid \text{used phone})\). Are the events "used phone" and "freshman" independent? Justify.
Show answer
\(P(\text{used phone}) = 90/150 = 0.60\). \(P(\text{used phone} \mid \text{freshman}) = 36/60 = 0.60\) (restrict to the 60 freshmen). \(P(\text{freshman} \mid \text{used phone}) = 36/90 = 0.40\) (restrict to the 90 phone users). Since \(P(\text{used phone} \mid \text{freshman}) = 0.60 = P(\text{used phone})\), knowing a student is a freshman does not change the probability of phone use: the events are independent. (Check: \(P(\text{used phone} \mid \text{senior}) = 54/90 = 0.60\) as well.)
A factory makes 60% of its bolts on Machine A, which produces 2% defective bolts, and 40% on Machine B, which produces 5% defective. A bolt is chosen at random from the day's output. (a) Find the probability it is defective. (b) Given that it is defective, find the probability it came from Machine A. Show a tree diagram or the rules you used.
Show answer
Tree: first branches \(P(A) = 0.6\), \(P(B) = 0.4\); second branches are defective or not given the machine. Multiply along branches: \(P(A \cap D) = 0.6(0.02) = 0.012\) and \(P(B \cap D) = 0.4(0.05) = 0.020\). (a) \(P(D) = 0.012 + 0.020 = 0.032\). (b) Conditional probability: \[P(A \mid D) = \frac{P(A \cap D)}{P(D)} = \frac{0.012}{0.032} = 0.375.\] Although Machine A makes most of the bolts, only 37.5% of the defective ones come from it, because its defect rate is so much lower.
Let \(X\) = the number of cars owned by a randomly selected household in a town.
x
0
1
2
3
P(X = x)
0.10
0.35
0.40
0.15
Find and interpret \(\mu_X\). Find \(\sigma_X\). Find \(P(X \ge 2)\).
Show answer
\(\mu_X = 0(0.10) + 1(0.35) + 2(0.40) + 3(0.15) = 1.6\) cars. Over many randomly selected households, the average number of cars would be about 1.6. Variance: \[\sigma_X^2 = (0 - 1.6)^2(0.10) + (1 - 1.6)^2(0.35) + (2 - 1.6)^2(0.40) + (3 - 1.6)^2(0.15) = 0.256 + 0.126 + 0.064 + 0.294 = 0.74,\] so \(\sigma_X = \sqrt{0.74} \approx 0.86\) cars: a household's car count typically differs from 1.6 by about 0.86. \(P(X \ge 2) = 0.40 + 0.15 = 0.55\).
A bakery's daily number of online orders \(X\) has mean 40 and standard deviation 5; its daily number of phone orders \(Y\) has mean 25 and standard deviation 3. The two are independent. (a) Find the mean and standard deviation of the total number of orders \(T = X + Y\). (b) Find the mean and standard deviation of \(X - Y\). (c) A delivery service charges the bakery $3 per online order plus a flat $2 per day, so the daily cost is \(C = 3X + 2\). Find the mean and standard deviation of \(C\).
Show answer
(a) \(\mu_T = 40 + 25 = 65\) orders; \(\sigma_T^2 = 5^2 + 3^2 = 34\), so \(\sigma_T = \sqrt{34} \approx 5.83\) orders (variances add because \(X\) and \(Y\) are independent). (b) \(\mu_{X-Y} = 40 - 25 = 15\); the variance of a difference is still the sum \(25 + 9 = 34\), so \(\sigma_{X-Y} \approx 5.83\) as well. (c) \(\mu_C = 3(40) + 2 = \$122\); \(\sigma_C = 3(5) = \$15\): the flat $2 shifts the mean but does not change the spread. Never add standard deviations.
Fifteen percent of flights at an airport are delayed, independently of one another. Let \(X\) = the number of delayed flights among 12 randomly selected flights. (a) Explain why \(X\) is binomial. (b) Find \(P(X = 2)\). (c) Find \(P(X \ge 3)\). (d) Find and interpret the mean and standard deviation of \(X\).
Show answer
(a) Binary (delayed or not), Independent (given), fixed Number of trials \(n = 12\), Same probability \(p = 0.15\) each flight: \(X \sim B(12, 0.15)\). (b) \(P(X = 2) = \binom{12}{2}(0.15)^2(0.85)^{10} = 66(0.0225)(0.1969) \approx 0.292\) (binompdf(12, 0.15, 2)). (c) \(P(X \ge 3) = 1 - P(X \le 2) = 1 - 0.736 = 0.264\) (1 − binomcdf(12, 0.15, 2)). (d) \(\mu_X = 12(0.15) = 1.8\) flights and \(\sigma_X = \sqrt{12(0.15)(0.85)} \approx 1.24\) flights: over many sets of 12 flights, about 1.8 would be delayed on average, typically within about 1.24 of that.
A basketball player makes 40% of her three-point attempts, independently. She shoots until she makes one. Let \(X\) = the number of the attempt on which she makes her first shot. (a) Explain why \(X\) is geometric rather than binomial. (b) Find \(P(X = 3)\). (c) Find \(P(X \le 3)\) and \(P(X \gt 4)\). (d) Find the expected number of attempts.
Show answer
(a) The trials are binary, independent, with the same \(p = 0.4\), but there is no fixed number of trials: the count of attempts until the first success is the random variable. (b) \(P(X = 3) = (0.6)^2(0.4) = 0.144\): miss, miss, make. (c) \(P(X \le 3) = 1 - (0.6)^3 = 1 - 0.216 = 0.784\) (geometcdf(0.4, 3)); \(P(X \gt 4) = (0.6)^4 = 0.1296\), the probability the first four attempts all miss. (d) \(\mu_X = 1/p = 1/0.4 = 2.5\) attempts.
Events \(A\) and \(B\) have \(P(A) = 0.4\), \(P(B) = 0.5\), and \(P(A \cap B) = 0.2\). Which statement is true?
Mutually exclusive means the events cannot happen together, which requires \(P(A \cap B) = 0\). Here \(P(A \cap B) = 0.2\), so the events overlap; mutually exclusive and independent are different ideas, and these events are the second, not the first.
\(P(A \mid B) = \dfrac{P(A \cap B)}{P(B)} = \dfrac{0.2}{0.5} = 0.4 = P(A)\), so knowing \(B\) occurred does not change the chance of \(A\); equivalently \(P(A)P(B) = (0.4)(0.5) = 0.2 = P(A \cap B)\).
0.9 is \(P(A) + P(B)\) with the overlap forgotten, which counts the intersection twice. The general addition rule subtracts it: \(P(A \cup B) = 0.4 + 0.5 - 0.2 = 0.7\).
0.5 is \(P(B)\), not \(P(A \mid B)\). The conditional probability is \(P(A \cap B)/P(B) = 0.2/0.5 = 0.4\); mixing up the conditioning event with the event of interest is the usual slip.
Lesson 5.1 · Unit 5 · CED topics 5.1–5.2
Normal calculations revisited and combining normal random variables
Unit 4 gave you the rules for the mean and variance of a sum or difference
of random variables. One more fact makes those rules powerful: if the
variables are normal and independent, their sum or difference is normal
too. That means you can answer probability questions about totals and
differences with the same z-score machinery you used in Unit 1.
Rule
If \(X \sim N(\mu_X, \sigma_X)\) and \(Y \sim N(\mu_Y, \sigma_Y)\) are
independent, then
\[X + Y \sim N\!\left(\mu_X + \mu_Y,\; \sqrt{\sigma_X^2 + \sigma_Y^2}\right)
\qquad
X - Y \sim N\!\left(\mu_X - \mu_Y,\; \sqrt{\sigma_X^2 + \sigma_Y^2}\right).\]
Variances add in both cases, even for a difference. Standard deviations
never add directly.
Worked example · A sum
Maya's drive to school takes \(T \sim N(24, 4)\) minutes and the walk from
the parking lot takes \(W \sim N(6, 1.5)\) minutes, independently. What is
the probability her total trip exceeds 35 minutes?
Let \(S = T + W\). Then \(\mu_S = 24 + 6 = 30\) and
\(\sigma_S = \sqrt{4^2 + 1.5^2} = \sqrt{18.25} \approx 4.272\) minutes, so
\(S \sim N(30, 4.272)\).
\[z = \frac{35 - 30}{4.272} \approx 1.17, \qquad P(S \gt 35) = P(Z \gt 1.17) \approx 0.1209.\]
About a 12% chance the trip runs over 35 minutes.
Worked example · A difference
Brand A cereal boxes have weights \(N(510, 8)\) grams; Brand B boxes are
\(N(500, 6)\) grams. If one box of each is chosen independently, what is
the probability the Brand A box is lighter?
Let \(D = A - B\). \(\mu_D = 510 - 500 = 10\) and
\(\sigma_D = \sqrt{8^2 + 6^2} = \sqrt{100} = 10\) grams. "A is lighter"
means \(D \lt 0\):
\[z = \frac{0 - 10}{10} = -1, \qquad P(D \lt 0) = P(Z \lt -1) \approx 0.1587.\]
Calculator
normalcdf(lower, upper, μ, σ): for the first example,
normalcdf(35, 1E99, 30, 4.272) returns 0.1209. On the exam,
write the distribution of the new variable (\(S \sim N(30, 4.272)\)) and
the probability statement before you type; a bare calculator number earns
no credit.
Try it
A lab technician takes two independent readings of a sample's mass, each
distributed \(N(1.20, 0.05)\) grams. Find the probability that the two
readings sum to more than 2.50 grams.
Take a random sample, compute a statistic, and you get one number. Take a
different random sample and you get a different number. Inference is
possible only because that sample-to-sample variation has a predictable
pattern, and that pattern is the sampling distribution.
Definitions
A parameter (\(\mu\), \(p\), \(\sigma\)) is a number that
describes a population. A statistic (\(\bar{x}\),
\(\hat{p}\), \(s\)) is a number computed from a sample. The
sampling distribution of a statistic is the distribution
of its values in all possible samples of the same size from the same
population.
A statistic is an unbiased estimator if the mean of its
sampling distribution equals the parameter it estimates. The
variability of a statistic is the spread of its sampling
distribution; larger samples give less variability, and (as long as the
population is at least 10 times the sample) the population size barely
matters.
Worked example · A tiny population
A population consists of four values: 2, 4, 6, 8, so \(\mu = 5\) and the
population range is 6. List every sample of size 2 (without replacement)
and compute the sample mean and sample range.
Sample
{2, 4}
{2, 6}
{2, 8}
{4, 6}
{4, 8}
{6, 8}
Sample mean
3
4
5
5
6
7
Sample range
2
4
6
2
4
2
Mean of the six sample means: \(\dfrac{3 + 4 + 5 + 5 + 6 + 7}{6} = 5 = \mu\),
so \(\bar{x}\) is unbiased. Mean of the six sample ranges:
\(\dfrac{20}{6} \approx 3.33\), well below 6: the sample range is a
biased estimator that systematically underestimates.
Worked example · Bias vs. variability
Two polling firms each repeatedly sample 1,000 voters when the true
support for a candidate is \(p = 0.52\). Firm A's estimates cluster
tightly around 0.47; Firm B's are centered at 0.52 but range from 0.45 to
0.59. Firm A has low variability but is biased (probably a flawed
sampling frame). Firm B is unbiased but highly variable. Taking a bigger
sample would shrink Firm B's spread; it would do nothing for Firm A's
bias.
Exam tip: when asked to "describe the sampling distribution," give shape,
center, and spread: all three, in context.
Try it
A statistic used to estimate a population proportion has a sampling
distribution with mean 0.42 when the true proportion is 0.40. Is the
statistic unbiased? If the sample size is quadrupled, what happens to its
bias and its variability?
Show answer
It is biased: on average it overestimates \(p\) by 0.02. Quadrupling
\(n\) roughly halves the variability (spread scales like
\(1/\sqrt{n}\)) but leaves the bias unchanged: the estimates cluster
more tightly around the wrong value.
Lesson 5.3 · Unit 5 · CED topic 5.5
Sampling distribution of a sample proportion
A sample proportion \(\hat{p}\) is a count of successes divided by \(n\), and
that count is binomial. Divide the binomial mean and standard deviation by
\(n\) and you have the center and spread of \(\hat{p}\). When the counts are
large enough, the shape is approximately normal, and probability questions
about \(\hat{p}\) become z-score problems.
Formula and conditions
For an SRS of size \(n\) from a population with proportion \(p\):
\[\mu_{\hat p} = p, \qquad \sigma_{\hat p} = \sqrt{\frac{p(1-p)}{n}}.\]
The standard deviation formula requires the 10% condition
(\(n \le \frac{1}{10}N\)). The shape is approximately normal when the
Large Counts condition holds: \(np \ge 10\) and
\(n(1-p) \ge 10\).
Worked example
At a university with 28,000 students, 35% commute. A researcher takes an
SRS of 200 students. Describe the sampling distribution of \(\hat{p}\), the
sample proportion who commute, and find \(P(\hat{p} \ge 0.40)\).
Shape: \(np = 200(0.35) = 70 \ge 10\) and
\(n(1-p) = 130 \ge 10\), so approximately normal.
Center: \(\mu_{\hat p} = 0.35\).
Spread: \(200 \le 2{,}800\), so
\(\sigma_{\hat p} = \sqrt{\dfrac{0.35(0.65)}{200}} \approx 0.0337\).
\[z = \frac{0.40 - 0.35}{0.0337} \approx 1.48, \qquad P(\hat p \ge 0.40) = P(Z \ge 1.48) \approx 0.0691.\]
In about 7% of all samples of 200, at least 40% of the students would be
commuters.
Calculator
normalcdf(0.40, 1E99, 0.35, 0.0337) = 0.0691. Better: type the
standard deviation unrounded, normalcdf(0.40, 1E99, 0.35, √(0.35·0.65/200)),
and round only the final probability.
Exam tip: graders look for all three conditions-and-parameters in words.
Write "approximately normal because 70 and 130 are both at least 10," not
just "normal."
Try it
Suppose 62% of the 40,000 adults in a city approve of a new transit plan.
In an SRS of 150 adults, what is the probability that fewer than 55% of
the sample approve?
Show answer
\(np = 93\) and \(n(1-p) = 57\), both \(\ge 10\), so \(\hat{p}\) is
approximately normal with mean 0.62 and
\(\sigma_{\hat p} = \sqrt{0.62(0.38)/150} \approx 0.0396\)
(150 is under 10% of 40,000). \(z = \dfrac{0.55 - 0.62}{0.0396} \approx -1.77\),
so \(P(\hat p \lt 0.55) \approx 0.0387\).
Lesson 5.4 · Unit 5 · CED topic 5.6
Sampling distribution of a difference of two proportions
Most interesting questions compare two groups: do seniors at School A work
more than seniors at School B? The statistic is \(\hat{p}_1 - \hat{p}_2\),
and its sampling distribution follows directly from the rules for
differences of independent random variables.
Formula and conditions
For independent random samples of sizes \(n_1\) and \(n_2\) from
populations with proportions \(p_1\) and \(p_2\):
\[\mu_{\hat p_1 - \hat p_2} = p_1 - p_2, \qquad
\sigma_{\hat p_1 - \hat p_2} = \sqrt{\frac{p_1(1-p_1)}{n_1} + \frac{p_2(1-p_2)}{n_2}}.\]
Conditions: the samples are independent random samples (or two groups in
a randomized experiment); each sample satisfies the 10% condition; and
Large Counts holds for both groups: \(n_1p_1\), \(n_1(1-p_1)\),
\(n_2p_2\), \(n_2(1-p_2)\) all at least 10.
Worked example
At School A, 60% of the 1,400 seniors have part-time jobs; at School B,
50% of its 1,200 seniors do. A counselor takes an SRS of 100 seniors from
each school. What is the probability that the sample from School B shows
a higher proportion with jobs than the sample from School A?
Shape: counts \(60, 40, 50, 50\) are all \(\ge 10\), so
\(\hat p_A - \hat p_B\) is approximately normal.
Center: \(0.60 - 0.50 = 0.10\).
Spread: both samples are under 10% of their schools, so
\[\sigma_{\hat p_A - \hat p_B} = \sqrt{\frac{0.6(0.4)}{100} + \frac{0.5(0.5)}{100}} = \sqrt{0.0049} = 0.07.\]
"B higher than A" means \(\hat p_A - \hat p_B \lt 0\):
\[z = \frac{0 - 0.10}{0.07} \approx -1.43, \qquad P(Z \lt -1.43) \approx 0.0766.\]
Even though School A's true rate is 10 points higher, roughly 1 sample
pair in 13 would point the other way. That's the kind of chance variation
inference has to account for.
Common error: subtracting the standard deviations, or forgetting to add
the variances for a difference. The variance of a difference is a
sum.
Try it
In one large city 30% of households own a dog; in another, 24% do.
Independent SRSs of 250 and 300 households are taken. Find the probability
that the difference in sample proportions (first city minus second)
exceeds 0.10.
Show answer
Counts 75, 175, 72, 228 are all \(\ge 10\). Mean \(= 0.06\);
\(\sigma = \sqrt{0.3(0.7)/250 + 0.24(0.76)/300} \approx 0.0381\).
\(z = \dfrac{0.10 - 0.06}{0.0381} \approx 1.05\), so
\(P(\hat p_1 - \hat p_2 \gt 0.10) \approx 0.1466\).
Lesson 5.5 · Unit 5 · CED topic 5.7
Sampling distribution of a sample mean and the Central Limit Theorem
Averages are less variable than individual observations, and, this is the
remarkable part, averages of enough observations are approximately normal
no matter what shape the population has. That second fact is the Central
Limit Theorem, and it's the reason so much of inference works.
Formula and conditions
For an SRS of size \(n\) from a population with mean \(\mu\) and standard
deviation \(\sigma\):
\[\mu_{\bar x} = \mu, \qquad \sigma_{\bar x} = \frac{\sigma}{\sqrt{n}} \quad(\text{requires the 10\% condition}).\]
Shape: if the population is normal, \(\bar{x}\) is exactly
normal for any \(n\). If not, the Central Limit Theorem
says \(\bar{x}\) is approximately normal when \(n \ge 30\).
Worked example · One box vs. sixteen boxes
Cereal box weights are \(N(510, 8)\) grams. (a) Find the probability one
box weighs under 500 g. (b) Find the probability the mean of an SRS of 16
boxes is under 500 g.
(a) \(z = \dfrac{500 - 510}{8} = -1.25\), so \(P \approx 0.1056\).
(b) \(\bar{x}\) is normal (the population is) with mean 510 and
\(\sigma_{\bar x} = 8/\sqrt{16} = 2\). \(z = \dfrac{500 - 510}{2} = -5\), so
\(P \approx 0.0000003\). A single light box is routine; a light
average of 16 is essentially impossible unless the machine has
drifted.
Worked example · A skewed population
Hold times at a call center are strongly right-skewed with mean 4.2
minutes and standard deviation 3.5 minutes. Find the probability that the
mean hold time of 50 randomly selected calls exceeds 5 minutes.
You cannot find \(P(\text{one call} \gt 5)\): the population isn't normal.
But \(n = 50 \ge 30\), so by the CLT \(\bar{x}\) is approximately normal
with mean 4.2 and \(\sigma_{\bar x} = 3.5/\sqrt{50} \approx 0.495\).
\[z = \frac{5 - 4.2}{0.495} \approx 1.62, \qquad P(\bar x \gt 5) \approx 0.0530.\]
Calculator
normalcdf(5, 1E99, 4.2, 3.5/√50) = 0.0530. Always divide
\(\sigma\) by \(\sqrt{n}\) inside the command: using \(\sigma\) alone is
the most common Unit 5 error.
Try it
A tire model's tread life is right-skewed with mean 45,000 miles and
standard deviation 6,000 miles. Find the probability that the mean tread
life of 36 randomly selected tires is less than 43,000 miles. Explain why
the calculation is valid.
Show answer
\(n = 36 \ge 30\), so the CLT makes \(\bar{x}\) approximately normal
despite the skew. \(\sigma_{\bar x} = 6000/\sqrt{36} = 1000\);
\(z = \dfrac{43000 - 45000}{1000} = -2\), so
\(P(\bar x \lt 43000) \approx 0.0228\).
Lesson 5.6 · Unit 5 · CED topic 5.8
Sampling distribution of a difference of two means
The last piece of the unit combines two ideas you already have: the
sampling distribution of \(\bar{x}\), and the rule that variances of
independent variables add. Put them together and you can describe
\(\bar{x}_1 - \bar{x}_2\), the statistic behind every two-sample comparison
of means in Unit 7.
Formula and conditions
For independent random samples of sizes \(n_1\) and \(n_2\):
\[\mu_{\bar x_1 - \bar x_2} = \mu_1 - \mu_2, \qquad
\sigma_{\bar x_1 - \bar x_2} = \sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}.\]
Conditions: independent random samples (or randomized groups); the 10%
condition for each sample; and the shape is approximately normal if each
population is normal or each sample has \(n \ge 30\) (Normal/Large Sample
condition, applied to both groups).
Worked example
Brand X batteries last a mean of 52 hours with standard deviation 4 hours;
Brand Y batteries last a mean of 50 hours with standard deviation 5 hours.
A tester takes independent random samples of 40 Brand X and 50 Brand Y
batteries. What is the probability the Brand Y sample mean is higher than
the Brand X sample mean?
Shape: both samples have \(n \ge 30\), so
\(\bar x_X - \bar x_Y\) is approximately normal.
Center: \(52 - 50 = 2\) hours.
Spread:
\[\sigma_{\bar x_X - \bar x_Y} = \sqrt{\frac{4^2}{40} + \frac{5^2}{50}} = \sqrt{0.4 + 0.5} \approx 0.949 \text{ hours}.\]
"Y higher" means \(\bar x_X - \bar x_Y \lt 0\):
\[z = \frac{0 - 2}{0.949} \approx -2.11, \qquad P(Z \lt -2.11) \approx 0.0175.\]
Fewer than 2% of sample pairs would rank the brands backwards. Compare
this with Lesson 5.4's job example (7.7%): larger samples and a bigger
gap relative to the spread make a reversal much rarer.
Exam tip: the formula sheet gives \(\sigma_{\bar x_1 - \bar x_2}\), but you
must still state the conditions that justify using a normal model. "Both
sample sizes are at least 30, so the CLT applies" is the sentence graders
want.
Try it
Commute times in City 1 have mean 28 minutes and standard deviation 9;
in City 2, mean 25 minutes and standard deviation 8. Independent random
samples of 60 and 45 commuters are taken. Find the probability that the
City 1 sample mean exceeds the City 2 sample mean by more than 5 minutes.
Show answer
Both \(n \ge 30\), so approximately normal. Mean \(= 3\);
\(\sigma = \sqrt{81/60 + 64/45} \approx 1.665\).
\(z = \dfrac{5 - 3}{1.665} \approx 1.20\), so
\(P(\bar x_1 - \bar x_2 \gt 5) \approx 0.1148\).
Unit 5 practice · 10 problems
Unit 5 practice: Sampling Distributions
Ten problems covering the whole unit, in roughly exam order. For every probability,
write the distribution of the statistic (shape, center, spread) and the z-score before
you reach for normalcdf: that's where the points are.
A machine cuts rods of type A with lengths \(N(50.0, 0.4)\) mm and rods of type B with lengths \(N(30.0, 0.3)\) mm, independently. One rod of each type is joined end to end. What is the probability the assembled length exceeds 80.8 mm?
Show answer
Let \(L = A + B\). \(\mu_L = 50.0 + 30.0 = 80.0\) mm and \(\sigma_L = \sqrt{0.4^2 + 0.3^2} = \sqrt{0.25} = 0.5\) mm; since \(A\) and \(B\) are independent normals, \(L \sim N(80.0, 0.5)\). \[z = \frac{80.8 - 80.0}{0.5} = 1.6, \qquad P(L \gt 80.8) = P(Z \gt 1.6) \approx 0.0548.\] normalcdf(80.8, 1E99, 80, 0.5) = 0.0548.
Heights of adult men are approximately \(N(70, 3)\) inches and heights of adult women approximately \(N(64.5, 2.5)\) inches. A man and a woman are selected independently at random. What is the probability the woman is taller than the man?
Show answer
Let \(D = M - W\). \(\mu_D = 70 - 64.5 = 5.5\) in; \(\sigma_D = \sqrt{3^2 + 2.5^2} = \sqrt{15.25} \approx 3.905\) in (variances add for a difference). \(D\) is normal because \(M\) and \(W\) are independent normals. "Woman taller" means \(D \lt 0\): \[z = \frac{0 - 5.5}{3.905} \approx -1.41, \qquad P(D \lt 0) \approx 0.0795.\] About 8% of randomly paired couples would have the woman taller.
A population consists of three values: 3, 6, and 9, so \(\mu = 6\) and the population maximum is 9. List all samples of size 2 drawn without replacement, compute the sample mean and sample maximum for each, and determine whether each statistic is an unbiased estimator of the corresponding parameter.
Show answer
Samples: {3, 6}, {3, 9}, {6, 9}. Sample means: 4.5, 6, 7.5; their mean is \((4.5 + 6 + 7.5)/3 = 6 = \mu\), so \(\bar x\) is an unbiased estimator of \(\mu\). Sample maximums: 6, 9, 9; their mean is \(24/3 = 8 \ne 9\), so the sample maximum is a biased estimator that systematically underestimates the population maximum (it can never exceed 9 and usually falls short).
Four statistics are proposed for estimating a parameter. Their sampling distributions: I is centered at the parameter with a wide spread; II is centered at the parameter with a narrow spread; III is centered well above the parameter with a narrow spread; IV is centered well above the parameter with a wide spread. Which statistic is unbiased but has high variability?
“Unbiased” means the sampling distribution is centered at the parameter (true for I and II), and “high variability” means a wide spread (true for I and IV); only I has both. A larger sample would shrink its spread without moving its center.
II is unbiased, but it has low variability: it is the ideal estimator, precise and centered on the target. The question asks for the one that is centered correctly but widely scattered.
III has low variability but is biased: its center sits well above the parameter, so it is consistently wrong in the same direction. Precision is not accuracy, and a larger sample would not fix the bias.
IV does have the high variability the question describes, but it is also biased, since it is centered above the parameter. It combines the worst of both properties rather than being the unbiased-but-variable case.
About 12% of adults are left-handed. A researcher takes an SRS of 300 adults from a large city. (a) Describe the sampling distribution of \(\hat p\), the sample proportion who are left-handed. (b) Find \(P(\hat p \ge 0.15)\). (c) Explain why part (b) could not be done the same way with an SRS of 50 adults.
Show answer
(a) Shape: approximately normal, because \(np = 300(0.12) = 36 \ge 10\) and \(n(1-p) = 264 \ge 10\). Center: \(\mu_{\hat p} = 0.12\). Spread: 300 is under 10% of the city's adults, so \(\sigma_{\hat p} = \sqrt{0.12(0.88)/300} \approx 0.0188\). (b) \(z = \dfrac{0.15 - 0.12}{0.0188} \approx 1.60\), so \(P(\hat p \ge 0.15) \approx 0.0549\) (normalcdf(0.15, 1E99, 0.12, √(0.12·0.88/300))). (c) With \(n = 50\), \(np = 6 \lt 10\); the Large Counts condition fails, the sampling distribution is skewed right, and a normal approximation is not justified: you would use the binomial distribution directly.
Among the customers of Store 1, 45% are members of its rewards program; at Store 2, 40% are. Independent SRSs of 200 customers from Store 1 and 250 from Store 2 are taken. What is the probability that the sample from Store 2 shows a higher proportion of members than the sample from Store 1?
Show answer
Let the statistic be \(\hat p_1 - \hat p_2\). Shape: the counts \(200(0.45) = 90\), \(110\), \(250(0.40) = 100\), and \(150\) are all at least 10, so approximately normal. Center: \(0.45 - 0.40 = 0.05\). Spread (both samples under 10% of their stores' customers): \[\sigma_{\hat p_1 - \hat p_2} = \sqrt{\frac{0.45(0.55)}{200} + \frac{0.40(0.60)}{250}} \approx 0.0469.\] "Store 2 higher" means \(\hat p_1 - \hat p_2 \lt 0\): \(z = \dfrac{0 - 0.05}{0.0469} \approx -1.07\), so the probability is about \(0.143\). Even with a real 5-point gap, about 1 sample pair in 7 would point the wrong way.
A filling machine puts \(N(20.0, 0.3)\) ounces of juice in each bottle. (a) Find the probability that one randomly selected bottle contains less than 19.8 oz. (b) Find the probability that the mean contents of an SRS of 9 bottles is less than 19.8 oz. (c) Explain why the normal calculation in (b) is valid even though \(n = 9 \lt 30\).
Show answer
(a) \(z = \dfrac{19.8 - 20.0}{0.3} \approx -0.67\), so \(P \approx 0.2525\). (b) \(\bar x\) has mean 20.0 and \(\sigma_{\bar x} = 0.3/\sqrt{9} = 0.1\) oz; \(z = \dfrac{19.8 - 20.0}{0.1} = -2\), so \(P(\bar x \lt 19.8) \approx 0.0228\). (c) When the population itself is normal, the sampling distribution of \(\bar x\) is exactly normal for any sample size; the \(n \ge 30\) rule (Central Limit Theorem) is only needed when the population shape is unknown or non-normal. One light bottle is common (25%); a light average of nine is rare (2%).
Tips left at a restaurant are strongly right-skewed with mean $5.20 and standard deviation $3.10. A server waits on 40 randomly selected tables in a week. (a) Find the probability that the server's mean tip exceeds $6.00, and justify the method. (b) Explain why you cannot use the same method to find the probability that a single tip exceeds $6.00.
Show answer
(a) \(n = 40 \ge 30\), so by the Central Limit Theorem \(\bar x\) is approximately normal despite the skewed population, with mean $5.20 and \(\sigma_{\bar x} = 3.10/\sqrt{40} \approx 0.490\). \[z = \frac{6.00 - 5.20}{0.490} \approx 1.63, \qquad P(\bar x \gt 6.00) \approx 0.0513.\] (b) A single tip comes from the population distribution, which is strongly skewed, not normal: the CLT says nothing about individual observations. Without knowing the population's shape, \(P(X \gt 6)\) cannot be computed.
Scores on a district math assessment have mean 68 and standard deviation 10 at School 1, and mean 65 with standard deviation 12 at School 2. Independent random samples of 50 students from School 1 and 40 from School 2 are taken. Find the probability that the School 1 sample mean exceeds the School 2 sample mean by more than 5 points.
Show answer
Shape: both \(n \ge 30\), so \(\bar x_1 - \bar x_2\) is approximately normal by the CLT. Center: \(68 - 65 = 3\) points. Spread: \[\sigma_{\bar x_1 - \bar x_2} = \sqrt{\frac{10^2}{50} + \frac{12^2}{40}} = \sqrt{2 + 3.6} \approx 2.366 \text{ points}.\] \(z = \dfrac{5 - 3}{2.366} \approx 0.85\), so \(P(\bar x_1 - \bar x_2 \gt 5) \approx 0.199\).
A researcher plans to estimate the mean commute time of a city's 400,000 workers with an SRS. (a) If the sample size is increased from 100 to 400, what happens to the mean and to the standard deviation of the sampling distribution of \(\bar x\)? (b) Would the answer change if the city had 4,000,000 workers instead? (c) Commute times are right-skewed; what can you say about the shape of the sampling distribution of \(\bar x\) for \(n = 400\)?
Show answer
(a) The mean stays at \(\mu\): \(\bar x\) is unbiased regardless of \(n\). The standard deviation \(\sigma/\sqrt{n}\) is cut in half, since \(\sqrt{400} = 2\sqrt{100}\); quadrupling the sample halves the spread. (b) No. As long as the sample is under 10% of the population, the population size has essentially no effect on \(\sigma_{\bar x}\): 400 out of 400,000 and 400 out of 4,000,000 give the same spread. (c) With \(n = 400 \ge 30\), the Central Limit Theorem says the sampling distribution of \(\bar x\) is approximately normal even though individual commute times are skewed.
Lesson 6.1 · Unit 6 · CED topics 6.1–6.2
Confidence intervals for a population proportion
A sample proportion is a single guess at \(p\), and Unit 5 told you how far
off that guess typically is. A confidence interval turns that knowledge
into a range of plausible values for \(p\), with a stated level of
confidence. Every interval in this course has the same shape: estimate
± (critical value)(standard error).
Formula and conditions
A one-sample z interval for \(p\):
\[\hat p \pm z^* \sqrt{\frac{\hat p(1 - \hat p)}{n}}.\]
Random: the data come from a random sample.
10%: \(n \le \frac{1}{10}N\) when sampling without replacement.
Large Counts: at least 10 successes and 10 failures in the
sample: \(n\hat p \ge 10\) and \(n(1 - \hat p) \ge 10\).
Critical values come from invNorm: for 95% confidence,
\(z^* = \text{invNorm}(0.975, 0, 1) = 1.960\). Memorize 1.645 (90%),
1.960 (95%), and 2.576 (99%).
Worked example · Four-step write-up
In an SRS of 400 U.S. high school seniors, 148 said they applied to five
or more colleges. Construct and interpret a 95% confidence interval.
State: \(p\) = the true proportion of all U.S. high school
seniors who applied to five or more colleges. We want a 95% confidence
interval for \(p\).
Plan: one-sample z interval for \(p\). Random: SRS ✓.
10%: 400 is far less than 10% of the several million U.S. seniors ✓.
Large Counts: 148 successes and 252 failures are both at least 10 ✓.
Conclude: We are 95% confident that the true proportion of
all U.S. high school seniors who applied to five or more colleges is
between 0.323 and 0.417.
Calculator
STAT → TESTS → A:1-PropZInt. Enter x = 148, n = 400,
C-Level = .95, Calculate → (0.3227, 0.4173). The x must be a whole-number
count; if you're given \(\hat p\), multiply by \(n\) and round. Write the
formula with numbers substituted even when you use the calculator.
Try it
A random sample of 250 customers at a large grocery chain found 60 who
brought reusable bags. Construct a 90% confidence interval for the true
proportion of the chain's customers who bring reusable bags.
Show answer
Conditions: random sample; 250 is under 10% of the chain's customers;
60 successes and 190 failures ≥ 10. \(\hat p = 0.24\);
\(0.24 \pm 1.645\sqrt{0.24(0.76)/250} = 0.24 \pm 1.645(0.0270) = 0.24 \pm 0.044\),
so (0.196, 0.284). We are 90% confident that the true proportion of the
chain's customers who bring reusable bags is between 0.196 and 0.284.
Lesson 6.2 · Unit 6 · CED topic 6.3
Interpreting confidence intervals and choosing a sample size
Computing an interval is the easy part. The exam cares far more about
whether you can say what it means, and what it doesn't. Two different
sentences are involved: one interprets the interval, the other interprets
the confidence level.
Interpretations
The interval: "We are 95% confident that the true
[parameter, in context] is between \(a\) and \(b\)."
The level: "If we took many random samples of this size
and built a 95% interval from each, about 95% of those intervals would
capture the true [parameter]." Confidence is a statement about the
method, not about one interval.
Wrong: "There is a 95% probability that \(p\) is in the
interval" (\(p\) is fixed; either it's in there or it isn't). "95% of the
sample falls in the interval." "95% of sample proportions will fall in
this interval."
Formula · Sample size
The margin of error is \(ME = z^*\sqrt{\hat p(1-\hat p)/n}\). To guarantee
a margin of error no larger than a target, solve for \(n\):
\[n \ge \left(\frac{z^*}{ME}\right)^2 \hat p(1 - \hat p),\]
using a prior estimate for \(\hat p\) if you have one and \(\hat p = 0.5\)
(the worst case, since it maximizes \(\hat p(1-\hat p)\)) if you don't.
Always round up.
Worked example · No prior estimate
A pollster wants to estimate the proportion of voters who support a bond
measure to within ±3 percentage points with 95% confidence. How many
voters must be sampled?
\(n \ge \left(\dfrac{1.960}{0.03}\right)^2(0.5)(0.5) = 1067.1\), so
\(n = 1068\). Notice the 3-point margin costs about a thousand people;
halving it to 1.5 points would quadruple the sample.
Worked example · With a prior estimate
A clinic knows from past records that about 20% of patients miss
appointments. To estimate this year's rate to within ±4 points with 90%
confidence: \(n \ge \left(\dfrac{1.645}{0.04}\right)^2(0.2)(0.8) = 270.6\),
so \(n = 271\). The prior estimate cut the required sample sharply
compared with using 0.5 (which would give 423).
The margin of error covers only random sampling variability. It says
nothing about undercoverage, nonresponse, or badly worded questions: a
biased sample of 10,000 has a tiny margin of error and a wrong answer.
Try it
A school board wants a 99% confidence interval for the proportion of
parents who favor a later start time, with a margin of error of at most
0.05. No prior estimate is available. What sample size is needed?
Show answer
\(n \ge \left(\dfrac{2.576}{0.05}\right)^2(0.5)(0.5) = 663.5\), so at
least \(n = 664\) parents.
Lesson 6.3 · Unit 6 · CED topics 6.4–6.6
Significance tests for a population proportion
A confidence interval estimates a parameter. A significance test weighs the
evidence against a specific claim about it. The logic: assume the claim is
true, ask how surprising the sample would be, and if it's surprising
enough, reject the claim.
Hypotheses, statistic, and conditions
\(H_0\!: p = p_0\) versus \(H_a\!: p \lt p_0\), \(p \gt p_0\), or
\(p \ne p_0\). Hypotheses are always about the parameter, never about
\(\hat p\). The test statistic is
\[z = \frac{\hat p - p_0}{\sqrt{\dfrac{p_0(1 - p_0)}{n}}},\]
and the p-value is the probability, assuming \(H_0\) is
true, of a statistic at least as extreme as the one observed. Conditions
are Random, 10%, and Large Counts: checked with \(p_0\), not \(\hat p\):
\(np_0 \ge 10\) and \(n(1-p_0) \ge 10\).
Worked example · Four-step write-up
A school district claims that 80% of its graduates enroll in a four-year
college. A counselor suspects the figure is lower and takes an SRS of 150
graduates from the district's several thousand recent graduates; 108
enrolled in a four-year college. Test at \(\alpha = 0.05\).
State: \(p\) = the true proportion of the district's
recent graduates who enrolled in a four-year college.
\(H_0\!: p = 0.80\) (the district's claim is accurate);
\(H_a\!: p \lt 0.80\) (the true proportion is lower than claimed).
\(\alpha = 0.05\).
Plan: one-sample z test for \(p\). Random: SRS ✓.
10%: 150 is less than 10% of several thousand graduates ✓.
Large Counts: \(150(0.80) = 120\) and \(150(0.20) = 30\), both \(\ge 10\) ✓.
Conclude: Because the p-value of 0.0072 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the
true proportion of the district's recent graduates who enrolled in a
four-year college is less than 0.80.
Worked example · Two-sided
A candy maker says 20% of its candies are red; a random sample of 250 has
62 red (\(\hat p = 0.248\)). With \(H_a\!: p \ne 0.20\),
\(z = \dfrac{0.248 - 0.20}{\sqrt{0.2(0.8)/250}} \approx 1.90\) and the
two-sided p-value is \(2P(Z \ge 1.90) \approx 0.0578\). Since
\(0.0578 \gt 0.05\), we fail to reject \(H_0\): not convincing evidence
that the proportion of red candies differs from 0.20.
Calculator
STAT → TESTS → 5:1-PropZTest: p₀ = .8, x = 108,
n = 150, prop: <p₀ → z = −2.449, p = 0.0072. By hand the p-value is
normalcdf(−1E99, −2.449, 0, 1). For a two-sided test the calculator
doubles the tail for you; by hand, you must.
Try it
A streaming service believes 55% of its subscribers watch on a phone. A
random sample of 500 subscribers finds 300 who do. Is there convincing
evidence at \(\alpha = 0.05\) that more than 55% watch on a phone?
Show answer
\(H_0\!: p = 0.55\), \(H_a\!: p \gt 0.55\). Conditions: random sample,
500 under 10% of subscribers, \(500(0.55) = 275\) and \(500(0.45) = 225\)
both ≥ 10. \(\hat p = 0.60\);
\(z = \dfrac{0.60 - 0.55}{\sqrt{0.55(0.45)/500}} = \dfrac{0.05}{0.0222} \approx 2.25\);
p-value \(\approx 0.0123\). Since \(0.0123 \lt 0.05\), reject \(H_0\):
convincing evidence that more than 55% of subscribers watch on a phone.
Lesson 6.4 · Unit 6 · CED topic 6.7
Conclusions, errors, and power
A significance test is a decision made under uncertainty, so it can be
wrong in two ways. Knowing which error matters more tells you how to
choose \(\alpha\); knowing what "power" means tells you how to design a
study that can actually detect what you're looking for.
Definitions
A Type I error is rejecting \(H_0\) when \(H_0\) is
actually true; its probability is \(\alpha\). A Type II
error is failing to reject \(H_0\) when \(H_a\) is actually true;
its probability is \(\beta\). The power of a test is
\(1 - \beta\): the probability of correctly rejecting a false \(H_0\)
for a particular alternative value of the parameter.
Power increases with a larger sample size, a larger \(\alpha\), and a
larger effect size (a true value farther from \(p_0\)). Only the first
raises power without also raising the chance of a Type I error, but it
costs time and money.
Worked example · Errors in context
A district will adopt an expensive reading program if a pilot study gives
convincing evidence that more than 60% of students improve.
\(H_0\!: p = 0.60\), \(H_a\!: p \gt 0.60\).
Type I: the data convince the district that more than
60% improve when in fact they don't. Consequence: money spent on a
program that isn't better than the current one.
Type II: the program really does help more than 60% of
students, but the pilot fails to show it. Consequence: students miss a
program that works. Graders want both the decision and the consequence,
in context.
Worked example · Computing power
Suppose \(n = 100\) and \(\alpha = 0.05\). The test rejects when
\(z \gt 1.645\), i.e. when
\(\hat p \gt 0.60 + 1.645\sqrt{0.6(0.4)/100} = 0.681\).
If the program truly helps 70% of students, \(\hat p\) is approximately
normal with mean 0.70 and standard deviation
\(\sqrt{0.7(0.3)/100} = 0.0458\), so
\[\text{Power} = P(\hat p \gt 0.681) = P\!\left(Z \gt \frac{0.681 - 0.70}{0.0458}\right) = P(Z \gt -0.42) \approx 0.66.\]
\(\beta \approx 0.34\): a one-in-three chance of missing a program that
works. With \(n = 400\) the power rises to about 0.995; with
\(\alpha = 0.10\) it rises to about 0.79.
Wording that loses points: never write "accept \(H_0\)": failing to
reject means the data are consistent with \(H_0\), not that it's true.
Conclusions are about the alternative: "there is (or is not) convincing
evidence that …".
Try it
A factory's quality team tests \(H_0\!: p = 0.02\) against
\(H_a\!: p \gt 0.02\), where \(p\) is the defect rate; if they reject, they
shut the line down for repairs. Describe a Type I and a Type II error and
their consequences, and explain how sampling 400 items instead of 100
affects the power of the test.
Show answer
Type I: shut the line down when the defect rate is really 2%; lost
production for no reason. Type II: keep running when the defect rate
has really risen; defective products ship to customers. Quadrupling
\(n\) halves the standard deviation of \(\hat p\), so a real increase
in defects is much more likely to be detected: power increases
(\(\beta\) decreases) while \(\alpha\) stays fixed.
Lesson 6.5 · Unit 6 · CED topics 6.8–6.9
Confidence intervals for a difference of two proportions
Comparing two groups is the most common inference task on the exam. The
parameter is now a difference, \(p_1 - p_2\), and the standard error is
Lesson 5.4's formula with the sample proportions standing in for the
unknown \(p_1\) and \(p_2\).
Formula and conditions
A two-sample z interval for \(p_1 - p_2\):
\[(\hat p_1 - \hat p_2) \pm z^* \sqrt{\frac{\hat p_1(1 - \hat p_1)}{n_1} + \frac{\hat p_2(1 - \hat p_2)}{n_2}}.\]
Random: two independent random samples, or two groups
formed by random assignment. 10%: for each sample (not
needed for a randomized experiment). Large Counts: the
observed successes and failures in both groups are all at least 10.
Worked example · Four-step write-up
In separate random samples from a large district, 126 of 300 ninth
graders and 75 of 250 twelfth graders reported sleeping at least 8 hours
on school nights. Construct and interpret a 95% confidence interval for
the difference in proportions.
State: \(p_1 - p_2\), where \(p_1\) and \(p_2\) are the
true proportions of the district's ninth and twelfth graders who sleep
at least 8 hours on school nights. 95% confidence.
Plan: two-sample z interval for \(p_1 - p_2\). Random:
independent random samples ✓. 10%: 300 and 250 are each under 10% of the
district's ninth and twelfth graders ✓. Large Counts: 126, 174, 75, 175
are all at least 10 ✓.
Conclude: We are 95% confident that the true difference
(ninth minus twelfth) in the proportions of the district's students who
sleep at least 8 hours on school nights is between 0.040 and 0.200.
Because the entire interval is above 0, there is convincing evidence
that ninth graders are more likely to get 8 hours.
Worked example · An interval that contains 0
A website randomly shows one of two ad designs to each of 400 visitors;
45 of 200 click design A and 38 of 200 click design B. A 90% interval for
\(p_A - p_B\) is \(0.035 \pm 1.645(0.0405) = 0.035 \pm 0.067\), or
(−0.032, 0.102). Since 0 is a plausible value of the difference, there is
not convincing evidence that the designs differ in click rate.
Calculator
STAT → TESTS → B:2-PropZInt: x1 = 126, n1 = 300, x2 = 75,
n2 = 250, C-Level = .95 → (0.0403, 0.1997). Keep track of which group is
"1" so your interpretation says which direction the difference runs.
Try it
Independent random samples of 400 adults in each of two states found
180 and 150 who had visited a national park in the past year. Construct
a 99% confidence interval for the difference in proportions and say
whether it gives convincing evidence of a difference.
Show answer
Counts 180, 220, 150, 250 all ≥ 10. \(\hat p_1 - \hat p_2 = 0.45 - 0.375 = 0.075\);
\(SE = \sqrt{0.45(0.55)/400 + 0.375(0.625)/400} \approx 0.0347\);
\(0.075 \pm 2.576(0.0347) = 0.075 \pm 0.089\), giving (−0.014, 0.164).
Since the interval contains 0, there is not convincing evidence of a
difference at the 99% level.
Lesson 6.6 · Unit 6 · CED topics 6.10–6.11
Significance tests for a difference of two proportions
The two-sample test asks whether an observed gap between two sample
proportions is too large to be chance. One twist: the null hypothesis says
the two population proportions are equal, so the best estimate of
that common proportion pools both samples together.
Hypotheses, pooled proportion, and statistic
\(H_0\!: p_1 - p_2 = 0\) (equivalently \(p_1 = p_2\)); \(H_a\!: p_1 - p_2 \gt 0\),
\(\lt 0\), or \(\ne 0\). The pooled (combined) proportion is
\(\hat p_c = \dfrac{X_1 + X_2}{n_1 + n_2}\), and
\[z = \frac{\hat p_1 - \hat p_2}{\sqrt{\hat p_c(1 - \hat p_c)\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}}.\]
Large Counts is checked with \(\hat p_c\): \(n_1\hat p_c\), \(n_1(1-\hat p_c)\),
\(n_2\hat p_c\), \(n_2(1-\hat p_c)\) all at least 10.
Worked example · Four-step write-up
A clinic randomly assigned 500 patients to receive either a text-message
reminder or no reminder about flu shots. Of the 250 who got texts, 105
were vaccinated; of the 250 who didn't, 80 were. Do text reminders
increase vaccination rates? Use \(\alpha = 0.05\).
State: \(p_1\) = the true proportion of patients like
these who would be vaccinated if sent a text reminder; \(p_2\) = the same
with no reminder. \(H_0\!: p_1 - p_2 = 0\) (the reminder makes no
difference); \(H_a\!: p_1 - p_2 \gt 0\) (the reminder raises the
vaccination rate). \(\alpha = 0.05\).
Plan: two-sample z test for \(p_1 - p_2\). Random:
treatments were randomly assigned ✓ (random assignment, so the 10%
condition does not apply). Large Counts: \(\hat p_c = 185/500 = 0.37\);
\(250(0.37) = 92.5\) and \(250(0.63) = 157.5\) for each group, all ≥ 10 ✓.
Conclude: Because the p-value of 0.0103 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that
text reminders increase the proportion of patients like these who get a
flu shot. Because treatments were randomly assigned, this is a
cause-and-effect conclusion.
Calculator
STAT → TESTS → 6:2-PropZTest: x1 = 105, n1 = 250, x2 = 80,
n2 = 250, p1: >p2 → z = 2.316, p = 0.0103, and the output shows
\(\hat p\) (the pooled proportion, 0.37). Report \(\hat p_c\) in your Plan
step: it's how graders see you checked the right condition.
Try it
Random samples of 150 students at each of two large high schools found
62 and 48 who bike to school at least once a week. Is there convincing
evidence at \(\alpha = 0.05\) that the proportions differ?
Show answer
\(H_0\!: p_1 = p_2\), \(H_a\!: p_1 \ne p_2\). \(\hat p_c = 110/300 = 0.367\);
\(150(0.367) = 55\) and \(150(0.633) = 95\) ≥ 10 for both groups; samples
are random and under 10% of each school.
\(z = \dfrac{0.413 - 0.320}{\sqrt{0.367(0.633)(1/150 + 1/150)}} = \dfrac{0.093}{0.0556} \approx 1.68\);
two-sided p-value \(\approx 0.0935\). Since \(0.0935 \gt 0.05\), fail to
reject \(H_0\): not convincing evidence that the proportions of students
who bike differ between the two schools.
Unit 6 practice · 10 problems
Unit 6 practice, Inference for Categorical Data: Proportions
Ten problems covering the whole unit, in roughly exam order. Problems 2, 5, and 8 are
full four-step write-ups: State, Plan (with conditions), Do, Conclude in context. Use
1-PropZInt, 1-PropZTest, 2-PropZInt, and 2-PropZTest to check your arithmetic, but
write the formulas with numbers substituted.
Find the critical value \(z^*\) for a 98% confidence interval, and explain why the input to invNorm is 0.99 rather than 0.98.
Show answer
A 98% interval leaves 2% in the two tails combined, 1% in each. The critical value is the z-score with area 0.99 to its left (0.98 in the middle plus 0.01 in the lower tail): \(z^* = \text{invNorm}(0.99, 0, 1) \approx 2.326\). Entering 0.98 would give the value with 2% in the upper tail alone, which corresponds to a 96% interval. For reference: 1.645 (90%), 1.960 (95%), 2.576 (99%).
In a random sample of 300 residents of a city of 60,000 adults, 174 said they favor a proposed park levy. Construct and interpret a 95% confidence interval for the proportion of all adult residents who favor the levy.
Show answer
State: \(p\) = the true proportion of all adult residents of the city who favor the park levy. We want a 95% confidence interval for \(p\).
Plan: one-sample z interval for \(p\). Random: random sample ✓. 10%: \(300 \le 6{,}000\) ✓. Large Counts: 174 successes and 126 failures are both at least 10 ✓.
Do: \(\hat p = 174/300 = 0.58\). \[0.58 \pm 1.960\sqrt{\frac{0.58(0.42)}{300}} = 0.58 \pm 1.960(0.0285) = 0.58 \pm 0.056 \;\Rightarrow\; (0.524,\, 0.636).\] (1-PropZInt with x = 174, n = 300 gives (0.5241, 0.6359).)
Conclude: We are 95% confident that the true proportion of all adult residents of the city who favor the park levy is between 0.524 and 0.636. Because the whole interval is above 0.50, the data give convincing evidence that a majority favor the levy.
Refer to the interval (0.524, 0.636) from the previous problem. Which statement is a correct interpretation of the 95% confidence level?
The population proportion \(p\) is a fixed number, so this particular interval either contains it or it doesn't: there is no 95% probability about it. The 95% belongs to the method over many samples, not to this one interval.
A population proportion is a single number describing the whole city; individual residents don't each “have” a support percentage. This statement misreads a parameter as a property of individuals and makes no sense.
The confidence level describes the long-run capture rate of the method: in repeated sampling, about 95% of the intervals produced this way would contain the true proportion, and about 5% would miss it.
This describes the sampling distribution of \(\hat p\), which is centered at the true \(p\), not at this sample's 0.58. The confidence level is about intervals capturing \(p\), not about other sample proportions landing inside this particular interval.
A state agency wants to estimate the proportion of drivers who use a phone while driving, with a margin of error of at most 2 percentage points at 95% confidence. (a) What sample size is needed if no prior estimate is available? (b) A previous study found about 30% of drivers use a phone. What sample size does that estimate suggest?
Show answer
(a) With \(\hat p = 0.5\) (the conservative choice, since \(\hat p(1-\hat p)\) is largest there): \(n \ge \left(\dfrac{1.960}{0.02}\right)^2(0.5)(0.5) = 2401\). (b) With \(\hat p = 0.3\): \(n \ge \left(\dfrac{1.960}{0.02}\right)^2(0.3)(0.7) = 2016.84\), so \(n = 2017\): always round up. The prior estimate saves nearly 400 drivers. Note the 2-point margin is expensive: halving it to 1 point would quadruple the sample.
A manufacturer claims that only 5% of its LED bulbs fail within the first year. A consumer group suspects the true failure rate is higher and tests an SRS of 400 bulbs from the manufacturer's large output; 30 fail within a year. Is there convincing evidence that the failure rate exceeds 5%? Use \(\alpha = 0.05\).
Show answer
State: \(p\) = the true proportion of the manufacturer's LED bulbs that fail within the first year. \(H_0\!: p = 0.05\) (the claim is accurate); \(H_a\!: p \gt 0.05\) (the failure rate is higher than claimed). \(\alpha = 0.05\).
Plan: one-sample z test for \(p\). Random: SRS ✓. 10%: 400 bulbs is far less than 10% of the manufacturer's output ✓. Large Counts (using \(p_0\)): \(400(0.05) = 20\) and \(400(0.95) = 380\), both at least 10 ✓.
Conclude: Because the p-value of 0.0109 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that more than 5% of the manufacturer's LED bulbs fail within the first year.
For the bulb test in the previous problem: (a) Describe a Type I error and a Type II error in context, and give a consequence of each. (b) Name two changes that would increase the power of the test, and state the cost of each.
Show answer
(a) Type I: the consumer group concludes the failure rate exceeds 5% when it really is 5%, the manufacturer is wrongly accused, perhaps facing a recall or bad publicity it doesn't deserve. Type II: the failure rate really is higher than 5%, but the sample doesn't provide convincing evidence, so the group fails to reject \(H_0\); consumers keep buying bulbs that fail more often than advertised. (b) Testing more bulbs (larger \(n\)) shrinks the standard deviation of \(\hat p\), making a real excess easier to detect; the cost is time and money. Using a larger \(\alpha\), say 0.10, makes rejection easier; the cost is a higher probability of a Type I error. (A larger true failure rate also raises power, but that isn't under the group's control.)
A community college compares pass rates for an online section and an in-person section of the same statistics course. Of 200 randomly sampled students who took the course online, 152 passed; of 240 randomly sampled in-person students, 168 passed. Construct a 95% confidence interval for the difference in pass rates (online minus in-person) and state whether it gives convincing evidence of a difference.
Show answer
\(\hat p_1 = 152/200 = 0.76\), \(\hat p_2 = 168/240 = 0.70\). Conditions: independent random samples; each sample is under 10% of the students who take the course; successes and failures 152, 48, 168, 72 are all at least 10. \[(0.76 - 0.70) \pm 1.960\sqrt{\frac{0.76(0.24)}{200} + \frac{0.70(0.30)}{240}} = 0.06 \pm 1.960(0.0423) = 0.06 \pm 0.083 \;\Rightarrow\; (-0.023,\, 0.143).\] We are 95% confident that the true difference in pass rates (online minus in-person) is between −0.023 and 0.143. Because the interval contains 0, a difference of zero is plausible: there is not convincing evidence that the pass rates differ.
A nursery randomly assigned 300 tomato seedlings to receive a fungicide treatment (150 plants) or no treatment (150 plants). After six weeks, 120 treated seedlings and 102 untreated seedlings had survived. Is there convincing evidence at \(\alpha = 0.05\) that the fungicide increases the survival rate?
Show answer
State: \(p_1\) = the true proportion of seedlings like these that would survive with the fungicide; \(p_2\) = the same without it. \(H_0\!: p_1 - p_2 = 0\) (the fungicide has no effect on survival); \(H_a\!: p_1 - p_2 \gt 0\) (the fungicide increases survival). \(\alpha = 0.05\).
Plan: two-sample z test for \(p_1 - p_2\). Random: treatments were randomly assigned ✓ (so the 10% condition does not apply). Large Counts, using the pooled proportion \(\hat p_c = 222/300 = 0.74\): \(150(0.74) = 111\) and \(150(0.26) = 39\) for each group, all at least 10 ✓.
Conclude: Because the p-value of 0.0089 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the fungicide increases the proportion of seedlings like these that survive six weeks. Because treatments were randomly assigned, this is a cause-and-effect conclusion.
The LED bulb test in problem 5 produced a p-value of 0.0109. (a) Interpret this p-value in context. (b) Suppose instead the sample had produced a p-value of 0.18. Write the conclusion, and explain why "we accept \(H_0\); the failure rate is 5%" would be wrong.
Show answer
(a) Assuming the true failure rate is 0.05, there is a 0.0109 probability of getting a sample proportion of 0.075 or greater in a random sample of 400 bulbs purely by chance. (b) Because 0.18 is greater than \(\alpha = 0.05\), we fail to reject \(H_0\): the data do not provide convincing evidence that more than 5% of the bulbs fail within a year. Failing to reject means the data are consistent with a 5% failure rate: it does not show that the rate is exactly 5%. Many other values (4%, 6%) would also be consistent with the data; a confidence interval would show which. "Accept \(H_0\)" claims more than the test can deliver and loses the conclusion point.
A pollster wants a narrower 95% confidence interval for a population proportion. Which change will make the interval narrower?
Raising the confidence level raises the critical value from \(z^* = 1.960\) to \(2.576\), which widens the interval. More confidence always costs width; you cannot get both narrower and more confident without more data.
Halving \(n\) multiplies the standard error by \(\sqrt{2}\), so the interval gets wider, not narrower. Sample size sits in the denominator of the margin of error, so less data means less precision.
The margin of error is \(z^*\sqrt{\hat p(1-\hat p)/n}\), so quadrupling \(n\) divides it by \(\sqrt{4} = 2\) and the interval is half as wide.
\(\hat p(1-\hat p)\) is largest at 0.5, so substituting 0.5 gives the widest possible interval, which is exactly why 0.5 is the conservative choice when planning a sample size, not a way to tighten an interval.
Lesson 7.1 · Unit 7 · CED topics 7.1–7.2
The t-distribution and confidence intervals for a mean
To build an interval for \(\mu\) you need \(\sigma_{\bar x} = \sigma/\sqrt{n}\),
but in real life you never know \(\sigma\). You substitute the sample
standard deviation \(s\), and that substitution adds uncertainty. The
statistic \((\bar x - \mu)/(s/\sqrt{n})\) no longer follows the standard
normal curve; it follows a t-distribution.
Definition
The t-distribution with \(n - 1\) degrees of
freedom is symmetric and bell-shaped like the normal curve but
with heavier tails; as df grows, it approaches the standard normal. For
df = 9 the 95% critical value is \(t^* = 2.262\) versus \(z^* = 1.960\);
for df = 99 it's 1.984.
Formula and conditions
A one-sample t interval for \(\mu\):
\[\bar x \pm t^* \frac{s}{\sqrt{n}}, \qquad \text{df} = n - 1.\]
Random: random sample. 10%: \(n \le N/10\).
Normal/Large Sample: the population is normal, or
\(n \ge 30\), or (for small \(n\)) a graph of the sample shows no strong
skewness and no outliers.
Worked example · Four-step write-up
A dietitian measures the sodium content (mg) of 10 randomly selected
servings of a brand's tomato soup:
Serving
1
2
3
4
5
6
7
8
9
10
Sodium (mg)
480
495
470
510
488
502
476
491
499
483
State: \(\mu\) = the true mean sodium content per serving
of this soup. 95% confidence interval.
Plan: one-sample t interval for \(\mu\). Random: random
sample of servings ✓. 10%: 10 servings is far under 10% of production ✓.
Normal/Large Sample: \(n = 10 \lt 30\), but a dotplot of the data is
roughly symmetric with no outliers ✓.
Conclude: We are 95% confident that the true mean sodium
content per serving of this soup is between 480.5 mg and 498.3 mg.
Calculator
invT(0.975, 9) = 2.262 gives \(t^*\) (2nd → DISTR).
STAT → TESTS → 8:TInterval; choose Data with the list in L1,
or Stats with \(\bar x\) = 489.4, Sx = 12.46, n = 10; C-Level .95 →
(480.49, 498.31). Use Sx, never σx.
Try it
A random sample of 16 commuters in a large city had a mean one-way
commute of 24.3 minutes with standard deviation 3.8 minutes. A boxplot
shows no outliers. Construct a 90% confidence interval for the mean
commute time.
Show answer
df = 15, \(t^* = \text{invT}(0.95, 15) = 1.753\).
\(24.3 \pm 1.753\,\dfrac{3.8}{\sqrt{16}} = 24.3 \pm 1.67\), so
(22.6, 26.0). We are 90% confident that the true mean one-way commute
for the city's commuters is between 22.6 and 26.0 minutes.
Lesson 7.2 · Unit 7 · CED topic 7.3
Interpreting intervals for a mean and choosing a sample size
The interpretation rules from Unit 6 carry over word for word: the
parameter is just \(\mu\) instead of \(p\). What's new is thinking about
what controls the width of a t interval, and how to plan a sample size
before collecting data.
Interpretations and width
Interval: "We are 95% confident that the true mean
[variable, in context] is between \(a\) and \(b\)." Level:
"In repeated sampling, about 95% of intervals built this way would capture
the true mean." The interval is about \(\mu\), never about individual
values or about \(\bar x\).
The margin of error is \(t^* s/\sqrt{n}\). Raising the confidence level
raises \(t^*\) and widens the interval; raising \(n\) shrinks
\(s/\sqrt{n}\) (and lowers \(t^*\) slightly), narrowing it.
Worked example · Using the soup interval
The label on the soup in Lesson 7.1 says 500 mg of sodium per serving.
The 95% interval was (480.5, 498.3). Because 500 is not in the interval,
the data give convincing evidence at the 5% level that the true mean
differs from 500 mg: in fact, that it's lower. And "95% of servings
contain between 480.5 and 498.3 mg" is wrong: individual
servings vary far more than that (the sample itself ranged 470 to 510).
Worked example · What changes the width
Suppose \(s = 12.8\). Compare margins of error:
n
Level
t*
Margin of error
10
95%
2.262
9.16
10
99%
3.250
13.15
40
95%
2.023
4.09
Quadrupling \(n\) cut the margin by more than half: the extra comes from
the smaller \(t^*\) at df = 39.
Formula · Sample size
To reach a target margin of error, you don't know \(n\) (so you don't
know df); use \(z^*\) and an estimate of \(\sigma\) from a pilot study:
\[n \ge \left(\frac{z^* \sigma}{ME}\right)^2, \quad\text{rounded up.}\]
A hospital wants to estimate mean ER wait time to within ±2 minutes with
95% confidence; past data suggest \(\sigma \approx 12\) minutes.
\(n \ge (1.960 \cdot 12 / 2)^2 = 138.3\), so sample 139 patients.
Try it
A researcher wants a 99% confidence interval for the mean daily screen
time of teenagers with a margin of error of at most 3 minutes. A pilot
study gave \(s \approx 15\) minutes. How many teenagers are needed?
Show answer
\(n \ge \left(\dfrac{2.576 \cdot 15}{3}\right)^2 = 165.9\), so at least
166 teenagers.
Lesson 7.3 · Unit 7 · CED topics 7.4–7.6
Significance tests for a population mean
Same logic as the test for a proportion: assume the claimed value of
\(\mu\), measure how many standard errors away the sample mean landed, and
convert that to a p-value. The only changes are the t statistic and its
degrees of freedom.
Hypotheses, statistic, and conditions
\(H_0\!: \mu = \mu_0\) versus \(H_a\!: \mu \lt \mu_0\), \(\mu \gt \mu_0\),
or \(\mu \ne \mu_0\).
\[t = \frac{\bar x - \mu_0}{s/\sqrt{n}}, \qquad \text{df} = n - 1.\]
The p-value is a tail area under the t curve with \(n - 1\) df; for a
two-sided test, double the one-tail area. Conditions: Random, 10%, and
Normal/Large Sample, the same as for the interval. Decide one-sided or
two-sided from the question before looking at the data.
Worked example · Four-step write-up
A coffee shop advertises 12-ounce lattes. A customer suspects the cups are
underfilled and measures a random sample of 20 lattes over several weeks:
\(\bar x = 11.82\) oz, \(s = 0.41\) oz, with no outliers or strong skew in
the sample. Test at \(\alpha = 0.05\).
State: \(\mu\) = the true mean volume of lattes served by
this shop. \(H_0\!: \mu = 12\) (lattes average 12 oz as advertised);
\(H_a\!: \mu \lt 12\) (lattes are underfilled on average). \(\alpha = 0.05\).
Plan: one-sample t test for \(\mu\). Random: random sample
of lattes ✓. 10%: 20 is far below 10% of all lattes the shop serves ✓.
Normal/Large Sample: \(n = 20 \lt 30\), but the sample shows no outliers
or strong skewness ✓.
Conclude: Because the p-value of 0.0322 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that
the true mean volume of this shop's lattes is less than 12 ounces.
Worked example · Two-sided
Brake pads are supposed to be 5.00 mm thick. A random sample of 15 pads
has \(\bar x = 5.06\) mm and \(s = 0.12\) mm. With \(H_a\!: \mu \ne 5.00\),
\(t = \dfrac{5.06 - 5.00}{0.12/\sqrt{15}} \approx 1.94\), df = 14, and the
two-sided p-value is \(2P(t_{14} \ge 1.94) \approx 0.0733\). Since
\(0.0733 \gt 0.05\), we fail to reject \(H_0\): not convincing evidence
that the mean thickness differs from 5.00 mm.
Calculator
STAT → TESTS → 2:T-Test, Stats: μ₀ = 12, \(\bar x\) = 11.82,
Sx = 0.41, n = 20, μ: <μ₀ → t = −1.963, p = 0.0322. By hand:
tcdf(−1E99, −1.963, 19). Whichever route you take, write the formula
with numbers, the df, and the p-value.
Try it
A seed company says its tomato plants reach a mean height of 50 cm after
six weeks. A gardener using a new soil mix measures 25 randomly chosen
plants: \(\bar x = 53.1\) cm, \(s = 8.5\) cm, no outliers. Is there
convincing evidence at \(\alpha = 0.05\) that her plants grow taller than
50 cm on average?
Show answer
\(H_0\!: \mu = 50\), \(H_a\!: \mu \gt 50\). Conditions: random sample,
10% clearly met, no outliers with \(n = 25\).
\(t = \dfrac{53.1 - 50}{8.5/\sqrt{25}} = \dfrac{3.1}{1.7} \approx 1.82\),
df = 24, p-value \(\approx 0.0404\). Since \(0.0404 \lt 0.05\), reject
\(H_0\): convincing evidence that the mean height with the new soil mix
exceeds 50 cm.
Lesson 7.4 · Unit 7 · CED topic 7.7
Paired data: the matched-pairs t procedure
Before-and-after measurements on the same people are not two independent
samples: each "after" belongs to a specific "before." The right move is
to subtract within each pair and analyze the single list of differences
with the one-sample t procedures you already know.
Method
Data are paired when each observation in one group is
naturally matched to one in the other: the same subject measured twice,
twins, or units matched by a blocking variable. Compute
\(d = \) (second − first) for every pair and treat the differences as one
sample: the parameter is \(\mu_d\), the test statistic is
\[t = \frac{\bar d - 0}{s_d/\sqrt{n}}, \qquad \text{df} = n - 1,\]
and the conditions (Random, 10%, Normal/Large Sample) are checked on the
differences. Two separate, unrelated groups call for the
two-sample procedures in Lessons 7.5–7.6 instead.
Worked example · Four-step write-up
Ten randomly selected students took a typing test (words per minute)
before and after a two-week keyboarding course.
Student
1
2
3
4
5
6
7
8
9
10
Before
38
42
35
50
44
39
47
41
36
45
After
43
41
40
52
49
37
50
46
35
48
d = After − Before
5
−1
5
2
5
−2
3
5
−1
3
State: \(\mu_d\) = the true mean change (after − before)
in typing speed for students like these who take the course.
\(H_0\!: \mu_d = 0\) (no change on average); \(H_a\!: \mu_d \gt 0\) (speed
increases on average). \(\alpha = 0.05\).
Plan: paired t test for \(\mu_d\). Random: students were
randomly selected ✓. 10%: 10 is under 10% of the school's students ✓.
Normal/Large Sample: \(n = 10\), but a dotplot of the ten differences
shows no outliers or strong skewness ✓.
Conclude: Because the p-value of 0.0119 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that
students' typing speeds increase, on average, after the course. (With no
control group, practice or familiarity with the test could also explain
the gain: the course can't be credited alone.)
Calculator
Enter Before in L1 and After in L2, then define L3 = L2 − L1. Run
T-Test with Data, List L3, μ₀ = 0, μ: >μ₀ → t = 2.714,
p = 0.0119. TInterval on L3 gives the 95% interval for
\(\mu_d\): (0.40, 4.40) wpm.
Try it
Fifteen randomly selected adults had their resting heart rate measured
before and after a month of daily meditation. The differences
(before − after) had mean 2.4 bpm and standard deviation 4.5 bpm, with no
outliers. Explain why this is a paired design, then test at
\(\alpha = 0.05\) whether mean heart rate changed.
Show answer
Each adult supplies both measurements, so the pairs are linked and the
differences are the data. \(H_0\!: \mu_d = 0\), \(H_a\!: \mu_d \ne 0\).
\(t = \dfrac{2.4}{4.5/\sqrt{15}} \approx 2.07\), df = 14, two-sided
p-value \(\approx 0.0579\). Since \(0.0579 \gt 0.05\), fail to reject
\(H_0\): not quite convincing evidence that mean resting heart rate
changed.
Lesson 7.5 · Unit 7 · CED topics 7.7–7.8
Confidence intervals for a difference of two means
When two groups are genuinely independent (different people, different
batteries, no pairing) you estimate \(\mu_1 - \mu_2\) directly. The
standard error comes from Lesson 5.6 with the sample standard deviations
substituted, and there's a wrinkle about degrees of freedom.
Formula and conditions
A two-sample t interval for \(\mu_1 - \mu_2\):
\[(\bar x_1 - \bar x_2) \pm t^* \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}.\]
Degrees of freedom: the calculator uses a complicated
formula that gives a non-integer df (here 24.98). By hand, the
conservative choice is the smaller of \(n_1 - 1\) and
\(n_2 - 1\); it gives a slightly wider interval. Either is accepted if you
say which you used.
Conditions: independent random samples (or random assignment); 10% for
each sample; Normal/Large Sample for each group.
Worked example · Four-step write-up
Independent random samples of AA batteries: 12 of Brand A lasted a mean of
8.4 hours (\(s = 0.9\)); 15 of Brand B lasted a mean of 7.8 hours
(\(s = 1.1\)). Neither sample shows outliers or strong skew. Construct a
95% confidence interval for the difference in mean lifetimes.
State: \(\mu_A - \mu_B\), the difference in true mean
lifetimes of Brand A and Brand B batteries. 95% confidence.
Plan: two-sample t interval for \(\mu_A - \mu_B\). Random:
independent random samples ✓. 10%: both samples are tiny fractions of
production ✓. Normal/Large Sample: both \(n \lt 30\), but neither sample
shows outliers or strong skewness ✓.
Conclude: We are 95% confident that the true difference in
mean lifetimes (Brand A minus Brand B) is between −0.25 and 1.45 hours.
Because the interval contains 0, there is not convincing evidence that
the brands' mean lifetimes differ.
Calculator
STAT → TESTS → 0:2-SampTInt, Stats: \(\bar x_1\) = 8.4,
Sx1 = 0.9, n1 = 12, \(\bar x_2\) = 7.8, Sx2 = 1.1, n2 = 15, C-Level .95,
Pooled: No → (−0.193, 1.393), df = 24.98. Always choose
Pooled: No; pooling assumes equal population variances, which the AP
course never asks you to assume.
Try it
Random samples of students from two large high schools reported daily
screen time (minutes): School 1, \(n = 30\), \(\bar x = 72.5\), \(s = 8.2\);
School 2, \(n = 35\), \(\bar x = 68.1\), \(s = 9.4\). Construct a 90%
confidence interval for \(\mu_1 - \mu_2\) using the conservative df.
Show answer
Both \(n \ge 30\), so Normal/Large Sample is met. Conservative df = 29,
\(t^* = 1.699\). \(SE = \sqrt{8.2^2/30 + 9.4^2/35} \approx 2.183\);
\(4.4 \pm 1.699(2.183) = 4.4 \pm 3.71\), so (0.69, 8.11) minutes. We are
90% confident that School 1's mean daily screen time exceeds School 2's
by between 0.69 and 8.11 minutes.
Lesson 7.6 · Unit 7 · CED topics 7.9–7.10
Significance tests for a difference of two means
The two-sample t test is the workhorse of experimental comparisons: did the
treatment group's mean differ from the control group's by more than chance
would produce? Unlike the two-proportion test, there is no pooling: each
group keeps its own standard deviation.
Hypotheses and statistic
\(H_0\!: \mu_1 - \mu_2 = 0\); \(H_a\!: \mu_1 - \mu_2 \gt 0\), \(\lt 0\), or \(\ne 0\).
\[t = \frac{(\bar x_1 - \bar x_2) - 0}{\sqrt{\dfrac{s_1^2}{n_1} + \dfrac{s_2^2}{n_2}}},\]
with df from the calculator or the conservative smaller \(n - 1\).
Conditions are the same as for the two-sample interval, checked for both
groups.
Worked example · Four-step write-up
Does background music hurt reading comprehension? Thirty-six volunteers
were randomly assigned to read a passage with music (18) or in silence
(18) and then took a 40-point quiz. Music: \(\bar x_1 = 27.4\), \(s_1 = 5.1\);
silence: \(\bar x_2 = 30.9\), \(s_2 = 4.6\). Dotplots of both groups show no
outliers or strong skew. Test at \(\alpha = 0.05\).
State: \(\mu_1 - \mu_2\), where \(\mu_1\) and \(\mu_2\) are
the true mean quiz scores for volunteers like these reading with music
and in silence. \(H_0\!: \mu_1 - \mu_2 = 0\) (music makes no difference);
\(H_a\!: \mu_1 - \mu_2 \ne 0\) (music changes mean comprehension).
\(\alpha = 0.05\).
Plan: two-sample t test for \(\mu_1 - \mu_2\). Random:
treatments randomly assigned ✓ (so no 10% condition). Normal/Large
Sample: both \(n = 18 \lt 30\), but neither group shows outliers or strong
skewness ✓.
Conclude: Because the p-value of 0.0378 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that
background music changes the mean quiz score of volunteers like these:
and since the music group scored lower, that it hurts comprehension.
Random assignment permits this cause-and-effect conclusion.
Calculator
STAT → TESTS → 4:2-SampTTest, Stats, μ1: ≠μ2, Pooled: No.
The output reads t = −2.162, p = 0.0378, df = 33.64, then \(\bar x_1\),
\(\bar x_2\), Sx1, Sx2, n1, n2. Copy t, p, and df onto your paper; the df
line is what tells a grader you didn't pool.
Try it
Random samples of 40 students who eat breakfast and 45 who don't took a
20-point memory test: breakfast \(\bar x = 15.2\), \(s = 4.1\); no
breakfast \(\bar x = 13.6\), \(s = 3.8\). Is there convincing evidence at
\(\alpha = 0.05\) that breakfast eaters score higher on average? Can you
conclude breakfast causes higher scores?
Show answer
\(H_0\!: \mu_1 - \mu_2 = 0\), \(H_a\!: \mu_1 - \mu_2 \gt 0\); both
\(n \ge 30\). \(t = \dfrac{15.2 - 13.6}{\sqrt{4.1^2/40 + 3.8^2/45}} = \dfrac{1.6}{0.861} \approx 1.86\);
p-value \(\approx 0.0334\) (calculator df = 79.97) or 0.0353
(conservative df = 39). Either way \(p \lt 0.05\): reject \(H_0\);
convincing evidence that breakfast eaters score higher on average. This
is an observational study, so no causal conclusion: students who eat
breakfast may differ in sleep, income, or other confounders.
Unit 7 practice · 10 problems
Unit 7 practice, Inference for Quantitative Data: Means
Ten problems covering the whole unit, in roughly exam order. Problems 2, 5, 6, and 9
are full four-step write-ups. Use invT, TInterval, T-Test, 2-SampTInt, and
2-SampTTest (Pooled: No) to check the arithmetic; always report the degrees of freedom.
(a) Find the critical value \(t^*\) for a 95% confidence interval for a mean based on a sample of size 12. (b) Explain why a t critical value is used instead of \(z^* = 1.960\). (c) What happens to \(t^*\) as the sample size grows?
Show answer
(a) df \(= 12 - 1 = 11\); \(t^* = \text{invT}(0.975, 11) \approx 2.201\). (b) The population standard deviation \(\sigma\) is unknown and is estimated by the sample standard deviation \(s\). That substitution adds variability to the standardized statistic \((\bar x - \mu)/(s/\sqrt{n})\), which therefore follows a t-distribution with heavier tails than the standard normal, so a larger critical value is needed to reach 95% confidence. (c) As \(n\) (and df) increases, the t-distribution approaches the standard normal and \(t^*\) decreases toward 1.960, at df = 99 it is already 1.984.
A consumer lab measured the caffeine content (mg) of 8 randomly selected 12-ounce cups of coffee from a national chain: 95, 102, 88, 110, 97, 105, 92, 99. Construct and interpret a 95% confidence interval for the mean caffeine content of the chain's 12-ounce coffees.
Show answer
State: \(\mu\) = the true mean caffeine content (mg) of the chain's 12-ounce cups of coffee. 95% confidence interval.
Plan: one-sample t interval for \(\mu\). Random: random sample of cups ✓. 10%: 8 cups is far under 10% of all cups the chain serves ✓. Normal/Large Sample: \(n = 8 \lt 30\), but a dotplot of the data (88, 92, 95, 97, 99, 102, 105, 110) is roughly symmetric with no outliers ✓.
Conclude: We are 95% confident that the true mean caffeine content of the chain's 12-ounce coffees is between 92.6 mg and 104.4 mg.
The chain in the previous problem advertises 100 mg of caffeine per 12-ounce cup. (a) Does the interval (92.6, 104.4) give convincing evidence that the true mean differs from 100 mg? Explain. (b) A student writes, "95% of the chain's cups contain between 92.6 and 104.4 mg of caffeine." Explain why this is incorrect.
Show answer
(a) No. 100 mg is inside the interval, so it is a plausible value for the true mean; at the 5% significance level the data are consistent with the advertised amount. (b) The interval estimates the mean caffeine content, not the contents of individual cups. Individual cups vary far more than the sample mean does, the sample itself ranged from 88 to 110 mg, so far fewer than 95% of cups would fall in such a narrow range. Confidence intervals are statements about parameters, never about individual observations or about \(\bar x\).
A hospital wants to estimate the mean time patients spend in its emergency-room waiting area to within ±1.5 minutes with 95% confidence. A pilot study suggests \(\sigma \approx 8\) minutes. How many patients should be sampled?
Show answer
Since \(n\) (and therefore df) is unknown in advance, use \(z^*\): \[n \ge \left(\frac{z^*\sigma}{ME}\right)^2 = \left(\frac{1.960 \times 8}{1.5}\right)^2 = 109.3,\] so sample at least 110 patients. Round up: 109 patients would leave the margin slightly above 1.5 minutes.
A pharmacy claims its average prescription wait time is 10 minutes. A reporter times a random sample of 35 customers and finds \(\bar x = 11.2\) minutes with \(s = 3.6\) minutes. Is there convincing evidence at \(\alpha = 0.05\) that the true mean wait time exceeds 10 minutes?
Show answer
State: \(\mu\) = the true mean prescription wait time at this pharmacy. \(H_0\!: \mu = 10\) (the claim is accurate); \(H_a\!: \mu \gt 10\) (waits are longer than claimed on average). \(\alpha = 0.05\).
Plan: one-sample t test for \(\mu\). Random: random sample of customers ✓. 10%: 35 is far under 10% of the pharmacy's customers ✓. Normal/Large Sample: \(n = 35 \ge 30\), so the sampling distribution of \(\bar x\) is approximately normal by the CLT ✓.
Conclude: Because the p-value of 0.0284 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the true mean prescription wait time at this pharmacy is greater than 10 minutes.
Eight randomly selected volunteers completed a reaction-time test (milliseconds) after a normal night's sleep and again after 24 hours without sleep.
Volunteer
1
2
3
4
5
6
7
8
Rested
245
268
232
290
255
240
275
262
Sleep-deprived
262
280
239
301
271
238
296
274
Is there convincing evidence at \(\alpha = 0.05\) that sleep deprivation increases mean reaction time? Explain why a paired procedure is appropriate.
Show answer
Each volunteer supplies both measurements, so the two lists are linked pair by pair; analyze the differences \(d = \) sleep-deprived − rested: 17, 12, 7, 11, 16, −2, 21, 12.
State: \(\mu_d\) = the true mean increase in reaction time (sleep-deprived minus rested) for volunteers like these. \(H_0\!: \mu_d = 0\) (no change on average); \(H_a\!: \mu_d \gt 0\) (reaction time increases). \(\alpha = 0.05\).
Plan: paired t test for \(\mu_d\). Random: volunteers were randomly selected ✓. 10%: 8 is far under 10% of the population of interest ✓. Normal/Large Sample: \(n = 8\), but a dotplot of the eight differences shows no outliers or strong skewness ✓.
Conclude: Because the p-value of 0.0010 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that 24 hours of sleep deprivation increases mean reaction time for volunteers like these.
For each study, state whether a paired t procedure or a two-sample t procedure is appropriate, and explain why. (a) Ten cars are each driven on a fixed course once with regular gasoline and once with premium, in random order, and fuel economy is recorded each time. (b) Random samples of 40 freshmen and 40 seniors report their hours of sleep last night. (c) Twenty pairs of identical twins are recruited; within each pair, one twin is randomly assigned to a memory-training program and the other is not, and both take a memory test afterward.
Show answer
(a) Paired: each car is measured twice, so each premium result is naturally matched to a regular result from the same car; analyze the ten differences. Pairing removes car-to-car variation from the comparison. (b) Two-sample: the freshmen and seniors are separate, independent groups with no natural link between any freshman and any senior. (c) Paired (matched pairs): the twins within a pair are alike, so the difference within each pair isolates the training effect; analyze the twenty differences. The question to ask: is there a natural one-to-one link between the observations in the two groups?
Independent random samples of two brands of rechargeable batteries were tested. Brand 1: \(n_1 = 20\), \(\bar x_1 = 34.2\) hours, \(s_1 = 3.1\) hours. Brand 2: \(n_2 = 25\), \(\bar x_2 = 31.8\) hours, \(s_2 = 3.9\) hours. Neither sample shows outliers or strong skewness. Construct a 95% confidence interval for \(\mu_1 - \mu_2\) using the conservative degrees of freedom, interpret it, and state what it says about the brands.
Show answer
Conditions: independent random samples; each sample is a tiny fraction of production; both \(n \lt 30\) but neither sample shows outliers or strong skew, so Normal/Large Sample is met. Conservative df = \(20 - 1 = 19\), \(t^* = \text{invT}(0.975, 19) = 2.093\). \[(34.2 - 31.8) \pm 2.093\sqrt{\frac{3.1^2}{20} + \frac{3.9^2}{25}} = 2.4 \pm 2.093(1.044) = 2.4 \pm 2.18 \;\Rightarrow\; (0.22,\, 4.58).\] We are 95% confident that the true mean lifetime of Brand 1 exceeds that of Brand 2 by between 0.22 and 4.58 hours. Because the entire interval is above 0, there is convincing evidence that Brand 1 batteries last longer on average. (2-SampTInt with df = 43.0 gives (0.30, 4.50): same conclusion.)
Fifty students in an introductory course were randomly assigned to learn a chapter with a new interactive study technique (24 students) or with the standard textbook approach (26 students), then took the same test. Interactive: \(\bar x_1 = 78.3\), \(s_1 = 8.1\). Standard: \(\bar x_2 = 72.9\), \(s_2 = 9.4\). Both groups' scores are roughly symmetric with no outliers. Is there convincing evidence at \(\alpha = 0.05\) that the mean test scores differ between the two techniques?
Show answer
State: \(\mu_1 - \mu_2\), where \(\mu_1\) and \(\mu_2\) are the true mean test scores for students like these using the interactive and standard techniques. \(H_0\!: \mu_1 - \mu_2 = 0\) (the techniques produce the same mean score); \(H_a\!: \mu_1 - \mu_2 \ne 0\) (the mean scores differ). \(\alpha = 0.05\).
Plan: two-sample t test for \(\mu_1 - \mu_2\). Random: students were randomly assigned to techniques ✓ (so no 10% condition). Normal/Large Sample: both \(n \lt 30\), but both groups are roughly symmetric with no outliers ✓.
Conclude: Because the p-value of 0.034 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the mean test scores differ between the two techniques for students like these: and since the interactive group scored higher, that the interactive technique raised scores. Random assignment permits this cause-and-effect conclusion.
A 2-SampTTest with \(H_a\!: \mu_1 \ne \mu_2\) (Pooled: No) reports t = 2.35, p = 0.0236, df = 41.3. At \(\alpha = 0.01\), which conclusion is correct?
Rejecting requires \(p \le \alpha\), and \(0.0236 \gt 0.01\). This conclusion would be correct at \(\alpha = 0.05\), so a student who picks it has probably compared p to the default 5% rather than the stated 1%.
The p-value 0.0236 is greater than \(\alpha = 0.01\), so we fail to reject \(H_0\): the data do not provide convincing evidence at the 1% level that the means differ. (The non-integer df = 41.3 confirms the unpooled procedure was used.)
A significance test can never prove \(H_0\) true; failing to reject means the evidence against it is insufficient, not that the means are equal. “Accept \(H_0\)” is language graders specifically penalize.
The decision compares the p-value with \(\alpha\), not \(t\) with an arbitrary cutoff like 2. A \(t\) of 2.35 is fairly large, but at this df the two-sided 1% threshold is about 2.70, so it falls short.
Lesson 8.1 · Unit 8 · CED topics 8.1–8.3
Chi-square goodness-of-fit
Unit 6 handled a categorical variable with two outcomes. When there are
three or more categories (flavors, colors, days of the week) a single
proportion won't do. The chi-square goodness-of-fit test compares the whole
set of observed counts with the counts a claimed distribution predicts.
Hypotheses, statistic, and conditions
\(H_0\): the population distribution of the variable matches the claim
(\(p_1 = 0.40\), \(p_2 = 0.30\), …). \(H_a\): at least one proportion
differs from the claim. Each expected count is
\(E_i = n p_i\), and
\[\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}, \qquad \text{df} = \text{(number of categories)} - 1.\]
The p-value is the area to the right of \(\chi^2\) (χ²cdf).
Conditions: Random; 10%; Large
Counts: every expected count is at least 5.
Worked example · Four-step write-up
A snack company says its trail mix is 40% peanuts, 30% raisins, 20%
almonds, and 10% chocolate by count. A random sample of 200 pieces from
a large shipment contains 68 peanuts, 66 raisins, 44 almonds, and 22
chocolates. Is there convincing evidence the mix differs from the claim?
Use \(\alpha = 0.05\).
State: \(H_0\): the distribution of pieces in the
shipment is 40% peanuts, 30% raisins, 20% almonds, 10% chocolate.
\(H_a\): at least one of these proportions is different. \(\alpha = 0.05\).
Plan: chi-square goodness-of-fit test. Random: random
sample ✓. 10%: 200 pieces is under 10% of a large shipment ✓. Large
Counts: expected counts 80, 60, 40, 20 are all at least 5 ✓.
Conclude: Because the p-value of 0.392 is greater than
\(\alpha = 0.05\), we fail to reject \(H_0\). There is not convincing
evidence that the distribution of pieces in the shipment differs from the
company's claim.
Worked example · Equal proportions
A die rolled 60 times shows the faces 1–6 a total of 6, 14, 9, 11, 8, and
12 times. If the die is fair, each expected count is 10, so
\(\chi^2 = \frac{16 + 16 + 1 + 1 + 4 + 4}{10} = 4.2\) with df = 5 and
p-value \(\approx 0.521\). No evidence against fairness.
Calculator
Observed counts in L1, expected counts in L2, then STAT → TESTS →
D:χ²GOF-Test, df = 3 → χ² = 3.0, p = 0.3916; the CNTRB
list shows each term's contribution. By hand: χ²cdf(3.0, 1E99, 3).
Try it
A bakery claims its four muffin flavors sell equally well. A random
sample of 120 sales shows 38 blueberry, 22 bran, 33 chocolate, and 27
lemon. Test the claim at \(\alpha = 0.05\).
Show answer
\(H_0\!: p_i = 0.25\) for all four flavors; \(H_a\): at least one differs.
Expected counts all 30 (≥ 5). \(\chi^2 = \frac{64 + 64 + 9 + 9}{30} \approx 4.87\),
df = 3, p-value \(\approx 0.182\). Since \(0.182 \gt 0.05\), fail to
reject \(H_0\): not convincing evidence that the flavors sell unequally.
Lesson 8.2 · Unit 8 · CED topics 8.4–8.5, 8.7
Chi-square test for homogeneity
Now compare several populations on one categorical variable: do students at
three schools prefer the same lunch options? The data form a two-way
table, and the expected counts come from the table itself rather than from
a claimed distribution.
Hypotheses, expected counts, and conditions
Design: separate random samples from two or more
populations (or groups in a randomized experiment), one categorical
response. \(H_0\): the distribution of the variable is the same in every
population. \(H_a\): the distributions are not all the same.
\[E = \frac{(\text{row total})(\text{column total})}{\text{grand total}},
\qquad \chi^2 = \sum \frac{(O - E)^2}{E}, \qquad \text{df} = (r - 1)(c - 1).\]
Conditions: Random (each sample); 10% (each sample, when sampling);
Large Counts: all expected counts at least 5.
Worked example · Four-step write-up
Independent random samples of students at three large high schools were
asked their preferred cafeteria lunch.
School
Pizza
Salad
Sandwich
Total
A
50
20
30
100
B
48
42
30
120
C
30
18
32
80
Total
128
80
92
300
State: \(H_0\): the distribution of lunch preference is
the same at all three schools. \(H_a\): the distributions are not all the
same. \(\alpha = 0.05\).
Plan: chi-square test for homogeneity. Random: three
independent random samples ✓. 10%: each sample is under 10% of its
school ✓. Large Counts: the expected counts below are all at least 5 ✓
(smallest is 21.33).
Conclude: Because the p-value of 0.0287 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the
distribution of lunch preference differs among students at the three
schools.
Calculator
Enter the 3×3 observed counts (no totals) in matrix [A] (2nd → MATRIX →
EDIT). STAT → TESTS → C:χ²-Test, Observed: [A],
Expected: [B] → χ² = 10.817, p = 0.0287, df = 4. Then view [B] to read
and report the expected counts: graders need to see them.
Try it
In a randomized experiment, 60 patients received Treatment A and 60
received Treatment B. Under A, 45 improved and 15 did not; under B, 32
improved and 28 did not. Is there convincing evidence at \(\alpha = 0.05\)
that the improvement distribution differs between treatments?
Show answer
Expected counts: improved \(= (60)(77)/120 = 38.5\), not improved
\(= 21.5\) for each treatment (all ≥ 5); random assignment ✓.
\(\chi^2 = 2\left[\frac{6.5^2}{38.5} + \frac{6.5^2}{21.5}\right] \approx 6.13\),
df = 1, p-value \(\approx 0.0133\). Reject \(H_0\): convincing evidence
that the distribution of improvement differs between Treatments A and B.
(With one df, this is equivalent to a two-sided two-proportion z test.)
Lesson 8.3 · Unit 8 · CED topics 8.6–8.7
Chi-square test for independence
Same table, same arithmetic, different question. When a single
random sample is classified by two categorical variables, you're not
comparing populations: you're asking whether the two variables are
associated in the one population you sampled.
Hypotheses and how the wording differs
Design: one random sample, each individual classified by
two categorical variables. \(H_0\): there is no association between
[variable 1] and [variable 2] in the population (they are independent).
\(H_a\): there is an association. Expected counts, \(\chi^2\),
df = \((r-1)(c-1)\), and conditions are identical to the homogeneity
test. Only the hypotheses and the conclusion sentence change.
Worked example · Four-step write-up
A random sample of 250 adults in a city was asked their age group and
primary source of news.
Age
TV
Online
Print
Total
18–34
20
55
5
80
35–54
35
45
10
90
55+
45
20
15
80
Total
100
120
30
250
State: \(H_0\): there is no association between age group
and primary news source among adults in this city. \(H_a\): there is an
association. \(\alpha = 0.05\).
Plan: chi-square test for independence. Random: random
sample ✓. 10%: 250 is under 10% of the city's adults ✓. Large Counts:
expected counts (below) are all at least 5; the smallest is
\((80)(30)/250 = 9.6\) ✓.
Conclude: Because the p-value is far less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence of an
association between age group and primary news source among adults in
this city.
Compare the two conclusion sentences. Homogeneity: "the distribution of
lunch preference differs among the three schools." Independence:
"there is an association between age group and news source." Using
the wrong one signals that you misread the design: a common deduction.
Calculator
Identical to Lesson 8.2: observed counts in [A], χ²-Test
→ χ² = 31.18, p = 2.8E−6, df = 4, expected counts in [B].
Try it
A random sample of 200 students at a large university was classified by
whether they own a car and whether they hold a part-time job: 48 own a
car and work, 22 own a car and don't work, 52 don't own a car and work,
78 neither. Test for an association at \(\alpha = 0.01\).
Show answer
\(H_0\): no association between car ownership and having a job among the
university's students; \(H_a\): there is an association. Row totals 70
and 130; column totals 100 and 100; expected counts 35, 35, 65, 65
(all ≥ 5). \(\chi^2 = \frac{13^2}{35} + \frac{13^2}{35} + \frac{13^2}{65} + \frac{13^2}{65} \approx 14.86\),
df = 1, p-value \(\approx 0.0001\). Reject \(H_0\): convincing evidence
of an association; car owners in this sample are more likely to work.
Lesson 8.4 · Unit 8 · CED topic 8.7
Choosing the right chi-square test and following up
The three chi-square tests share one formula, so the exam tests whether you
can tell them apart by design, and whether you can say something
useful after rejecting \(H_0\). "The distributions differ" is a start; a
grader wants to know how.
Method · Which test?
Goodness-of-fit: one sample, one categorical variable,
compared with a claimed distribution. df = categories − 1.
Homogeneity: two or more samples (or treatment groups),
one categorical variable. Ask: were the row totals fixed by the
researcher? df = \((r-1)(c-1)\).
Independence: one sample, two categorical variables
measured on each individual. Only the grand total was fixed.
df = \((r-1)(c-1)\).
Follow-up: a significant \(\chi^2\) says a difference
exists, not where. Find the cells with the largest components
\((O-E)^2/E\) and describe them in context, saying whether the observed
count was above or below expected.
Worked example · Classifying designs
(a) A geneticist counts 400 pea plants and compares the four phenotype
counts with the 9:3:3:1 ratio Mendel predicts. Goodness-of-fit,
df = 3.
(b) Random samples of 150 voters in each of four counties are asked which
of three issues matters most. Homogeneity (four populations,
sample sizes fixed), df = \(3 \times 2 = 6\).
(c) One random sample of 500 shoppers records gender and preferred payment
method (cash, card, app). Independence, df = \(1 \times 2 = 2\).
Worked example · Follow-up on the lunch data
The homogeneity test in Lesson 8.2 gave \(\chi^2 = 10.82\), p = 0.029. The
components:
(O − E)²/E
Pizza
Salad
Sandwich
A
1.26
1.67
0.01
B
0.20
3.13
1.26
C
0.50
0.52
2.27
The largest contributor is School B's salad cell: 42 students chose salad
versus 32.0 expected, so School B students chose salad far more often
than the other schools (35% vs. 20% at A and 22.5% at C). The next is
School C's sandwich cell (32 observed, 24.5 expected): sandwiches are
unusually popular at C. School A's shortfall in salad (20 vs. 26.7) also
contributes. Together these cells account for 7.07 of the 10.82.
Exam tip: a chi-square test never proves causation and never tells you the
strength of an association. For that, compare the conditional distributions
(row percentages), as in Unit 2.
Try it
Name the appropriate chi-square test and its degrees of freedom:
(a) A random sample of 300 employees is classified by department (5) and
commute mode (3). (b) A hospital checks whether births are equally likely
on each day of the week using 700 randomly selected birth records.
(c) Random samples of 200 teens and 200 adults are asked which of four
social platforms they use most.
Show answer
(a) Independence (one sample, two variables) df = \((5-1)(3-1) = 8\).
(b) Goodness-of-fit against a uniform distribution, df = \(7 - 1 = 6\).
(c) Homogeneity (two separate samples, one variable) df = \((2-1)(4-1) = 3\).
Unit 8 practice · 10 problems
Unit 8 practice, Inference for Categorical Data: Chi-Square
Ten problems covering the whole unit, in roughly exam order. Problems 2, 4, and 5 are
full four-step write-ups. Use χ²GOF-Test and χ²-Test (with matrices) to check the
arithmetic, but show the expected counts and at least the first terms of the sum.
A candy company claims its bags contain 24% blue, 20% orange, 16% green, 14% yellow, 13% red, and 13% brown candies. A random sample of 250 candies will be used to test the claim. (a) Find the expected count for each color and the degrees of freedom. (b) Is the Large Counts condition met? (c) A student objects that an expected count of 32.5 candies is impossible. Respond.
Show answer
(a) Expected counts \(E_i = 250p_i\): blue 60, orange 50, green 40, yellow 35, red 32.5, brown 32.5 (they sum to 250). df \(= 6 - 1 = 5\). (b) Yes: every expected count is at least 5. (c) An expected count is a long-run average, not a prediction for one bag: if the claim is true, samples of 250 would average 32.5 red candies. Expected counts are never rounded to whole numbers.
A video game's publisher states that its loot boxes contain a common item 60% of the time, a rare item 30%, and an epic item 10%. Players who opened a random sample of 300 boxes recorded 165 common, 100 rare, and 35 epic items. Is there convincing evidence at \(\alpha = 0.05\) that the actual distribution differs from the publisher's claim?
Show answer
State: \(H_0\): the distribution of loot-box items is 60% common, 30% rare, 10% epic. \(H_a\): at least one of these proportions is different. \(\alpha = 0.05\).
Plan: chi-square goodness-of-fit test. Random: random sample of boxes ✓. 10%: 300 is under 10% of all boxes opened ✓. Large Counts: expected counts \(300(0.6) = 180\), \(300(0.3) = 90\), \(300(0.1) = 30\) are all at least 5 ✓.
Conclude: Because the p-value of 0.202 is greater than \(\alpha = 0.05\), we fail to reject \(H_0\). There is not convincing evidence that the distribution of loot-box items differs from the publisher's claim.
In a two-way table with 3 rows and 4 columns and a grand total of 400, one cell lies in a row with total 130 and a column with total 90. (a) Find the expected count for that cell under the null hypothesis. (b) Find the degrees of freedom for the chi-square test. (c) Explain in one sentence where the expected-count formula comes from.
Show answer
(a) \(E = \dfrac{(\text{row total})(\text{column total})}{\text{grand total}} = \dfrac{130 \times 90}{400} = 29.25\). (b) df \(= (3 - 1)(4 - 1) = 6\). (c) If the variables were unrelated, the column's share of the whole table, \(90/400\), would apply equally to every row, so that row's 130 individuals would be expected to include \(130 \times 90/400\) in this column.
Independent random samples of 120 urban and 80 rural residents of a state were asked their primary means of getting to work.
Residence
Drive
Public transit
Bike or walk
Total
Urban
60
40
20
120
Rural
60
8
12
80
Total
120
48
32
200
Is there convincing evidence at \(\alpha = 0.05\) that the distribution of commuting method differs between urban and rural residents of the state?
Show answer
State: \(H_0\): the distribution of primary commuting method is the same for urban and rural residents of the state. \(H_a\): the distributions differ. \(\alpha = 0.05\).
Plan: chi-square test for homogeneity (two separate samples, one categorical variable). Random: two independent random samples ✓. 10%: each sample is under 10% of its population ✓. Large Counts: expected counts; urban 72, 28.8, 19.2; rural 48, 19.2, 12.8 (e.g., \(120 \times 120/200 = 72\)): are all at least 5 ✓.
Conclude: Because the p-value of 0.0003 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence that the distribution of primary commuting method differs between urban and rural residents of the state.
A random sample of 300 adults in a county was asked whether they exercise at least three times a week and whether they rate their sleep quality as good or poor.
Exercise 3+ times/week
Good sleep
Poor sleep
Total
Yes
90
50
140
No
60
100
160
Total
150
150
300
Is there convincing evidence at \(\alpha = 0.01\) of an association between regular exercise and sleep quality among adults in this county?
Show answer
State: \(H_0\): there is no association between regular exercise and sleep quality among adults in the county. \(H_a\): there is an association. \(\alpha = 0.01\).
Plan: chi-square test for independence (one sample, two variables). Random: random sample ✓. 10%: 300 is under 10% of the county's adults ✓. Large Counts: expected counts \(140 \times 150/300 = 70\), 70, 80, 80 are all at least 5 ✓.
Conclude: Because the p-value is far less than \(\alpha = 0.01\), we reject \(H_0\). There is convincing evidence of an association between regular exercise and sleep quality among adults in this county: 64% of regular exercisers reported good sleep versus 38% of non-exercisers. (This is an observational study, so exercise cannot be said to cause better sleep.)
Name the appropriate chi-square test and its degrees of freedom for each study. (a) A random sample of 500 registered voters is classified by party (three categories) and opinion on a ballot measure (favor, oppose, undecided). (b) A blood bank compares the blood-type distribution (O, A, B, AB) of 200 randomly selected donors with the known national percentages. (c) Random samples of 100 residents from each of three cities are asked which of four coffee chains they prefer.
Show answer
(a) Test for independence: one sample, two categorical variables measured on each voter; df \(= (3 - 1)(3 - 1) = 4\). (b) Goodness-of-fit: one sample, one variable, compared with a claimed distribution; df \(= 4 - 1 = 3\). (c) Test for homogeneity: three separate samples (row totals fixed by the researcher), one categorical variable; df \(= (3 - 1)(4 - 1) = 6\). Homogeneity and independence share the arithmetic; the design decides the name and the wording of the conclusion.
Refer to the urban/rural commuting test in problem 4 (\(\chi^2 = 15.97\)). Identify the two cells that contribute most to the chi-square statistic and describe, in context and with conditional percentages, what they reveal about how the distributions differ.
Show answer
The components \((O - E)^2/E\) are: urban drive 2.00, urban transit 4.36, urban bike/walk 0.03, rural drive 3.00, rural transit 6.53, rural bike/walk 0.05. The largest is rural/public transit: only 8 rural residents used transit versus 19.2 expected. Next is urban/public transit: 40 observed versus 28.8 expected. In context, 33% of urban residents (40/120) commute by public transit compared with just 10% of rural residents (8/80), while rural residents are much more likely to drive (75% vs. 50%). The bike/walk proportions are similar (17% vs. 15%) and contribute almost nothing. A significant \(\chi^2\) says the distributions differ; the components say how.
A researcher plans a goodness-of-fit test with a random sample of 60 items across five categories whose claimed proportions are 0.50, 0.24, 0.12, 0.08, and 0.06. Check the Large Counts condition, and explain what the researcher could do if it fails.
Show answer
Expected counts are \(60(0.50) = 30\), 14.4, 7.2, 4.8, and \(60(0.06) = 3.6\). The last two are below 5, so the Large Counts condition fails and the chi-square distribution would not approximate the sampling distribution of the statistic well. Options: collect a larger sample (with \(n = 100\), the smallest expected count would be 6), or combine the two smallest categories into one "other" category with expected count \(60(0.14) = 8.4\), which reduces df from 4 to 3. Never drop the data or adjust the observed counts.
A chi-square test is carried out on a two-way table with 3 rows and 4 columns. The degrees of freedom are
12 is the number of cells in the table (\(3 \times 4\)), not the degrees of freedom. Once the row and column totals are fixed, only \((r-1)(c-1)\) of the cells are free to vary.
11 is cells minus one: the goodness-of-fit formula (categories − 1) misapplied to a two-way table. Goodness-of-fit uses \(k - 1\); tests of homogeneity and independence do not.
For a two-way table, df \(= (r - 1)(c - 1) = (3 - 1)(4 - 1) = 6\): fix the marginal totals and only six cell counts can be chosen freely before the rest are determined.
7 is \(3 + 4\), adding the dimensions instead of multiplying the reduced dimensions. Degrees of freedom for a two-way table always come from the product \((r-1)(c-1)\).
The loot-box test in problem 2 gave \(\chi^2 = 3.19\) and a p-value of 0.202. (a) Interpret the p-value in context. (b) Explain why the p-value for a chi-square test is always the area to the right of the statistic, even though the alternative hypothesis isn't "one-sided" in the usual sense.
Show answer
(a) Assuming the publisher's claimed distribution (60/30/10) is correct, there is a 0.202 probability of getting a chi-square statistic of 3.19 or larger from a random sample of 300 boxes purely by chance. That's not surprising, so the data are consistent with the claim. (b) The chi-square statistic measures total discrepancy between observed and expected counts, and squaring makes every deviation, above or below expected, contribute positively. Any departure from \(H_0\) in any direction makes \(\chi^2\) larger, so "at least as extreme as observed" always means "at least this large," a right-tail area. A tiny \(\chi^2\) means the observed counts match expectations unusually well, not evidence against \(H_0\).
Lesson 9.1 · Unit 9 · CED topics 9.1–9.2
Regression output and the sampling distribution of the slope
The least-squares line you fit in Unit 2 was computed from a sample. A
different random sample of cars would give a slightly different line, so
the slope \(b\) is a statistic with a sampling distribution, and that
makes it a candidate for confidence intervals and tests, just like
\(\bar x\) and \(\hat p\).
Definitions
The population regression line is \(\mu_y = \alpha + \beta x\):
for each \(x\), the mean response is a linear function of \(x\). The
sample slope \(b\) is an unbiased estimator of \(\beta\), and its sampling
distribution has standard deviation \(\sigma_b = \dfrac{\sigma}{\sigma_x\sqrt{n}}\),
which shrinks with a larger sample, more spread in \(x\), and less scatter
about the line. Computer output estimates it as SE Coef
(\(SE_b\)), and estimates \(\sigma\) with \(s\), the standard deviation of
the residuals. Inference for \(\beta\) uses \(t\) with df = \(n - 2\).
Conditions · LINER
Linear: the scatterplot is roughly linear and the residual plot shows no curve.
Independent: observations are independent; check the 10% condition when sampling.
Normal: a dotplot or histogram of the residuals shows no strong skew or outliers.
Equal SD: the residual plot's vertical spread is roughly constant across \(x\).
Random: data come from a random sample or randomized experiment.
Worked example · Reading the output
A random sample of 15 sedans was weighed (thousands of pounds) and
tested for highway fuel economy (mpg). Software produced:
Predictor
Coef
SE Coef
T
P
Constant
53.317
1.595
33.42
0.000
Weight
−7.616
0.455
−16.74
0.000
S = 1.131 R-Sq = 95.6% R-Sq(adj) = 95.2%
The equation is \(\widehat{\text{mpg}} = 53.317 - 7.616(\text{weight})\).
The slope \(b = -7.616\) estimates \(\beta\): each additional 1,000 lb is
associated with a predicted drop of about 7.6 mpg. \(SE_b = 0.455\) means
that in repeated samples of 15 sedans, the sample slope would typically
differ from the true slope by about 0.455 mpg per 1,000 lb.
\(s = 1.131\): actual mpg is typically about 1.1 mpg from the predicted
value. \(r^2 = 0.956\), and since the slope is negative,
\(r = -\sqrt{0.956} \approx -0.978\). The T column is Coef ÷ SE Coef
(\(-7.616/0.455 = -16.74\)); df = \(15 - 2 = 13\).
Exam tip: ignore the Constant row's T and P unless a question asks about the
intercept. Inference questions are almost always about the slope.
Try it
Output for 22 randomly selected employees, predicting salary (thousands
of dollars) from years of experience: Constant coef 41.20 (SE 3.10);
Experience coef 4.12 (SE 1.35); S = 8.4; R-Sq = 31.8%. Write the
regression equation, interpret the slope, state \(SE_b\) and the degrees
of freedom, and compute the t statistic for the slope.
Show answer
\(\widehat{\text{salary}} = 41.20 + 4.12(\text{years})\). Each additional
year of experience is associated with a predicted salary increase of
about $4,120. \(SE_b = 1.35\); df = 20;
\(t = 4.12/1.35 \approx 3.05\). (Also \(r = +\sqrt{0.318} \approx 0.56\).)
Lesson 9.2 · Unit 9 · CED topics 9.2–9.3
Confidence intervals for the slope
The interval for \(\beta\) has the familiar shape, estimate ± critical
value × standard error, with the estimate and standard error read straight
off the output. The skill the exam tests is interpreting it: an interval for
a slope is a statement about a rate.
Formula
A t interval for the slope of the population regression line:
\[b \pm t^* \, SE_b, \qquad \text{df} = n - 2.\]
Interpretation: "We are 95% confident that the true slope of the
population regression line relating [y] to [x] is between \(a\) and
\(c\)": that is, for each additional unit of \(x\), the predicted \(y\)
changes by an amount in that range.
Worked example · Four-step write-up
Using the sedan output from Lesson 9.1 (\(n = 15\), \(b = -7.616\),
\(SE_b = 0.455\)), construct and interpret a 95% confidence interval for
the slope. The scatterplot is linear, the residual plot shows no pattern
and roughly constant spread, and a dotplot of residuals has no outliers.
State: \(\beta\) = the true slope of the population
regression line relating highway mpg to weight (thousands of pounds) for
sedans of this type. 95% confidence.
Plan: t interval for the slope. Linear: scatterplot linear,
no curve in the residual plot ✓. Independent: 15 sedans is under 10% of
all sedans ✓. Normal: residual dotplot shows no strong skew or
outliers ✓. Equal SD: residual spread roughly constant ✓. Random: random
sample ✓.
Conclude: We are 95% confident that the true slope of the
population regression line relating highway mpg to weight is between
−8.60 and −6.63 mpg per thousand pounds. In other words, each additional
1,000 lb is associated with a decrease in predicted fuel economy of
between 6.63 and 8.60 mpg.
Because the interval lies entirely below 0, it also gives convincing
evidence at the 5% level that \(\beta \ne 0\): the same conclusion the
test in Lesson 9.3 reaches. A 99% interval (\(t^* = 3.012\)) would be
(−8.99, −6.25); more confidence, wider range.
Calculator
With the data in L1 and L2: STAT → TESTS → G:LinRegTInt,
Xlist L1, Ylist L2, C-Level .95 → (−8.599, −6.633), b = −7.616, df = 13,
s = 1.131. The TI-84 does not print \(SE_b\); if you need it, run
LinRegTTest and compute \(SE_b = b/t\), or use
\(SE_b = \dfrac{s}{s_x\sqrt{n-1}}\). When only output is given (no raw
data), work by hand: the calculator can't help.
Try it
A regression of exam score on weekly study hours for 20 randomly
selected students gave slope 2.35 points per hour with \(SE_b = 0.62\).
Conditions are met. Construct and interpret a 90% confidence interval for
the slope.
Show answer
df = 18, \(t^* = \text{invT}(0.95, 18) = 1.734\).
\(2.35 \pm 1.734(0.62) = 2.35 \pm 1.08\), so (1.27, 3.43). We are 90%
confident that the true slope of the population regression line
relating exam score to weekly study hours is between 1.27 and 3.43
points per hour.
Lesson 9.3 · Unit 9 · CED topics 9.4–9.6
Significance tests for the slope
"Is there a linear relationship at all?" is a question about \(\beta\): if
\(\beta = 0\), the mean response doesn't change with \(x\) and the line is
useless for prediction. The test asks whether the sample slope is far
enough from 0, in standard-error units, to rule out chance.
Hypotheses and statistic
\(H_0\!: \beta = 0\) (no linear relationship between \(x\) and \(y\) in the
population); \(H_a\!: \beta \ne 0\), \(\beta \gt 0\), or \(\beta \lt 0\).
\[t = \frac{b - 0}{SE_b}, \qquad \text{df} = n - 2.\]
Computer output's P column is the two-sided p-value; for a
one-sided alternative whose direction matches \(b\), halve it. Conditions
are LINER, as in Lesson 9.1.
Worked example · Four-step write-up
Twelve randomly selected students at a large school reported hours of
sleep the night before a 20-point quiz.
Sleep (h)
5.0
5.5
6.0
6.0
6.5
7.0
7.0
7.5
8.0
8.0
8.5
9.0
Score
11
14
12
15
13
16
14
15
18
15
17
16
Output: \(\widehat{\text{score}} = 5.814 + 1.265(\text{sleep})\),
\(SE_b = 0.321\), \(s = 1.32\), \(r^2 = 0.609\). Is there convincing
evidence that more sleep is associated with higher scores? \(\alpha = 0.05\).
State: \(\beta\) = the true slope of the population
regression line relating quiz score to hours of sleep for students at
this school. \(H_0\!: \beta = 0\) (no linear relationship);
\(H_a\!: \beta \gt 0\) (scores tend to rise with sleep). \(\alpha = 0.05\).
Plan: t test for the slope. Linear: scatterplot is
roughly linear with no curve in the residual plot ✓. Independent: 12 is
under 10% of the school's students ✓. Normal: residuals (all between
−1.4 and 2.1) show no strong skew or outliers ✓. Equal SD: residual
spread is similar across sleep values ✓. Random: random sample ✓.
Conclude: Because the p-value of 0.0014 is less than
\(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence of a
positive linear relationship between hours of sleep and quiz score for
students at this school.
Calculator
STAT → TESTS → F:LinRegTTest, Xlist L1, Ylist L2,
β & ρ: >0 → t = 3.945, p = 0.0014, df = 10, a = 5.814, b = 1.265,
s = 1.322, r² = 0.609, r = 0.780. Recover \(SE_b = b/t = 0.321\) if a
question asks for it.
Caution · What a significant slope means
It means a linear association exists in the population, not that sleep
causes higher scores (this is an observational study; motivated
students may both sleep more and study more). It says nothing about
strength: with large \(n\), a slope that explains 2% of the variation can
be highly significant. Use \(r^2\) for strength. And failing to reject
\(H_0\) doesn't prove \(\beta = 0\): the relationship may be nonlinear
or the sample too small.
Try it
For 25 randomly selected employees, a regression of job-satisfaction
score on one-way commute time (minutes) gave slope −0.85 with
\(SE_b = 0.31\); conditions are met. Test \(H_0\!: \beta = 0\) against
\(H_a\!: \beta \ne 0\) at \(\alpha = 0.05\).
Show answer
\(t = \dfrac{-0.85}{0.31} \approx -2.74\), df = 23, two-sided p-value
\(\approx 0.0116\). Since \(0.0116 \lt 0.05\), reject \(H_0\): convincing
evidence of a linear relationship between commute time and job
satisfaction; longer commutes are associated with lower satisfaction,
though causation can't be claimed from an observational study.
Unit 9 practice · 10 problems
Unit 9 practice, Inference for Quantitative Data: Slopes
Ten problems covering the whole unit, in roughly exam order. Problems 4 and 6 are full
four-step write-ups. Most of the unit is reading computer output, so the calculator
only helps with invT and tcdf here.
A dealer recorded the age (years) and asking price (thousands of dollars) of a random sample of 20 used cars of one model. Software output:
Predictor
Coef
SE Coef
T
P
Constant
24.83
1.12
22.17
0.000
Age
−1.940
0.210
−9.24
0.000
S = 2.15 R-Sq = 82.5%
Write the equation of the least-squares line, interpret the slope in context, explain what \(SE_b = 0.210\) measures, and state the degrees of freedom for inference about the slope.
Show answer
\(\widehat{\text{price}} = 24.83 - 1.94(\text{age})\). Slope: each additional year of age is associated with a predicted decrease of about $1,940 in asking price. \(SE_b = 0.210\) estimates the standard deviation of the sampling distribution of \(b\): in repeated random samples of 20 cars, the sample slope would typically differ from the true slope \(\beta\) by about 0.21 thousand dollars per year. Degrees of freedom \(= n - 2 = 18\). (Check: \(T = -1.94/0.210 \approx -9.24\), as printed.)
Using the output in problem 1, find the correlation \(r\) and interpret \(r^2\) and \(s\) in context.
Show answer
\(r^2 = 0.825\), and the slope is negative, so \(r = -\sqrt{0.825} \approx -0.91\): a strong negative linear association between age and price. \(r^2\): about 82.5% of the variation in the asking prices of these cars is explained by the linear relationship with age. \(s = 2.15\): actual asking prices are typically about $2,150 away from the price predicted by the regression line.
Before doing inference for the slope, an analyst examines the used-car data. The scatterplot shows a roughly linear pattern. The residual plot shows no curve, but the vertical spread of the residuals clearly increases as age increases. A histogram of the residuals is roughly symmetric with no outliers. The cars were a random sample from a very large inventory. Check each condition for inference about the slope and identify any that are violated.
Show answer
Linear: met: the scatterplot is linear and the residual plot shows no curved pattern. Independent: met: random sample of 20 from a very large inventory, so the 10% condition holds. Normal: met: the residual histogram is roughly symmetric with no outliers. Equal SD:violated: the fan shape (residual spread growing with age) means the scatter about the line is not constant across \(x\). Random: met. With Equal SD violated, the standard error \(SE_b\) and the t-based p-values and intervals are not fully trustworthy; a transformation of the response (such as \(\log(\text{price})\)) often fixes both the fan shape and any curvature.
Assume all conditions are met for the used-car data in problem 1. Construct and interpret a 95% confidence interval for the slope of the population regression line.
Show answer
State: \(\beta\) = the true slope of the population regression line relating asking price (thousands of dollars) to age (years) for used cars of this model. 95% confidence.
Plan: t interval for the slope. Conditions (Linear, Independent, Normal, Equal SD, Random) are given as met; \(n = 20\) is under 10% of the dealer's inventory.
Conclude: We are 95% confident that the true slope of the population regression line relating asking price to age is between −2.38 and −1.50 thousand dollars per year: that is, each additional year of age is associated with a decrease in predicted price of between $1,500 and $2,380.
(a) Does the interval (−2.38, −1.50) from problem 4 give convincing evidence, at the 5% significance level, of a linear relationship between age and price? Explain. (b) A student concludes, "Aging causes a car's price to fall by about $1,940 per year." Comment on the word causes.
Show answer
(a) Yes. The interval does not contain 0, so \(\beta = 0\) is not a plausible value; a two-sided test of \(H_0\!: \beta = 0\) at \(\alpha = 0.05\) would reject. There is convincing evidence of a negative linear relationship between age and asking price. (b) The data are observational, the dealer didn't randomly assign ages to cars, so the regression establishes an association, not causation. Older cars also tend to have more mileage and wear, and those variables are confounded with age. Say "is associated with," not "causes." (It also matters that \(b = -1.94\) is a sample estimate; the population slope could plausibly be anywhere in the interval.)
For a random sample of 16 houses in a county, a regression of monthly electricity cost (dollars) on living area (square feet) produced slope \(b = 0.0412\) with \(SE_b = 0.0135\). The scatterplot is linear, the residual plot shows no pattern and constant spread, and a dotplot of residuals has no outliers. Is there convincing evidence at \(\alpha = 0.05\) of a linear relationship between living area and monthly electricity cost?
Show answer
State: \(\beta\) = the true slope of the population regression line relating monthly electricity cost to living area for houses in this county. \(H_0\!: \beta = 0\) (no linear relationship); \(H_a\!: \beta \ne 0\) (there is a linear relationship). \(\alpha = 0.05\).
Plan: t test for the slope. Linear: scatterplot linear, no curve in the residual plot ✓. Independent: 16 houses is under 10% of the county's houses ✓. Normal: residual dotplot shows no outliers or strong skew ✓. Equal SD: residual spread roughly constant ✓. Random: random sample ✓.
Conclude: Because the p-value of 0.0086 is less than \(\alpha = 0.05\), we reject \(H_0\). There is convincing evidence of a linear relationship between living area and monthly electricity cost for houses in this county; since \(b \gt 0\), larger houses tend to have higher costs.
Regression output for predicting a plant's height from weekly hours of sunlight shows a positive slope with P = 0.036 in the slope row. (a) What p-value should you report for \(H_a\!: \beta \gt 0\)? (b) For \(H_a\!: \beta \lt 0\)? (c) What would you report for \(H_a\!: \beta \ne 0\)? Explain.
Show answer
Software prints the two-sided p-value. (a) The sample slope is positive, matching the direction of \(H_a\), so the one-sided p-value is half of 0.036: \(0.018\). (b) The sample slope points the wrong way for this alternative, so the p-value is \(1 - 0.018 = 0.982\), no evidence at all for a negative slope. (c) Report 0.036 as printed. Decide the alternative from the research question before looking at the output; you cannot switch to one-sided after seeing that it halves the p-value.
Part of a regression output for 22 randomly selected employees, predicting annual sales (thousands of dollars) from years of experience, has been smudged: the Experience row reads Coef 3.42, SE Coef ____, T 2.85. Recover \(SE_b\), state the degrees of freedom, and write the interpretation of the slope.
Show answer
Since \(T = b/SE_b\), \(SE_b = b/T = 3.42/2.85 = 1.20\). df \(= 22 - 2 = 20\). Slope: each additional year of experience is associated with a predicted increase of about $3,420 in annual sales. (The same trick recovers \(SE_b\) from a TI-84 LinRegTTest, which prints \(b\) and \(t\) but not \(SE_b\).)
A test of \(H_0\!: \beta = 0\) versus \(H_a\!: \beta \ne 0\) for a regression based on 400 randomly selected adults gives \(t = 4.1\) and a p-value less than 0.0001, with \(r^2 = 0.04\). Which statement is correct?
The adults were randomly selected, not randomly assigned to values of the explanatory variable, so this is observational data and no causal claim is justified, however small the p-value.
This confuses statistical significance with strength. With \(n = 400\), even a weak relationship produces a tiny p-value; \(r^2 = 0.04\) (so \(|r| = 0.2\)) says the linear relationship is weak.
A tiny p-value says the slope is distinguishable from 0 in the population, the relationship is real, but \(r^2 = 0.04\) means the line explains only 4% of the variation in the response, so it is weak. Significance and strength are separate questions.
This reverses \(r^2\): the line explains 4% of the variation, and 96% is the unexplained fraction, \(1 - r^2\), that remains in the residuals.
For the used-car data in problem 1, the standard deviation of the residuals is \(s = 2.15\) and the standard deviation of the ages is \(s_x = 2.35\) years, with \(n = 20\). (a) Use the formula \(SE_b = \dfrac{s}{s_x\sqrt{n - 1}}\) to verify the printed value of \(SE_b\). (b) Explain how each of \(s\), \(s_x\), and \(n\) affects the precision of the slope estimate.
Show answer
(a) \(SE_b = \dfrac{2.15}{2.35\sqrt{19}} = \dfrac{2.15}{10.24} \approx 0.210\), matching the output. (b) Less scatter about the line (smaller \(s\)) makes the slope more precise. A wider range of \(x\)-values (larger \(s_x\)) also makes it more precise: points spread far apart in \(x\) pin down the tilt of the line better than points bunched together. And a larger \(n\) shrinks \(SE_b\) in proportion to \(1/\sqrt{n - 1}\). To estimate a slope well, sample cars across a wide range of ages, not just 3- and 4-year-olds.
Unit recap
Unit recap
0:00 / 0:00
Animated recap with on-screen narration. Turn on Voice to
have it read aloud (uses your device's built-in voice). Pressing play
counts as your one free video.
Free preview complete
That's the end of the free preview.
You've opened five lessons (or watched a unit video), which is
as much as we can show without a subscription. Everything you've
already opened stays available; use the outline on the left to go
back to it.