MTH1W Data Literacy Grade 9 Unit 6 — practice questions with full solutions
Measures of central tendency, the effect of outliers, sampling methods, recognising bias, scatter plots and lines of best fit. An 18-hour examinable unit that is easy to under-prepare for.
Free MTH1W data literacy practice for Grade 9 students in Ontario. These 11 questions cover the same material as a typical unit test or exam review on this topic, and every one comes with a complete worked solution — not just an answer key. Useful whether you are preparing for a MTH1W unit test, catching up on a lesson, or reviewing before the final exam.
MTH1W · Grade 9Strand D — Data10 core1 challenge
1Core
Find the mean, median and mode of: 12, 15, 11, 15, 18, 20, 15
Show the full solution
Sort the data first: 11, 12, 15, 15, 15, 18, 20.
Mean: add all values (11+12+15+15+15+18+20 = 106) and divide by 7 → 106 ÷ 7 ≈ 15.14.
Median: with 7 values, the middle is the 4th → 15.
Mode: the most frequent value is 15 (appears three times).
Final answermean ≈ 15.14, median = 15, mode = 15
2Core
An extra value of 60 is added to the data set above. Recalculate the mean and median, and state which measure is more resistant to outliers.
Show the full solution
New sum: 106 + 60 = 166, with 8 values → mean = 166 ÷ 8 = 20.75.
Sorted data is now 11, 12, 15, 15, 15, 18, 20, 60 — with 8 values the median is the average of the 4th and 5th.
Median = (15 + 15) ÷ 2 = 15, unchanged.
The mean jumped by more than 5 while the median did not move at all.
Final answermean = 20.75, median = 15; the median is far more resistant to outliers
3Core
Find the range of: 11, 12, 15, 15, 15, 18, 20
Show the full solution
Range = largest value − smallest value.
Largest = 20, smallest = 11.
Subtract.
Final answer9
4Core
Study the scatter plot and its line of best fit. (a) Describe the correlation. (b) Use the line to estimate the score improvement for 4.5 hours of study. (c) A student claims the line proves that studying causes higher scores. Is that a fair claim?
Show the full solution
(a) The points rise from left to right in a fairly tight band, so there is a strong positive correlation.
(b) Read up from 4.5 on the horizontal axis to the dashed line, then across. The line is approximately y = 1.1x + 2.
So at x = 4.5: y ≈ 1.1(4.5) + 2 ≈ 7. Any estimate near 7 is acceptable when reading from a graph.
(c) No. Correlation shows the two variables move together, but it cannot establish cause.
Another factor could explain both — for example, more motivated students may both study more and attend class more.
Final answer(a) Strong positive correlation (b) ≈ 7 (c) No — correlation does not prove causation
5Core
Name the sampling method in each case: (a) surveying every 10th student on an alphabetical list; (b) splitting students by grade, then randomly selecting from each grade; (c) surveying whoever happens to walk past the cafeteria.
Show the full solution
(a) Choosing at a fixed regular interval from an ordered list is systematic sampling.
(b) Dividing the population into groups and sampling randomly within each is stratified sampling.
(c) Sampling whoever is easiest to reach is convenience sampling — the least reliable of the three.
Final answer(a) systematic (b) stratified (c) convenience
6Core
Identify the bias in this survey question and rewrite it fairly: “Don't you agree that our excellent cafeteria should stay open longer?”
Show the full solution
The phrase “Don't you agree” pressures the respondent toward saying yes.
The word “excellent” is a loaded adjective that pre-judges the cafeteria's quality.
This is a leading question — the wording steers the answer.
A fair version removes both the pressure and the loaded language.
Final answerLeading/loaded wording. Fair version: “Should the cafeteria's opening hours be extended?”
7Core
A scatter plot compares hours spent studying with test scores, and the points rise steadily from left to right in a fairly tight band. Describe the correlation in terms of direction and strength.
Show the full solution
Direction is given by the trend: points rising left to right means as one variable increases, so does the other.
That is a positive correlation.
Strength is given by how tightly points cluster around a line — a tight band means a strong relationship.
Important caution: correlation does not prove that studying caused the higher scores.
Final answerA strong positive correlation (but correlation alone does not establish causation)
8Core
Use the histogram of test scores to find: (a) the total number of students, (b) the modal interval, (c) how many students scored 80 or above, and (d) that group as a percent of the class, to one decimal place.
Show the full solution
(a) Add every bar: 3 + 7 + 12 + 9 + 4 = 35 students.
(b) The modal interval is the tallest bar — the 70–80 interval, with 12 students.
(c) Scores of 80 or above come from the last two bars: 9 + 4 = 13 students.
(d) As a percent: 13 ÷ 35 = 0.3714… × 100 ≈ 37.1%.
Final answer(a) 35 (b) 70–80 (c) 13 (d) ≈ 37.1%
9Core
Explain the difference between interpolation and extrapolation, and state which is generally more reliable.
Show the full solution
Interpolation means estimating a value inside the range of the data you actually collected.
Extrapolation means estimating beyond the range of the collected data.
Interpolation is more reliable because the trend is supported by real observations in that region.
Extrapolation assumes the pattern continues unchanged, which may simply not be true.
Final answerInterpolation estimates within the data range and is more reliable; extrapolation goes beyond it and is riskier
10Core
Distinguish between primary and secondary data, and give one example of each in a study of student sleep habits.
Show the full solution
Primary data is collected first-hand by the person doing the study.
Secondary data is collected by someone else and reused.
Primary example: handing out your own sleep survey to classmates.
Secondary example: using published Statistics Canada figures on teenage sleep.
Final answerPrimary = collected yourself (own survey); Secondary = collected by others (published statistics)
11Challenge
A line of best fit for a scatter plot of study hours (x) versus test score (y) is y = 6x + 52, based on data from 0 to 8 hours. (a) Predict the score for 5 hours. (b) Predict the score for 20 hours and comment on the reliability of that prediction.
Show the full solution
(a) Substitute x = 5: y = 6(5) + 52 = 30 + 52 = 82.
Since 5 is inside the 0–8 range, this is interpolation and is reasonably trustworthy.
(b) Substitute x = 20: y = 6(20) + 52 = 120 + 52 = 172.
A test score of 172 is impossible — and 20 hours lies far outside the data range.
This is extrapolation, and it shows exactly why extrapolating far beyond the data can produce nonsense.
Final answer(a) 82 (b) 172, which is unreliable — it is extrapolation well beyond the data and gives an impossible score
MTH1W Data Literacy — common questions
Short answers to the things students ask most about this unit.
Why is the median better than the mean when there is an outlier?
The mean uses every value, so one extreme number drags it far off. The median only depends on the middle position, so it barely moves.
What is the difference between interpolation and extrapolation?
Interpolation estimates inside the range of your data and is fairly reliable. Extrapolation goes beyond the data and assumes the trend continues, which often fails.
The classes teach the method behind every one of these.
The solutions above show the steps. The classes teach how to think about the problem in the first place — interactive, and worked through at the student’s own pace. Try 3 complete classes free — no payment required to begin.