Chapter 6 — Ratios, Percentages, Data and Probability
This is Problem-Solving and Data Analysis, about 15% of the SAT Math section.
It is the most readable part of the test and the easiest to lose marks on, because the mathematics is simple and the questions are long. Almost everything here is arithmetic you have known for years, wrapped in a paragraph about a survey, a recipe, or a scatterplot. The work is in reading carefully and keeping track of what each number means.
A calculator is permitted on every question, so none of these are testing whether you can divide. They are testing whether you divided the right two numbers.
Two habits pay for themselves throughout this chapter:
- Write the units on every number. "12 dollars per hour" and "12 hours per dollar" are different, and one of them is always an option.
- Ask what the answer should look like before you compute. Bigger or smaller? A percentage or an amount? Roughly what size? A rough expectation catches most wrong answers instantly.
Topics covered: ratio and proportion · unit rates and conversion · percentages, increase and decrease · mean, median and range · standard deviation · scatterplots and lines of best fit · probability and two-way tables · sampling and inference
Section 1 — Ratio and Proportion
A ratio compares quantities. A proportion says two ratios are equal:
which you solve by cross-multiplying: .
The trap is part versus whole. If a class has boys and girls in the ratio , then:
- boys to girls is
- boys to everyone is , because there are parts altogether
Read carefully which one is wanted. Both are always offered.
Setting up a proportion works whenever a relationship stays constant. Keep the same quantity on top of both fractions — miles over hours on the left means miles over hours on the right.
Topic: Part-to-whole from a ratio
In a bag, the ratio of red to blue marbles is . If there are 44 marbles altogether, how many are red?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
The ratio has parts in total. So red marbles are of everything:
Check: red 16, blue , and ✓ And simplifies to ✓
The quick way: 11 parts share 44 marbles, so one part is 4 marbles. Red is 4 parts, so . Finding the value of one part first works on every ratio question and is worth making automatic.
Why each wrong option is wrong:
- A, — the ratio number, not a count of marbles.
- C, — the blue marbles.
- B, — divided the total by 4 instead of by 11.
Takeaway: Add the ratio parts to get the whole, then find the value of one part. Everything else is multiplication after that.
Topic: Setting up a proportion
A biscuit recipe is written for a fixed batch size and uses 3 cups of flour for every 8 biscuits it makes. A baker needs to scale the recipe up for an order of 20 biscuits. How many cups of flour are needed for 20 biscuits?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
Keep the same quantity on top of both fractions — cups over biscuits, both sides:
Cross-multiply:
Check the size before accepting it. 20 biscuits is a bit more than twice 8, so the flour should be a bit more than twice 3 — around 7 or 8. fits ✓
Why each wrong option is wrong:
- A, — set the fractions up upside down on one side.
- B, — the rate of biscuits per cup, which is not what was asked.
- C, — doubled 3 to get 6 and then adjusted wrongly, or stopped at 60 divided by 4.
Takeaway: Same units on top on both sides, then cross-multiply. Estimate the answer's size first — it eliminates the inverted-ratio option immediately.
Topic: Converting units
A tap fills a tank at 250 millilitres per second. How many litres does it deliver in 4 minutes? (1 litre = 1000 millilitres)
There are no options — work the answer out and enter it yourself, in the form the grid accepts.
Show the worked solution
Answer: 60
Explanation
Two conversions, done one at a time.
First, minutes to seconds, because the rate is per second:
Now the volume in millilitres:
Finally millilitres to litres:
Check the direction of each conversion by asking whether the number should get bigger or smaller. Seconds are smaller than minutes, so there are more of them — multiply. Litres are bigger than millilitres, so there are fewer of them — divide. Getting a conversion upside down changes the answer by a factor of 3600 here, so it is never a close call.
This is a grid-in. Enter 60.
Takeaway: Convert one unit at a time and check each direction by asking whether the count should grow or shrink. Smaller unit, bigger number.
Section 2 — Percentages
Every percentage question becomes easy once you turn the percentage into a multiplier:
| The words | The multiplier |
|---|---|
| 20% of | |
| increase by 20% | |
| decrease by 20% | |
| what percent is of |
Three things the test relies on:
- Increase then decrease by the same percentage does not return you to the start. Up 10% then down 10% is — a 1% loss overall.
- Percentage change is always measured against the original, not the new value.
- Working backwards from a final amount means dividing by the multiplier, not subtracting the percentage.
Topic: Percentage increase
A jacket in a shop is priced at $80. At the start of the new season the shop raises the price of every item in that range by 15%. What is the new price of the jacket?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
An increase of 15% means the new price is 115% of the old:
Or find the increase and add it: of is , and ✓
Why each wrong option is wrong:
- A, — the size of the increase, not the new price.
- B, — decreased by 15%.
- C, — added $15. A percentage of 80 is not 15.
Takeaway: "Increase by %" means multiply by . One multiplication, no addition step to forget.
Topic: Working backwards from a percentage
A coat is offered in a sale at 20% off its original price, and after the reduction a customer pays $96 for it. What was the original price?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
The discount means the customer paid 80% of the original:
To undo a multiplication, divide:
Check: 20% of 120 is 24, and ✓
The tempting move is to add 20% to $96, and it does not work. 20% of 96 is 19.20, not 24, because the discount was 20% of the larger original price. That is why working backwards requires division, not a second percentage.
Why each wrong option is wrong:
- A, — added 20% to the sale price, which is the mistake just described.
- B, — added $20.
- C, — took another 20% off.
Takeaway: To undo a percentage change, divide by the multiplier. Adding the same percentage back always undershoots, because it is a percentage of a smaller number.
Topic: Successive percentage changes
A share price rises by 25% one year and falls by 20% the next. Over the two years, the price has
A)
B)
C)
D)
Show the worked solution
Answer: A
The options give the overall multiplier: 1 means the price is unchanged.
Explanation
Multiply the two multipliers together. Do not add the percentages.
The price ends exactly where it started.
Follow $100 through to see why. Up 25%: . Down 20%: 20% of 125 is 25, so . The same percentage applied to a bigger number is a bigger amount, which is exactly what makes the two changes cancel here.
Why each wrong option is wrong:
- B, — added the percentages, . Percentage changes never add.
- C, — subtracted in the other direction.
- D, — treated the two as a combined 55% loss.
Takeaway: Successive percentage changes multiply. Test with $100 — it takes ten seconds and settles the question completely.
Section 3 — One-Variable Data
| Measure | What it is | Careful of |
|---|---|---|
| Mean | total ÷ how many | dragged by extreme values |
| Median | the middle value, once sorted | sort first |
| Mode | the most common value | there can be several, or none |
| Range | largest − smallest | a spread, not a centre |
| Standard deviation | how spread out the values are | never negative |
Two facts the SAT tests repeatedly:
An outlier moves the mean far more than the median. A single very large value drags the mean up but shifts the median by at most one place. So if a question says a distribution is skewed, or mentions one unusually large value, the median is the more representative measure.
Standard deviation measures spread only. Two data sets can have the same mean and completely different standard deviations. Values bunched close together give a small standard deviation; values spread far apart give a large one. You will never be asked to calculate one — only to compare two sets and say which is larger.
Topic: Median of a data set
A researcher recorded six readings during a trial and wrote them down in the order they were taken: . What is the median of the data set?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
Sort the values first — this is the step people skip:
There are six values, an even number, so there is no single middle one. The median is the average of the middle two, which are 5 and 7:
Why each wrong option is wrong:
- A, — the mean, .
- C, — the mode, the most frequent value.
- D, — averaged the two middle values of the unsorted list.
Takeaway: Sort, then find the middle. Even count → average the middle two. And check which measure the question named; all four are offered.
Topic: The effect of an outlier
A quality inspector's data set is . A sixth reading of is then recorded. Which statement about the effect of that value is true?
A) The mean and median both increase by the same amount
B) The mean increases much more than the median
C) The median increases much more than the mean
D) Neither the mean nor the median changes
Show the worked solution
Answer: B
Explanation
Work out both, before and after.
Before: the values sorted are . The mean is and the median is the middle value, 14.
After adding 60: the values are . The mean is
and the median is the average of the middle two, .
So the mean jumped by about 7.7 while the median moved by 0.5.
The reason is structural. The mean uses the size of every value, so one huge number pulls hard. The median only uses position, so an extra value at the top shifts the middle by half a place regardless of how extreme it is. Replace the 60 with 6000 and the median would still be 14.5.
Why each wrong option is wrong:
- A — they change by very different amounts.
- C — backwards. The median is the resistant measure.
- D — both change; the median just changes very little.
Takeaway: Outliers drag the mean and barely move the median. When a question mentions one extreme value, or calls a distribution skewed, the median is the measure that represents it fairly.
Topic: Comparing standard deviations
Two laboratories each recorded five measurements of the same quantity. Set P is and set Q is . Which statement is true?
A) P has the greater standard deviation
B) They have equal standard deviations
C) Q has the greater standard deviation
D) Standard deviation cannot be compared without calculating it
Show the worked solution
Answer: C
Explanation
Both sets have a mean of 20. But standard deviation is not about the centre — it is about how far the values sit from the centre.
Every value in P is 20, so nothing deviates at all. P's standard deviation is zero — the smallest it can ever be.
Q's values sit 10, 5, 0, 5 and 10 away from the mean. There is real spread, so its standard deviation is clearly greater.
You never have to compute one on this test. Look at how tightly the values cluster: bunched means small, spread out means large.
A has it backwards. B ignores the spread entirely. D is wrong because comparison by inspection is exactly what the test wants here.
Takeaway: Standard deviation measures spread, not centre. Identical values give zero. Two sets can share a mean and differ completely in standard deviation.
Section 4 — Two-Variable Data and Probability
Scatterplots and lines of best fit. A line of best fit is a linear model drawn through a cloud of points. Its slope is a rate of change in context, and its -intercept is the predicted value at . Reading a value inside the data range is reliable; predicting far outside it is not, and the test sometimes asks you to say so.
Probability on this test is almost always counting:
The work is deciding what "altogether" means. In a two-way table, the denominator is whichever total the question restricts you to — the grand total, a row total, or a column total. Getting that wrong is the mistake the wrong options are built from.
Topic: Probability from a two-way table
A survey of 200 people recorded whether they cycle and whether they own a car.
| Owns a car | No car | Total | |
|---|---|---|---|
| Cycles | 30 | 50 | 80 |
| Does not cycle | 90 | 30 | 120 |
| Total | 120 | 80 | 200 |
One person is chosen at random from those who cycle. What is the probability that they own a car?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
Read the restriction: "from those who cycle". That fixes the pool. You are no longer choosing from all 200 people — only from the 80 who cycle.
Of those 80, the number who own a car is 30:
The phrase that sets the denominator is worth hunting for in every one of these questions. "Of those who cycle", "among car owners", "given that" — each one narrows the pool to a single row or column.
Why each wrong option is wrong:
- A, — used the grand total, 200. That answers "what is the probability a randomly chosen person both cycles and owns a car".
- C, — used 120, the car-owner total. That answers the reversed question: given they own a car, do they cycle?
- D, — the probability of being a cyclist at all.
Takeaway: Find the phrase that restricts the group; its total is your denominator. "Of those who X" means the X row, not the whole table.
Topic: Reading a line of best fit
The line of best fit for a scatterplot of study hours against test score is . What does the 8 represent?
A) The predicted increase in score for each extra hour studied
B) The predicted score for a student who does not study
C) The number of students in the study
D) The number of hours the average student studied
Show the worked solution
Answer: A
Explanation
The 8 multiplies , so it is the slope — the change in for each 1 added to . In context: each extra hour of study predicts 8 more points.
Check: and , a rise of 8 ✓
B describes the 42, the value at . C and D are quantities the model says nothing about — a line of best fit describes a relationship, not the size of the sample.
Note the word predicted. A line of best fit does not promise that an extra hour gives exactly 8 more marks; it describes the trend across the data.
Takeaway: In fitted to data, is "per one more " and is "at ". Check the units: a score per hour cannot be a number of students.
Topic: Inference from a sample
A researcher surveys 300 randomly selected students at a university and finds that 42% support a proposed change. Which conclusion is best supported?
A) Exactly 42% of all students at the university support the change
B) 42% of all students in the country support the change
C) It is plausible that close to 42% of students at that university support the change
D) No conclusion can be drawn, because only 300 students were asked
Show the worked solution
Answer: C
Explanation
Two rules decide every question of this kind.
Who was sampled? The survey drew from students at that university. A random sample supports conclusions about the population it was drawn from, and no further. Nothing here says anything about students elsewhere.
How precise can it be? A sample gives an estimate, not an exact figure. The right language is "approximately", "it is plausible that", "close to" — never "exactly".
Option C respects both. It stays inside the university and it hedges the number.
Why each wrong option is wrong:
- A — says exactly. A sample of 300 cannot pin down the whole university's opinion to the decimal.
- B — generalises beyond the population sampled. Students at one university are not a random sample of the country.
- D — too pessimistic. 300 randomly selected people support a reasonable estimate; that is the entire purpose of random sampling.
Takeaway: A random sample supports an estimate about the population it came from. Reject options that say "exactly", and reject options that reach beyond the group actually sampled.
Section 5 — Mixed Practice
Topic: Finding a percentage
A quality report needs one figure expressed as a percentage of another: 40 items out of a batch of 250 were set aside for inspection. What percent of 250 is 40?
A)
B)
C)
D)
Show the worked solution
Answer: A
Explanation
"What percent is of " means , with the part on top:
Check the size: 40 is well under half of 250, so the answer must be well under 50% ✓
B divided the wrong way round. C subtracted. D is the inverted division scaled up.
Takeaway: Part over whole, then times 100. If the answer comes out above 100% when the part is smaller than the whole, the division is upside down.
Topic: A unit rate
An office printer runs at a steady rate, producing 3 pages every 4 seconds without pausing between jobs. The office manager needs the output over a full minute. How many pages does it produce in one minute?
There are no options — work the answer out and enter it yourself, in the form the grid accepts.
Show the worked solution
Answer: 45
Explanation
Find the rate per second first:
Then a minute is 60 seconds:
Check by scaling instead: 60 seconds is 15 lots of 4 seconds, and each lot gives 3 pages, so ✓ Two routes, same answer.
This is a grid-in. Enter 45.
Takeaway: Reduce to a rate per one unit, then scale up. Or find how many whole blocks of the given time fit — often cleaner, and it avoids decimals.
Topic: Mean from a total
A teacher recorded five results and found that their mean was 18. Four of the five values were 12, 20, 15 and 25, and the fifth was lost. What is the fifth number?
A)
B)
C)
D)
Show the worked solution
Answer: C
Explanation
Turn the mean into a total. Five numbers with mean 18 must add to
The four given numbers add to . So the fifth is
Check: ✓
That the answer happens to equal the mean is a coincidence of these numbers, not a rule.
A subtracted wrongly. B is the total of all five. D is the total of the four.
Takeaway: Mean questions become easy the moment you convert to a total: total mean count.
Topic: A percentage of a percentage
In a town, 60% of residents own a bicycle. Of those bicycle owners, 25% ride daily. What percent of all residents ride daily?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
The 25% applies only to bicycle owners, not to everyone. So take a percentage of a percentage — which means multiply:
Follow 1000 residents through to see it. 60% own bicycles, so 600 people. A quarter of those ride daily: . And 150 out of 1000 is 15% ✓
A reported the 25% as though it applied to everyone. C added the percentages, D subtracted them — neither operation means anything here.
Takeaway: "Of those…" signals a percentage of a subgroup, so multiply the two. Testing with a round population of 100 or 1000 makes it concrete in seconds.
Topic: A percentage of a total from a table
The table shows how 400 people travel to work.
| Method | Car | Bus | Train | Cycle |
|---|---|---|---|---|
| People | 180 | 96 | 84 | 40 |
What percentage travel by bus?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
Check the size: 96 is a bit under a quarter of 400 ✓ And a quick sanity check on the whole table — , so nothing is missing.
A gave the raw count. D answered the opposite question.
Takeaway: Part over whole, times 100. Add the table up first; if it does not reach the stated total, you have misread a row.
Topic: Reading a rate from a table
A car's fuel use is recorded.
| Distance (km) | 0 | 50 | 100 | 150 |
|---|---|---|---|---|
| Fuel left (litres) | 60 | 55 | 50 | 45 |
How many litres does the car use per kilometre?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
Every 50 km uses 5 litres, so per kilometre:
Check the units the question asked for — litres per kilometre. Option B, 10, is kilometres per litre: the same relationship the other way up, and a perfectly sensible number that answers a different question.
A gave the fuel per 50 km without dividing.
Takeaway: Write the units on your answer and compare them with the units the question named. "Per" tells you which quantity goes on the bottom.
Topic: Probability of "at least one"
A bag has 3 red and 5 blue counters. One counter is drawn and not replaced, then a second is drawn. What is the probability that both are red?
A)
B)
C)
D)
Show the worked solution
Answer: C
Explanation
The first draw: 3 red out of 8, so .
The counter is not replaced, so for the second draw there are only 2 reds left out of 7 counters:
The words "not replaced" are doing all the work. A ignored them and used twice, which would be right with replacement.
Takeaway: "And then" means multiply. Check whether the first item goes back: without replacement, both the top and the bottom of the second fraction drop.
Topic: A weighted average
A class of 20 students averages 68 on a test. A second class of 30 students averages 78. What is the average across both classes?
A)
B)
C)
D)
Show the worked solution
Answer: A
Explanation
You cannot average two averages unless the groups are the same size — and here they are not. Go back to totals.
The plain average of 68 and 78 would be 73. The true answer is pulled towards 78, because more students sat in the class that scored 78. That pull is the whole point of the question.
Takeaway: Averages combine through totals, never by averaging the averages. The answer always leans towards the larger group.
Topic: Converting a rate between units
A tap fills a container at a steady rate of 3 litres per minute and is left running unattended. The volume delivered is wanted over a period of 2 hours. How many litres does it deliver in 2 hours?
There are no options — work the answer out and enter it yourself, in the form the grid accepts.
Show the worked solution
Answer: 360
Explanation
Convert to the unit the rate uses. Two hours is minutes, so
Check the direction: minutes are smaller than hours, so there are more of
them — multiply. Enter 360.
Takeaway: Convert time into the rate's own unit first. Smaller unit, bigger number.
Topic: Evaluating a statistical claim
A researcher wants to know the average number of hours students at a large school sleep. She surveys 40 students from the school's basketball team. What is the main problem with this study?
A) The sample is too small to say anything
B) Hours of sleep cannot be averaged
C) She should have surveyed exactly 100 students
D) The sample was not chosen at random from all students
Show the worked solution
Answer: D
Explanation
The problem is who was asked, not how many. A basketball team is not a random slice of the school: its members train, may keep different hours, and are selected for something. Conclusions from it apply to the team, not the school.
Forty people chosen at random from the whole school would support a reasonable estimate. Forty chosen from one team does not, however carefully the average is worked out.
A blames the size, which is not the fault here. C invents a magic number; there is no required sample size.
Takeaway: A sample supports conclusions about the group it was randomly drawn from. Ask who was left out before you ask how many were asked.
Topic: Percentage of a percentage
A shop reduces a coat by 20%, then takes a further 10% off the reduced price. What single percentage reduction is this equivalent to?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
Multiply the multipliers:
So 72% of the price remains, and the reduction is
Follow $100: down 20% to $80, then 10% off $80 is $8, leaving $72. A $28 reduction ✓
A added the percentages, which is the mistake the question exists for — the second 10% is taken off a smaller price, so it is worth less than 10% of the original. C gave what remains rather than what was taken off.
Takeaway: Successive percentage changes multiply. Then read whether the question wants the amount remaining or the amount removed.
Topic: Reading a scatterplot's fit
A scatterplot of 12 points has a line of best fit . Which is the best interpretation of the slope?
A) The 12 points all lie on the line
B) is always 2 less than
C) On average, falls by 2 for each 1 that rises
D) The largest value of is 30
Show the worked solution
Answer: C
Explanation
A slope is a rate: for each 1 added to , the fitted changes by . Because this is a line of best fit and not an exact rule, the honest wording is "on average".
A overstates it — a line of best fit passes near the points, not through them; if it did, it would not need to be a line of best fit. B describes , a different relationship entirely. D confuses the intercept with a maximum: 30 is the fitted value at , and the line goes above 30 for negative .
Takeaway: For a line of best fit, the slope is an average rate of change and the intercept is the fitted value at zero. Neither is a promise about any individual data point.
Topic: Ratio with a total to find
Concrete is mixed with cement, sand and stone in the ratio . If 12 kg of sand is used, what is the total mass of the mixture, in kilograms?
There are no options — work the answer out and enter it yourself, in the form the grid accepts.
Show the worked solution
Answer: 42
Explanation
Sand is 2 parts and weighs 12 kg, so one part is 6 kg.
The whole mixture is parts:
Check: cement 6, sand 12, stone 24, total 42 ✓ and simplifies to
✓ Enter 42.
Takeaway: Find the value of one part first. Every ratio question becomes multiplication after that.
Topic: Median from a frequency table
The table shows the number of siblings reported by 21 students.
| Siblings | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Students | 4 | 9 | 6 | 2 |
What is the median number of siblings?
A)
B)
C)
D)
Show the worked solution
Answer: A
Explanation
There are 21 students, so the median is the 11th value when they are lined up in order.
Count along: the first 4 students report 0 siblings (positions 1–4), the next 9 report 1 (positions 5–13). Position 11 falls inside that block, so the median is 1.
B is the mean, . C took the middle column of the table rather than the middle student — the table's columns are values, not people. D halved the number of students, which finds the position, not the value at it.
Takeaway: With a frequency table, find which position is the middle, then count along the frequencies to see which value sits there. The middle column is not the median.
Section 6 — Samples, Margins of Error and Claims
The test asks a small number of questions about what a study is entitled to conclude. They involve almost no calculation, and they are among the easiest marks on the paper once you know the three rules.
1. A margin of error gives a range. If a survey estimates 64% with a margin of error of 4%, the plausible values run from to . The estimate sits in the middle; the margin reaches out both ways.
The wording matters: it is likely that the true value lies in that range — never certain, and never exactly the estimate.
2. A bigger random sample gives a smaller margin of error. Asking more people does not change the estimate's direction, it narrows the range around it. Nothing else on this test changes the margin.
3. A conclusion may only reach as far as the sampling did. Two separate limits, and questions test both:
| What was done | What may be concluded |
|---|---|
| Randomly sampled from group X | something about group X, and nothing beyond it |
| Observed a relationship | that the two things are associated |
| Randomly assigned a treatment | that the treatment caused the difference |
That last row is the whole of "observational study versus experiment". If nobody assigned anything, no cause may be claimed — however strong the pattern.
Every option that says "proves", "exactly", or "causes" is suspect. The right answers use "likely", "plausible", "suggests" and "approximately".
Topic: Reading a margin of error
A poll of randomly selected voters estimates that 47% support a proposal, with an associated margin of error of 3%. Which is the most appropriate conclusion?
A) Exactly 47% of all voters support the proposal
B) It is likely that between 44% and 50% of all voters support the proposal
C) Between 44% and 50% of the voters polled support the proposal
D) At least 50% of all voters support the proposal
Show the worked solution
Answer: B
Explanation
The margin of error reaches out both ways from the estimate:
So the plausible range is 44% to 50%, and the conclusion is about all voters, not just the ones polled — that is the whole purpose of taking a random sample.
Why each wrong option is wrong:
- A — says exactly. A sample never pins down a population precisely; that is why a margin of error is quoted at all.
- C — applies the range to the people polled. Their support is a measured fact, not an estimate; the uncertainty is about everyone else.
- D — 50% is the extreme top of the range, not a floor. The poll is at least as consistent with 44%.
Takeaway: Estimate ± margin gives the range, and the claim is about the whole population. Reject any option saying "exactly" — and check whether a boundary value is being quoted as a minimum.
Topic: What makes a margin of error smaller
A researcher repeats a survey, using the same method but selecting 4,000 people at random instead of 400. Compared with the first survey, the second is most likely to have
A) a larger margin of error
B) the same margin of error
C) a smaller margin of error
D) no margin of error at all
Show the worked solution
Answer: C
Explanation
A larger random sample gives a better estimate, and "better" here means a narrower range around it — a smaller margin of error.
A has it backwards. B ignores the sample size, which is the one thing that changed. D overreaches: only surveying every single person would remove the uncertainty, and even a sample of 4,000 leaves some.
Note what does not change the margin: the estimate itself, the wording of the question, or how the results are presented. On this test, sample size is the lever.
Takeaway: Bigger random sample, smaller margin of error. Never zero unless everyone was asked.
Topic: Observational study versus experiment
Researchers recorded the sleep and the test scores of 500 randomly selected students, and found that students who slept longer tended to score higher. Which conclusion is best supported?
A) There is an association between hours of sleep and test scores among these students
B) Sleeping longer causes higher test scores
C) Higher test scores cause students to sleep longer
D) No relationship exists between sleep and test scores
Show the worked solution
Answer: A
Explanation
Nobody assigned anyone an amount of sleep. The researchers recorded what students already did, which makes this an observational study, and an observational study can establish that two things go together — an association — but not that one causes the other.
Both B and C claim a direction of cause, and the study cannot separate them. Nor can it rule out some third thing affecting both: a student who is ill, or working long hours, might sleep less and score lower for reasons that have nothing to do with sleep itself.
D contradicts the finding. A relationship was observed; it just cannot be called causal.
To support a causal claim, the researchers would have had to randomly assign students to different amounts of sleep. That is what makes something an experiment rather than an observation, and it is the only design on this test that licenses the word "cause".
Takeaway: Observed → association. Randomly assigned → cause. If nobody assigned anything, reject every option containing "causes".
Topic: How far a conclusion may be generalised
A random sample of 200 members of a large gym were asked how often they attend. The results are used to estimate attendance for the whole gym. Which statement best describes the limitation of this study?
A) The sample was too small to support any estimate
B) The estimate applies to every gym in the country
C) Random selection makes the estimate unreliable
D) The estimate applies to that gym's members, not to gym-goers generally
Show the worked solution
Answer: D
Explanation
The sample was drawn at random from one gym's members, so it supports conclusions about that gym's members and no further.
A blames the size. Two hundred people chosen at random support a perfectly reasonable estimate — the number is not the problem here.
B is the error the question is testing: stretching the conclusion past the population that was sampled.
C has it backwards. Random selection is what makes the estimate trustworthy; a sample of whoever happened to be there on a Monday morning would not be.
Takeaway: A random sample supports conclusions about the group it was drawn from. Ask "who could have been chosen?" — that group, and only that group, is what the estimate describes.
Topic: Choosing a model for two-variable data
The table shows five data points.
| 1 | 2 | 3 | 4 | 5 | |
|---|---|---|---|---|---|
| 12 | 19 | 29 | 41 | 52 |
Which is the most appropriate model for the relationship?
A) Increasing and roughly linear
B) Decreasing and roughly linear
C) Constant
D) Increasing and exponential, doubling each step
Show the worked solution
Answer: A
Explanation
Look at the differences between consecutive values:
They rise, and they are roughly the same size each time — so the relationship increases at a roughly constant rate. That is linear, allowing for the scatter real data always has.
B has the direction wrong; climbs throughout. C would need the values to stay put. D would mean multiplying by 2 each step: . The actual value at is 52, nowhere near 192.
Remember the distinction from Chapter 5: constant differences mean linear, constant ratios mean exponential. Here but , so the ratios are not constant.
Takeaway: Differences roughly constant → linear. Ratios roughly constant → exponential. Real data will never be exact; look for which is closer to constant.
Topic: Predicting inside and outside the data
A line of best fit for a scatterplot of house size (square metres) against price (thousands) is . The data covers houses from 50 to 200 square metres. Which prediction is least reliable?
A) A 60 square metre house costs about 220 thousand
B) A 120 square metre house costs about 400 thousand
C) A 200 square metre house costs about 640 thousand
D) A 900 square metre house costs about 2,740 thousand
Show the worked solution
Answer: D
Explanation
Every one of these four numbers is what the model gives — check them and they all come out right. So the arithmetic is not what separates them.
What separates them is where the prediction sits. The data covers houses from 50 to 200 square metres:
- A, B and C ask about 60, 120 and 200 — all inside that range, or at its edge. Predicting inside the data is what a line of best fit is for.
- D asks about 900 square metres, four and a half times the largest house ever measured. Nothing in the data says the relationship continues out there, and in reality it almost certainly does not.
Predicting beyond the data is called extrapolation, and it is the least reliable thing you can do with a model. The test asks about it directly.
Takeaway: A model is trustworthy inside the range of the data it was built from. A prediction far outside that range is the least reliable one, even when the arithmetic is perfect.
Topic: Comparing actual values with predicted ones
A line of best fit is . For the data point , how does the actual value compare with the predicted value?
A)
B)
C)
D)
Show the worked solution
Answer: B
Explanation
The model predicts
The actual value is 15, so the actual is 2 above the prediction:
A point sitting above the line of best fit means the model under-predicted it. Below the line means it over-predicted.
A subtracted predicted minus actual, which reverses the sign and so reverses the meaning. C and D each give one of the two values rather than the gap between them.
Takeaway: Substitute the into the model to get the prediction, then subtract it from the actual value. Above the line is positive; below is negative.
Topic: An experiment, where a causal claim is allowed
Two hundred volunteers with headaches were randomly assigned to receive either a new painkiller or a placebo. Those given the painkiller reported relief significantly more often. Which conclusion is best supported?
A) The painkiller is associated with relief, but no cause can be claimed
B) The painkiller will relieve every headache
C) For people like these volunteers, the painkiller caused more relief than the placebo
D) No conclusion is possible, because only 200 people took part
Show the worked solution
Answer: C
Explanation
This is the other side of Q29, and the difference is one word: assigned.
Because the researchers randomly assigned who got the painkiller and who got the placebo, the two groups were alike in every other respect on average. So when one group did better, the treatment is what explains it. That is an experiment, and an experiment does license a causal claim.
Why the other options fail:
- A would be right for an observational study — if the researchers had merely recorded which volunteers happened to take a painkiller. They did not; they assigned it.
- B overreaches in the other direction. "Significantly more often" is not "every time"; the study compares rates, not guarantees.
- D dismisses a perfectly workable study. Two hundred randomly assigned participants support a conclusion.
Note the careful wording of the correct option: "for people like these volunteers". The volunteers were not a random sample of everybody, so the causal claim holds for the kind of people who took part. Random assignment licenses cause; random selection licenses generalising to a wider population. This study had the first and not the second.
Takeaway: Random assignment → cause may be claimed. Random selection → the result may be generalised to the population sampled. They are different things, and a question will give you one, the other, or neither.
Topic: A conversion with a squared unit
A vehicle's speed increases at a rate of 8 metres per second squared. What is this rate in kilometres per second squared?
There are no options — work the answer out and enter it yourself, in the form the grid accepts.
Show the worked solution
Answer: .008 or 0.008
Explanation
The unit is metres per second squared — metres divided by (seconds × seconds). Only the metres are being changed; the seconds are the same on both sides, so they are left alone.
Check the direction: a kilometre is bigger than a metre, so the same rate is a smaller number of kilometres ✓
The "squared" is what makes this look harder than it is. It attaches to the seconds, and the seconds are not being converted — so it never enters the arithmetic at all. Had the question asked for metres per minute squared, the squared would matter and you would multiply by , not by 60.
This is a grid-in. Enter .008 or 0.008. Both fit the five-character
limit; 0.008 uses exactly five.
Takeaway: In a compound unit, convert one part at a time and leave the rest alone. A squared unit is squared only when that unit is the one being changed.
Topic: A two-step conversion
A storage tank has developed a slow leak, losing 3 litres per minute at a steady rate that shows no sign of changing. The maintenance report needs that figure expressed per day. What is the rate in litres per day?
A)
B)
C)
D)
Show the worked solution
Answer: D
Explanation
Two conversions, one at a time.
Minutes to hours:
Hours to a day:
Check the direction at each step: a day is much longer than a minute, so the number must get much bigger ✓
A stopped after one step, B used 24 without the 60, and C worked out how many minutes are in a day and forgot to multiply by the leak rate — a number that is correct about time and says nothing about litres.
Takeaway: Chain conversions one unit at a time, checking after each step whether the number should grow or shrink. Most wrong answers here are correct partial results.
Section 7 — Comparing and Transforming Data Sets
The test shows two distributions — as dot plots, histograms or lists — and asks you to compare them, or it changes every value in a set and asks what happened. Both are answered by thinking about centre and spread separately.
Adding the same number to every value slides the whole distribution along. Every measure of centre moves by that amount; every measure of spread stays exactly the same, because nothing has moved relative to anything else.
| Measure | Add to every value | Multiply every value by |
|---|---|---|
| Mean | increases by | multiplied by |
| Median | increases by | multiplied by |
| Range | unchanged | multiplied by |
| Standard deviation | unchanged | multiplied by |
That "unchanged" is the whole point, and it is worth seeing why: if every value moves 56 to the right, the gaps between them are exactly what they were.
Comparing two spreads by eye. You are never asked to calculate a standard deviation. You are asked which set is more spread out — so look at how tightly the values cluster around the middle. Values bunched near the centre give a small standard deviation; values pushed out to the extremes give a large one.
Topic: Adding a constant to every value
Data set B is made by adding 56 to each value in data set A. How do the median and the range of B compare with those of A?
A) The median and the range are both 56 greater
B) The median is 56 greater; the range is unchanged
C) The median is unchanged; the range is 56 greater
D) The median and the range are both unchanged
Show the worked solution
Answer: B
Explanation
Adding the same amount to every value slides the whole distribution. The middle value slides with it, so the median rises by 56.
The range is largest minus smallest. Both ends move up by 56, so the difference between them does not change at all:
Try it on five values: has median 8 and range 11. Adding 56 gives — median 64, which is ✓ and range , unchanged ✓
The same argument covers the standard deviation, which is also a measure of spread: sliding everything changes no gap between any two values.
Takeaway: Adding a constant shifts every measure of centre and leaves every measure of spread alone. Multiplying by a constant scales both.
Topic: Comparing the spread of two distributions
Two classes each recorded 10 values.
Class A: Class B:
Both have a mean of 8.1. Which statement is true?
A) Class A has the greater standard deviation
B) The standard deviations are equal, because the means are equal
C) Class B has the greater standard deviation
D) The standard deviation cannot be compared without calculating it
Show the worked solution
Answer: C
Explanation
Same mean, different spread — which is exactly what standard deviation is for.
Class A's values run from 6 to 10, a span of 4, and most sit at 8 or 9. Class B runs from 1 to 15, a span of 14, with values pushed out to both extremes.
B is far more spread out, so B has the greater standard deviation.
B's option confuses centre with spread: equal means say nothing about how tightly the values cluster. D is wrong because comparing by inspection is precisely what the test expects — you are never asked to compute one.
Takeaway: Look at how tightly the values cluster around the middle. Equal means with different clustering is the standard setup for this question.
Topic: The median from a grouped frequency table
The table groups 23 integers.
| Interval | 10–19 | 20–29 | 30–39 | 40–49 |
|---|---|---|---|---|
| Frequency | 4 | 7 | 9 | 3 |
Which interval contains the median?
A) 10–19
B) 20–29
C) 30–39
D) 40–49
Show the worked solution
Answer: C
Explanation
With 23 values, the median is the 12th when they are lined up in order.
Count along the intervals, keeping a running total:
| Interval | Frequency | Running total |
|---|---|---|
| 10–19 | 4 | 4 |
| 20–29 | 7 | 11 |
| 30–39 | 9 | 20 |
| 40–49 | 3 | 23 |
After two intervals you have reached only the 11th value, so the 12th is in the next one: 30–39.
B is the trap: 20–29 is the interval where the count is closest to halfway, but it stops one short. Run the total to 12, not to "about half".
Takeaway: Find which position is the middle, then run the frequencies until you pass it. The interval you are in when you pass it holds the median.
Topic: Multiplying every value in a data set
Every value in a data set is multiplied by 3. What happens to the mean and to the range?
A) Both are multiplied by 3
B) Both are unchanged
C) The mean is multiplied by 3; the range is unchanged
D) The mean is unchanged; the range is multiplied by 3
Show the worked solution
Answer: A
Explanation
Multiplying stretches the distribution rather than sliding it, so both the centre and the spread scale.
Try : the mean is 5 and the range is 7. Tripling gives — mean 15, which is ✓ and range , which is ✓
Contrast this with Q37. Adding slides everything and leaves the gaps alone, so spread is unchanged. Multiplying stretches the gaps too, so spread scales with everything else. Knowing which operation was applied decides the answer.
Takeaway: Add → centre moves, spread unchanged. Multiply → both scale. The two questions look alike and have different answers.