Data 1 — percentages, spread and what a sample proves

Thirteen questions across the whole domain — ratios and rates, percentages, one- and two-variable data, probability from a two-way table, and inference from a sample. The arithmetic in almost all of them is easy. The difficulty is reading precisely, and the questions are built around the specific places where reading loosely gives you a wrong answer that feels right.

Percent against percentage points. A rate that goes from 5% to 7% has risen by two percentage points and by forty percent, and those are different sentences describing the same event. A question asking "by what percent did it increase" has one answer, and the other number is sitting among the options. Both of the other wrong choices here are real arithmetic errors: dividing the change by the finishing value instead of the starting one, and giving the ratio of new to old rather than the increase.

Mean against median. An outlier drags the mean and leaves the median more or less where it was. Removing a low score raises the mean of what is left. Neither of these needs a calculation, and both appear as questions.

Standard deviation is spread. Two sets with the same mean and the same number of values can look nothing alike: 18, 19, 20, 21, 22 against 4, 12, 20, 28, 36. You are never asked to compute a standard deviation on this test, only to say which set is more spread out — and the wrong answers here are the two features the sets genuinely share, neither of which has anything to do with spread.

Conditional probability. "Of the people who exercise regularly, what fraction sleep well" has a denominator of the regular exercisers, not of everybody. Deciding what the denominator is before computing anything is the whole question.

What a study lets you say. The most tested idea in the statistics part of the SAT, in two halves, and both are here:

  • Random selection lets you generalise — but only to the population you selected from. A random sample of people leaving one train station tells you about people leaving that train station, and not about every adult in the city.
  • Random assignment is what licenses a causal claim. An observational study, however large and however clean, shows an association. The wrong answers overreach in exactly these two directions, every time.

And a third failure worth naming: a mean is not a statement about any individual. A mean commute of 47 minutes is entirely consistent with commutes of 10 minutes and 90 minutes.

Margin of error in one sentence: it is the range in which the true population value plausibly sits, and it shrinks as the sample grows.

The twelve flashcards carry the rules as sentences you can recite, which is what these questions actually reward — this domain is only 5 to 7 questions on the real test, and most of them are won or lost on a definition.

  • Compute a percent change against the starting value, and tell it apart from a change in percentage points
  • Predict what happens to the mean and to the median when an outlier is added or removed
  • Compare the spread of two data sets without computing a standard deviation
  • Decide the correct denominator before computing a conditional probability
  • Generalise a random sample only to the population it was actually drawn from
  • Refuse a causal claim from an observational study, however large
  • Read a mean as a property of the group and not of any individual in it
  • State what a margin of error does and does not tell you
  • Convert between rates and totals with the units under control

A researcher surveyed 200 people chosen at random from those leaving one train station on a weekday morning, and found a mean commute of 47 minutes. Which conclusion is best supported? — Random selection lets you generalise, but only to the population you selected FROM: people leaving that station on a weekday morning, not every adult in the city. Nobody was assigned to a means of transport, so no causal claim is available; and a mean of 47 minutes is entirely consistent with commutes of 10 minutes and 90 minutes.

Sample question

A store sells a laptop for 800 dollars. During a sale, the price is reduced by 20 percent. What is the new price of the laptop?

See the answer

640 dollars

A 20 percent reduction means the new price is 80 percent of the original. $800 \times 0.80 = 640$. 600 dollars is a common error of subtracting 200 instead of 20 percent, and 960 dollars is an error of adding 20 percent.

Try this quiz →

← Problem-Solving and Data Analysis

↑ SAT