AP Statistics — exploring data

Units 1 and 2 of AP Statistics cover describing one variable and then describing the relationship between two, and together they are worth roughly a fifth to a quarter of the exam. This set takes ten of their ideas, one question each. Every distribution and data set is described in words with the numbers written into the sentence, because there are no figures here.

Several questions are built so the answer can be computed rather than recalled. Nine ordered values have a tenth of 200 added to them: the median moves from 9 to 9.5 because it depends only on position, while the mean goes from 12 to 30.8 and the standard deviation rises sharply, because both depend on magnitude. That split is the whole of what resistance means, and seeing it in two arithmetic steps is more useful than the word. The fence question gives quartiles of 20 and 50, so the interquartile range is 30, one and a half of those is 45, and the fences sit at −25 and 95 — measured outward from the quartiles rather than from the centre, which is the error that would wrongly flag a value of 60.

Three questions are about what a statistic does not say. The correlation coefficient measures linear association and nothing else, so points lying exactly on a parabola can give r near zero while the relationship is perfect — which is why r = 0 rules out a straight-line trend and rules out nothing else. The coefficient of determination reports a proportion of variation, not a probability that the model is right, and a model can post a high r² while being the wrong shape entirely. And a U-shaped residual plot means the line has missed something systematic; the tempting wrong answer, that the residuals balance out, is true of every least-squares fit by construction and therefore carries no information at all.

The last question is the one worth carrying furthest. A point far out at one end of the x-range has leverage: the fitted line is dragged towards it, so removing it would visibly change the slope — and precisely because the line has been dragged towards it, its residual is often small. Judging influence by the size of a residual therefore misses the points that matter most, while a point sitting far from the line in the middle of the range mostly inflates the spread of residuals rather than tilting anything.

Every explanation names which misunderstanding produces which wrong answer. Every data set, context and number here is invented and is attributed to no real study, institution or place, and nothing is drawn from any College Board publication, equations sheet or reference table.

  • Choose between the mean and the median for a skewed distribution, and state resistance as the reason
  • Predict how a single extreme observation moves the mean, median, standard deviation and interquartile range differently, and compute the change
  • Interpret a standard deviation in the units of the variable measured, against the mean
  • Compare values from two different distributions by standardising, and explain why equal gaps above two means are not equal standing
  • Apply the $1.5 \times \text{IQR}$ rule outward from the quartiles to identify a flagged value
  • State what the correlation coefficient measures and the three things it does not — curvature, causation, and any dependence on units
  • Interpret a least-squares slope as a rate, and judge whether the intercept means anything in context
  • Read a residual plot, and explain why balanced residuals are guaranteed rather than informative
  • Interpret $r^2$ as a proportion of variation explained, distinguishing it from $r$ and from a probability
  • Identify an influential point by leverage in the $x$-direction, and explain why its residual is often small

Built against the published structure of AP Statistics, Unit 1: Exploring One-Variable Data and Unit 2: Exploring Two-Variable Data. The exam runs 3 hours and is taken digitally: 42 multiple-choice questions in 1 hour 30 minutes for 50 per cent of the score, then four free-response questions in 1 hour 30 minutes for the remaining 50 per cent, each worth ten points. A graphing calculator with statistical capabilities is expected, and an equations sheet and reference tables are supplied during the exam. Every data set, context and number in this material is invented, and none is attributed to a real study, institution or place. Nothing is reproduced from any College Board publication, released exam, course and exam description, scoring guideline, equations sheet or reference table. Zestly is an independent study tool. It is not affiliated with the College Board, which owns the AP Statistics exam, and it is not an exam centre.

Sample question

House prices in a small village are strongly skewed to the right: most houses cluster together and a handful of very large properties sit far out in the upper tail. Which measure of centre better describes what a typical house costs, and why?

See the answer

The median, because it depends on the position of the middle value rather than on the size of the extreme ones, and so is not dragged into the tail

In a right-skewed distribution the mean sits above the median, pulled up by the far values, so it can exceed what most houses actually cost — a figure that is arithmetically correct and misleading as a description of the typical case. The median splits the ordered values in half and moves only when the middle moves, so a mansion twice as expensive shifts it not at all. That property is called resistance, and it is the reason for the choice. Using every observation is a genuine merit of the mean and exactly why it fails here: the extreme values are included at full weight. Calling that sensitivity an advantage confuses describing the SHAPE with describing the CENTRE — the gap between mean and median is itself useful evidence of skew, but neither number is thereby a better summary of the typical value. And the mean and median coincide only in a symmetric distribution, which is precisely what this one is not.

Try this quiz →Try this exam →Try this written work →

← Statistics

↑ AP