AP Statistics — statistical inference

The last four units of AP Statistics are the ones candidates lose marks on for reasons that have nothing to do with arithmetic. Almost every mark here turns on saying precisely what a procedure does and does not claim, and the ten questions in this bank are built around the sentences that go wrong. A confidence interval comes first: the ninety-five per cent belongs to the method across many samples that were never taken, not to the one interval now sitting on the page, which either contains the parameter or does not. The second question sets confidence against precision — more data narrows an interval, more confidence widens it, and the tempting fourth option, drawing a second sample and keeping whichever interval came out narrower, destroys the very property being claimed.

Hypotheses follow, stated about the population parameter rather than about a sample mean already in hand, with the equality in the null because that is what the test assumes while working out how surprising the data would be. Then the p-value, asked so that all three of the usual misreadings appear side by side: that it is the probability the null is true, that its complement is the probability the alternative is true, and the loose everyday phrase about a result arising purely by chance — which is the right answer with the conditional assumption quietly dropped, and is the hardest of the three to shake. The decision question separates failing to reject from proving, and the error question is put in a quality manager's own words: a Type I error is stopping a line that was working properly, and the significance level is the long-run rate of exactly that false alarm, while letting a bad line run has a rate that depends on how far things have drifted and on how many components were inspected.

The closing four are about judgement. Power rises with the sample and falls when the significance level is tightened, so the two cannot both be improved without collecting more data. A study of five million people can make a difference of a few thousandths of a second significant, which is why the size of an effect has to be reported and judged separately from its p-value. One question asks which procedure fits a described study, with a proportion test, a mean test, a two-sample comparison and a chi-square test for independence all on offer. And the capstone is the point the whole sequence rests on: ten people asked in a cafeteria cannot support an interval for a whole college, and fifty people asked the same way would buy precision about the wrong population rather than repair anything. No procedure fixes a badly collected sample.

Every institute, firm and study named here is invented, and each question says so. Nothing is drawn from any College Board publication, released examination, course and exam description, scoring guideline, formula sheet or statistical table. There are no tables or computer output — every study is described in words, which keeps the questions readable anywhere but also means the output-reading that some real inference questions require is not practised here.

Zestly is an independent study tool. It is not affiliated with the College Board, which owns the AP Statistics exam, and it is not an exam centre.

  • Interpret a confidence interval as a statement about the method rather than about the one interval calculated
  • Say what widens and what narrows an interval, and why confidence and precision pull against each other
  • State a null and an alternative hypothesis about a population parameter, with the equality in the null
  • Define a p-value as a probability of data computed on the assumption that the null holds
  • Compare a p-value with a significance level, and distinguish failing to reject from proving
  • Describe a Type I and a Type II error in the words of a given situation, and say which one the significance level controls
  • Name what raises and what lowers the power of a test, including the cost of tightening the significance level
  • Separate statistical significance from practical importance, particularly in very large samples
  • Choose between a proportion test, a mean test, a two-sample comparison and a chi-square test for a described study
  • Check the conditions for inference, and explain why enlarging a convenience sample repairs nothing

At the invented firm Titan Manufacturing a quality manager tests each batch against the claim that the proportion of defective components is no higher than the permitted level, and stops the production line when the evidence says otherwise. What is a Type I error here, and which error does the significance level control?

Sample question

A researcher at the invented Zenith Institute reports a ninety-five per cent confidence interval for the mean height of a plant species. Which statement interprets that interval correctly?

See the answer

If the whole sampling procedure were repeated many times, about ninety-five per cent of the intervals it produced would contain the true mean height

The confidence level is a property of the procedure rather than of the one interval it happened to produce. Once the interval has been calculated it either contains the true mean or it does not, and there is nothing left to be probable about — which is what rules out attaching a ninety-five per cent probability to this particular interval. An interval for a mean says nothing about where individual plants fall, so the reading about the heights of ninety-five per cent of the plants confuses an estimate of a parameter with a range covering the population itself. And the interval is built around one sample mean rather than describing where other sample means would land, which is what the reading about all possible sample means supposes.

Try this quiz →Try this exam →Try this written work →

← Statistics

↑ AP