Skip to content
All articles

A-Level · 6 min read

The large data set: what you actually need to know

Students either ignore the large data set completely or try to memorise it. Neither is what the questions require.

By Joseph Eno ·

The large data set is a real data file, published by your exam board, that you are expected to have worked with during the course. Questions in the statistics section may refer to it, and familiarity gives you a genuine advantage on those marks.

What examiners actually ask

  • Interpreting summary statistics in the context of the variables in the file.
  • Recognising which variables are qualitative, discrete or continuous.
  • Commenting on whether a proposed model or correlation is sensible for that context.
  • Spotting units, missing values and the limitations of the sampling.

How to prepare efficiently

  1. Open the file in a spreadsheet and read the variable descriptions properly, including units.
  2. Produce a few plots yourself — scatter, box plot, histogram — for the variables your board favours.
  3. Note typical values so an implausible answer in the exam looks wrong to you immediately.
  4. Work through past questions that reference the set; the same interpretive themes recur.

An hour or two of genuine handling beats any amount of reading about it, and the context marks are among the most accessible in the statistics section.

Tell me what your child is finding difficult.

No pressure and no sales call — just an honest conversation about whether I can help, and how.