Skip to topic
    ← Back to course topics

    L: Data presentation and interpretation — AQA A-Level Mathematics

    Test yourself on L: Data presentation and interpretation with AQA A-Level practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    L: Data presentation and interpretation explained

    Single-variable data can be displayed in diagrams such as histograms, bar charts, box plots and cumulative frequency graphs.

    Read the full explanation

    In a histogram, bars are drawn so that area, not height, represents frequency; this matters when class widths are unequal. Frequency density is calculated as frequency divided by class width, and the area of each bar equals frequency. Interpreting these diagrams involves reading frequencies, identifying modal classes and estimating probabilities. Connecting to probability distributions means recognising that relative frequencies from data can estimate probabilities, and that the shape of a histogram can suggest an underlying distribution. For example, a histogram of waiting times with unequal classes: a taller bar does not necessarily mean a higher frequency unless width is considered. Always check class width before comparing bars.

    Interpret scatter diagrams and regression lines for bivariate data, including recognition of scatter diagrams which include distinct sections of the population (calculations involving regression lines are excluded). Understand informal interpretation of correlation. Understand that correlation does not imply causation.

    This topic develops your ability to read and describe relationships between two variables using scatter diagrams and given regression lines, without calculating the line yourself. You must identify overall patterns, spot distinct subgroups within the data, and describe correlation informally using terms such as strong, weak, positive or negative. For example, a scatter diagram of revision hours against test scores may show a positive correlation, but if points form two separate clusters, you should recognise that the population may contain distinct sections. You also need to understand that a regression line helps predict one variable from another, but only within the observed range. Crucially, you must explain why correlation does not imply causation, using clear reasoning about possible confounding variables.

    Interpret measures of central tendency and variation, extending to standard deviation. Be able to calculate standard deviation, including from summary statistics.

    This topic requires you to interpret and calculate measures of central tendency (mean, median, mode) and variation (range, interquartile range, standard deviation). You should understand that standard deviation measures the spread of data around the mean, with a larger value indicating greater variability. You must be able to calculate standard deviation from raw data using the formula, and also from summary statistics such as Σx, Σx² and n. For example, given Σx = 120, Σx² = 2000 and n = 10, you can find the mean and then the standard deviation. You should also compare data sets using these measures and explain what they tell you in context, including how outliers affect them.

    Recognise and interpret possible outliers in data sets and statistical diagrams. Select or critique data presentation techniques in the context of a statistical problem. Be able to clean data, including dealing with missing data, errors and outliers.

    Outliers are values that lie an unusual distance from the bulk of the data. A common rule is that a value is an outlier if it is more than 1.5 times the interquartile range below the lower quartile or above the upper quartile. You must recognise such values in lists, box plots, histograms or scatter diagrams, and interpret their possible causes. When choosing or criticising a presentation, consider whether it shows the distribution clearly, whether the scale is misleading, and whether the diagram suits the data type. Cleaning data means checking for missing entries, impossible values such as negative heights, and outliers, then deciding whether to remove, correct or keep them, and recording your decision.

    Your focus

    1. Calculate frequency density and use it to interpret or construct a histogram.
    2. Find frequencies from a histogram by calculating areas of bars.
    3. Estimate probabilities from relative frequencies shown in a diagram.
    Show all 14 objectives
    1. Describe how the shape of a data diagram can suggest features of a probability distribution.
    2. Describe the direction, strength and form of a relationship shown in a scatter diagram.
    3. Recognise and explain the effect of distinct sections of a population within a scatter diagram.
    4. Use a given regression line to make contextual predictions and explain why correlation does not imply causation.
    5. Calculate standard deviation from raw data and from summary statistics.
    6. Interpret measures of central tendency and variation in context, including standard deviation.
    7. Compare data sets using appropriate measures and explain the effect of outliers.
    8. Calculate quartiles and the interquartile range, and apply the 1.5 × IQR rule to identify possible outliers.
    9. Interpret outliers in context and justify whether to retain, correct or remove them.
    10. Evaluate and critique data presentation techniques, explaining how well a chosen diagram represents the data.
    11. Describe and apply a systematic process for cleaning data, including missing values, errors and outliers.

    L: Data presentation and interpretation exam tips

    Marking Points
    • Identifies that in a histogram the area of each bar represents the frequency of the class, not the height alone.
    • Calculates frequency density using frequency divided by class width, and uses it to draw or interpret histogram bars correctly.
    • Reads frequencies from a histogram by multiplying frequency density by class width for each bar.
    • Interprets other single-variable diagrams such as bar charts, box plots or cumulative frequency graphs, extracting relevant values such as median, quartiles or frequencies.
    • Connects relative frequency from a diagram to probability, for example estimating the probability that a randomly chosen item falls in a given class.
    • Recognises that the overall shape of a histogram or other diagram can suggest features of a probability distribution, such as skewness or the location of the mode.
    • Describe the direction of a relationship as positive or negative, and the strength as strong, moderate or weak, referring to how closely points follow a straight-line pattern.
    • Identify distinct sections or clusters within a scatter diagram and explain that these may represent different subgroups of the population, which can affect the overall interpretation.
    • Use a given regression line to make predictions or estimates, recognising that predictions are reliable only within the range of the data and that extrapolation may be unreliable.
    • Explain informally that correlation measures association, not cause, and give a plausible reason such as a lurking variable or coincidence.
    • Interpret the meaning of the variables in context, including units, and comment on whether a linear model is appropriate.
    • Calculate the mean, median and mode from raw data or summary statistics and interpret them in context.
    • Calculate standard deviation using the formula s = √(Σx²/n − (Σx/n)²) or the equivalent form involving deviations from the mean.
    • Use summary statistics such as Σx, Σx² and n to find the mean and standard deviation without access to the original data.
    • Compare two data sets using measures of central tendency and variation, explaining what the differences indicate about the data.
    • Recognise the effect of outliers on the mean, range and standard deviation, and contrast this with the median and interquartile range.
    • Correctly calculate the interquartile range and use the 1.5 × IQR rule to identify possible outliers.
    • Interpret an outlier in context, suggesting a plausible reason such as a recording error, a genuine extreme value, or a different population.
    • Evaluate a statistical diagram by commenting on its suitability, for example a histogram for continuous data or a box plot for comparing distributions.
    • Critique a presentation by identifying misleading features such as a truncated axis, unequal class widths without frequency density, or a 3D effect that distorts areas.
    • Describe a data-cleaning process that checks for missing values, impossible values and outliers, and states what action is taken for each.
    • Justify whether to remove or retain an outlier, referring to the context and the effect on summary statistics.
    Examiner Tips
    • 💡Before answering a histogram question, check whether the vertical axis is labelled frequency or frequency density; if it is frequency density, use area to find frequencies.
    • 💡When estimating a probability from a diagram, write the relative frequency as a fraction or decimal and relate it to the total frequency, not to the number of bars.
    • 💡If asked to connect a diagram to a probability distribution, comment on the shape, centre and spread, and explain how relative frequencies approximate probabilities for large samples.
    • 💡Always refer to the context and the variables when describing correlation, rather than just saying 'positive correlation'.
    • 💡When asked about causation, explicitly use the phrase 'correlation does not imply causation' and support it with a specific alternative explanation.
    • 💡Check whether the scatter diagram has distinct groups before making a single overall judgement; mention them if they are present.
    • 💡Show all steps in standard deviation calculations, including the mean and the sum of squares, to gain method marks.
    • 💡When comparing data sets, always quote numerical values and relate them to the context.
    • 💡Use the correct notation and units throughout, and round final answers appropriately as indicated by the question.
    • 💡When asked to identify outliers, show the quartiles, the IQR and the boundary calculations clearly so the method is visible.
    • 💡In critique questions, make at least two distinct points about the diagram, such as axis scale, choice of average, or clarity of labels.
    • 💡For data cleaning, structure your answer as a checklist: missing data, errors, outliers, and the action taken for each.
    • 💡Use the context in your interpretation; a numerical answer without a contextual comment may not gain full credit.
    Common Mistakes
    • Assuming the height of a histogram bar is the frequency. Correction: frequency is represented by area; height is frequency density, so frequency = frequency density × class width.
    • Comparing histogram bars of different widths by height alone. Correction: compare areas, or calculate frequencies, because a wider class with the same height contains more data.
    • Treating a bar chart and a histogram as interchangeable. Correction: bar charts are for discrete categories with gaps and height represents frequency; histograms are for continuous data with no gaps and area represents frequency.
    • Assuming that a strong correlation proves that one variable causes the other. Correction: state that correlation only shows association and suggest a possible confounding variable.
    • Treating a regression line as a perfect summary and using it to predict far outside the observed data. Correction: restrict predictions to the range of the data and note that extrapolation is unreliable.
    • Ignoring distinct clusters and describing the whole scatter diagram as one homogeneous group. Correction: identify the separate sections and explain how they might change the overall correlation.
    • Forgetting to take the square root when calculating standard deviation, leaving the variance instead. Correction: always check that the final step is to square root the variance.
    • Mixing up the formulas for population and sample standard deviation. Correction: in AQA A-level Mathematics, use the formula for standard deviation as given in the specification, which divides by n.
    • Misinterpreting a larger standard deviation as meaning the data are 'better' or 'more consistent'. Correction: a larger standard deviation means greater spread or variability, not necessarily better.
    • Using the range instead of the interquartile range when applying the 1.5 × IQR rule. Correction: always find Q1, Q3 and IQR first, then test values below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR.
    • Assuming every outlier is an error and deleting it automatically. Correction: investigate the context; an outlier may be a genuine extreme value that should be kept and discussed.
    • Choosing a pie chart for continuous data or a histogram for categorical data. Correction: match the diagram to the data type—pie charts for proportions of a whole, histograms for continuous grouped data, bar charts for categories.
    • Ignoring missing data or treating blanks as zero. Correction: identify missing entries, decide whether to exclude them or estimate them, and state the effect on the analysis.