Skip to topic
    โ† Back to course topics

    Data presentation and interpretation โ€” Edexcel A-Level Mathematics

    Test yourself on Data presentation and interpretation with PEARSON EDEXCEL A-Level practice questions.

    Start free

    7 days Premium ยท Then free forever ยท No card, no charge

    Data presentation and interpretation explained

    Single-variable diagrams summarise one measured variable.

    Read the full explanation

    In a histogram, unequal class widths mean frequency is represented by the area of each bar, not its height. Frequency density is frequency divided by class width, so a bar of width 5 and height 4 represents a frequency of 20. To interpret such a diagram, read each bar's width and height, multiply them to recover frequency, and compare areas when judging which class is most common. This links directly to probability distributions: relative frequency, found by dividing each frequency by the total, estimates probability, and a histogram of relative frequency density has total area 1. Thus the same area principle underlies continuous probability density functions, where probability equals the area under the curve between two values.

    2.2 Interpret scatter diagrams and regression lines for bivariate data, including recognition of scatter diagrams which include distinct sections of the population (calculations involving regression lines are excluded). Understand informal interpretation of correlation. Understand that correlation does not imply causation.

    Bivariate data pairs two variables for each individual. A scatter diagram plots one variable against the other, and a regression line summarises a linear trend, allowing a prediction of one variable from the other within the observed range. Interpretation means describing direction, strength and form, and recognising distinct sections of the population, such as two clusters or a group of points that follow a different pattern; these may indicate subgroups that should be described separately rather than forced into one line. Correlation is an informal measure of how closely points follow a straight line, from strong positive through weak to strong negative. Crucially, correlation does not imply causation: a third variable or coincidence may explain the association, so causal claims need evidence beyond the scatter diagram.

    2.3 Interpret measures of central tendency and variation, extending to standard deviation. Be able to calculate standard deviation, including from summary statistics.

    Measures of central tendency, such as the mean, median and mode, describe a typical value, while measures of variation, such as range, interquartile range and standard deviation, describe spread. Standard deviation measures how far values typically lie from the mean, in the same units as the data. For this specification, standard deviation uses division by n. From summary statistics, use the formula standard deviation equals the square root of the quantity (sum of x squared divided by n) minus the mean squared. For example, if sum of x = 20, sum of x squared = 130, and n = 5, the mean is 4 and standard deviation is sqrt(130/5 - 16) = sqrt(26 - 16) = sqrt(10) = 3.16. Interpret a larger standard deviation as greater spread about the mean.

    2.4 Recognise and interpret possible outliers in data sets and statistical diagrams. Select or critique data presentation techniques in the context of a statistical problem. Be able to clean data, including dealing with missing data, errors and outliers.

    An outlier is a value that lies far from the rest of the data. A common rule flags values below the lower quartile minus 1.5 times the interquartile range or above the upper quartile plus 1.5 times the interquartile range. Outliers may be genuine extreme values or errors, so interpret them in context before deciding what to do. Selecting or critiquing presentation means choosing a diagram or summary that suits the data, for example a box plot for spread and outliers or a histogram for distribution shape, and explaining the strengths and limitations of that choice. Cleaning data means checking for missing values, impossible or misrecorded entries and outliers, then deciding whether to correct, exclude or retain each one and stating the effect on the analysis.

    Your focus

    1. Calculate frequency from the area of a histogram bar and frequency density from frequency and class width.
    2. Compare classes in a histogram using areas and explain the comparison in context.
    3. Convert a frequency distribution to relative frequencies and relate the total area to probability.
    Show all 12 objectives
    1. Describe direction, strength and form of association from a scatter diagram.
    2. Identify and describe distinct sections of a population shown in a scatter diagram.
    3. Explain why an observed correlation does not by itself establish a causal relationship.
    4. Calculate a mean and a standard deviation from raw data or from summary statistics using division by n.
    5. Select an appropriate measure of central tendency and of variation for a given data set.
    6. Interpret and compare standard deviations of two data sets in the context of the problem.
    7. Identify possible outliers using quartiles and the 1.5 times interquartile range rule.
    8. Justify the choice of a data presentation technique for a given statistical problem.
    9. Clean a data set by handling missing values, errors and outliers and explain the decisions made.

    Data presentation and interpretation exam tips

    Marking Points
    • States that in a histogram with unequal class widths, frequency equals frequency density multiplied by class width, so area represents frequency.
    • Calculates a missing frequency by multiplying a bar's height by its class width, and a missing height by dividing frequency by class width.
    • Compares classes correctly by area rather than by bar height when class widths differ.
    • Explains that relative frequency estimates probability and that total relative frequency is 1.
    • Links histogram area to probability density, where probability corresponds to area under a curve between two values.
    • Describes a scatter diagram in terms of direction, strength and form of association between the two variables.
    • Interprets a regression line as a model of a linear trend and uses it to predict one variable from the other within the data range.
    • Recognises distinct sections or clusters in a scatter diagram and describes each subgroup separately.
    • Explains informally what a correlation coefficient indicates about the closeness of points to a straight line.
    • States that correlation does not imply causation and suggests a possible confounding or alternative explanation.
    • Distinguishes measures of central tendency from measures of variation and states what each describes.
    • Calculates a mean from raw data or summary totals and uses it in further work.
    • Applies the standard deviation formula, ensuring division by n as required by the specification.
    • Uses summary statistics such as sum of x, sum of x squared and n to find a standard deviation without the raw data.
    • Interprets standard deviation as a measure of spread in the original units and compares two data sets using it.
    • Identifies possible outliers using a stated rule such as quartiles and 1.5 times the interquartile range.
    • Interprets an outlier in context, distinguishing a genuine extreme value from a likely error.
    • Selects a suitable data presentation technique for a given statistical problem and justifies the choice.
    • Critiques a given presentation by commenting on its suitability, clarity or possible distortion.
    • Cleans a data set by identifying missing values, errors and outliers and explaining the treatment of each.
    Examiner Tips
    • ๐Ÿ’กLabel the vertical axis frequency density and show one area calculation to justify a frequency you read from a histogram.
    • ๐Ÿ’กWhen asked to compare two classes, quote both areas and state the conclusion in the context of the data.
    • ๐Ÿ’กConvert frequencies to relative frequencies before discussing probability, and check that they sum to 1.
    • ๐Ÿ’กUse the words positive, negative, strong, weak and linear when describing association, and quote the context of the variables.
    • ๐Ÿ’กWhen a diagram shows clusters, state how many sections you see and describe the trend within each.
    • ๐Ÿ’กFor any causal-sounding question, add a sentence explaining why correlation alone cannot establish cause.
    • ๐Ÿ’กWrite the formula you are using before substituting values, and keep the mean to full accuracy until the final rounding.
    • ๐Ÿ’กRemember that standard deviation calculations in this specification always use division by n, not n minus 1.
    • ๐Ÿ’กWhen comparing two data sets, quote both standard deviations and state which is more spread out in context.
    • ๐Ÿ’กShow the quartiles and the 1.5 times interquartile range boundaries when identifying outliers.
    • ๐Ÿ’กWhen critiquing a diagram, give one strength and one limitation linked to the data and the question.
    • ๐Ÿ’กFor cleaning data, state clearly what you did with each problem value and the effect on your conclusion.
    Common Mistakes
    • Treating bar height as frequency when class widths are unequal; the correction is to multiply height by width to obtain frequency.
    • Forgetting to divide frequency by class width when drawing or completing a histogram; the correction is to plot frequency density on the vertical axis.
    • Assuming a histogram's total area equals the total frequency; the correction is that total area equals total frequency only when frequency density is used, whereas relative frequency density gives total area 1.
    • Describing a correlation as proof that one variable causes the other; the correction is to say the association may be explained by a third variable or by chance.
    • Ignoring separate clusters and fitting one description to all points; the correction is to identify the distinct sections and describe each one.
    • Predicting far outside the range of the observed data; the correction is to restrict predictions to values within the data range.
    • Dividing by n minus 1 instead of n when calculating standard deviation; correct this by always dividing by n for standard deviation in this specification.
    • Forgetting to take the square root at the end of the calculation; the correction is to take the square root of the variance to obtain the standard deviation.
    • Mixing up sum of x squared with the square of sum of x; the correction is to keep the two summary quantities distinct and use the correct one in the formula.
    • Removing every outlier automatically without checking whether it is genuine; the correction is to interpret each value in context before deciding.
    • Using the mean and range to identify outliers when the data is skewed; the correction is to use quartiles and the interquartile range, which are resistant to extreme values.
    • Ignoring missing data and treating the remaining values as complete; the correction is to state how many values are missing and how they were handled.