Skip to topic
    ← Back to course topics

    C1b — AQA GCSE Statistics

    Test yourself on C1b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. population pyramid

    C1b exam tips

    Quick Revision Summary (Key Takeaway)

    C1b in AQA GCSE Statistics covers the collection, organisation, and representation of data, including sampling methods, types of data, and graphical techniques such as histograms and cumulative frequency diagrams. It is a foundational topic that underpins all subsequent statistical analysis and interpretation.

    Topic Overview

    C1b is a core component of the AQA GCSE Statistics specification, focusing on the methods used to collect and represent data accurately. It covers sampling techniques, types of data, and graphical representations such as histograms, cumulative frequency diagrams, and box plots. Understanding these concepts is essential for conducting valid statistical investigations and interpreting data correctly.

    This topic provides the foundation for later statistical analysis, including measures of central tendency and dispersion, correlation, and probability. Mastery of C1b ensures students can critically evaluate data sources, choose appropriate sampling methods, and present data in ways that reveal patterns and trends. It is a key skill for both exams and real-world data handling.

    Key Concepts
    • →Sampling methods: simple random, systematic, stratified, and quota sampling, each with advantages and disadvantages.
    • →Types of data: qualitative vs quantitative, discrete vs continuous, and primary vs secondary data.
    • →Graphical representation: histograms (with unequal class widths), cumulative frequency diagrams, box plots, and scatter diagrams.
    • →Measures of central tendency and spread: mean, median, mode, range, interquartile range, and standard deviation.
    • →Interpretation of graphs: identifying skewness, outliers, and comparing distributions.
    Examiner Tips
    • 💡Always define the sampling frame and explain why your chosen sampling method is suitable for the context; this shows understanding beyond rote learning.
    • 💡When drawing graphs, label axes clearly with units and a title; examiners award marks for correct labelling and scaling.
    • 💡In interpretation questions, refer to specific data values and use comparative language (e.g., 'the median for group A is higher than group B, indicating...').
    Common Mistakes
    • Students often think that a larger sample is always better, but a well-designed smaller sample can be more accurate than a large biased sample.
    • Many confuse frequency with frequency density in histograms, leading to incorrect area calculations and misinterpretation of the data.
    • Students sometimes believe that correlation implies causation, which is not true; correlation only indicates a relationship, not a cause-and-effect link.
    Revision Plan
    1. 1Day 1-2: Review sampling methods and types of data; create a summary table of advantages and disadvantages for each method.
    2. 2Day 3-4: Practice drawing histograms with unequal class widths and cumulative frequency diagrams; focus on calculating frequency density and plotting points accurately.
    3. 3Day 5-6: Work through exam-style questions on interpreting graphs, including comparing distributions and identifying outliers.
    4. 4Day 7-8: Complete a past paper section on C1b under timed conditions; mark your answers and note common errors.
    5. 5Day 9-10: Revise weak areas using flashcards and online quizzes; focus on command words and mark scheme requirements.
    Exam Question Types
    • 📋Multiple-choice questions on sampling methods and data types: read each option carefully and eliminate obviously wrong answers.
    • 📋Calculation questions on stratified sampling: show your working clearly, including the sampling fraction and each stratum calculation.
    • 📋Graph drawing and interpretation: ensure axes are labelled, scales are consistent, and you refer to specific data points in your answers.
    • 📋Comparison questions: use comparative language and quote statistics (e.g., medians, IQRs) to support your conclusions.
    Command Word Expectations (AQA)
    Describe

    Give a detailed account of the main features; for example, describe the shape of a distribution by mentioning skewness, outliers, and the median.

    Explain

    Provide reasons or justifications; for example, explain why a stratified sample is more representative than a simple random sample, linking to the population structure.

    Compare

    Identify similarities and differences, using comparative language and specific data values; for example, compare the interquartile ranges of two box plots.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often confuse the different sampling methods and fail to explain why a particular method is appropriate for a given scenario.
    ❌ Weak Answer (Loses Marks):A student might say 'systematic sampling is when you pick people randomly' without specifying the ordered list and regular interval.
    Example improved answer:Systematic sampling involves selecting every nth item from an ordered sampling frame, starting at a random point within the first n items. It is appropriate when the population is large and a sampling frame is available, as it is quick and ensures even coverage.
    Examiner Tip: Always link the sampling method to the context: mention the sampling frame, the interval, and why it reduces bias or is practical.
    Pitfall: When drawing histograms, students frequently use equal class widths and treat frequency as the height, ignoring frequency density.
    ❌ Weak Answer (Loses Marks):A student might plot frequency on the y-axis for a histogram with unequal class widths, leading to incorrect area representation.
    Example improved answer:For a histogram with unequal class widths, the vertical axis must be frequency density, calculated as frequency divided by class width. The area of each bar represents the frequency, so the height must be adjusted accordingly.
    Examiner Tip: Always check if class widths are equal. If not, calculate frequency density and label the axis clearly as 'Frequency density'.
    Step-by-Step Worked Solutions

    Question: A school has 1200 students. A researcher wants to select a stratified sample of 60 students by year group. The numbers in each year group are: Year 7: 200, Year 8: 220, Year 9: 240, Year 10: 260, Year 11: 280. Calculate the number of students to sample from each year group.

    1. 1.Step 1: Identify the total population (1200) and sample size (60). Calculate the sampling fraction: 60/1200 = 1/20.
    2. 2.Step 2: Apply the sampling fraction to each stratum: Year 7: 200 × (1/20) = 10; Year 8: 220 × (1/20) = 11; Year 9: 240 × (1/20) = 12; Year 10: 260 × (1/20) = 13; Year 11: 280 × (1/20) = 14.
    3. 3.Step 3: Check that the sum equals 60: 10+11+12+13+14 = 60. State the final sample sizes for each year group.
    Final Answer: Year 7: 10, Year 8: 11, Year 9: 12, Year 10: 13, Year 11: 14.

    Question: The table shows the speeds of 80 cars on a road, with class intervals: 0-20 (frequency 10), 20-30 (frequency 20), 30-40 (frequency 25), 40-60 (frequency 15), 60-80 (frequency 10). Draw a histogram and estimate the number of cars travelling between 25 and 45 mph.

    1. 1.Step 1: Calculate frequency density for each class: 0-20: 10/20=0.5; 20-30: 20/10=2; 30-40: 25/10=2.5; 40-60: 15/20=0.75; 60-80: 10/20=0.5.
    2. 2.Step 2: Draw the histogram with frequency density on the y-axis and speed on the x-axis, using the calculated heights.
    3. 3.Step 3: To estimate cars between 25 and 45 mph, find the area under the histogram between these speeds. From 25-30: width 5, height 2, area=10. From 30-40: width 10, height 2.5, area=25. From 40-45: width 5, height 0.75, area=3.75. Total area = 10+25+3.75 = 38.75, so approximately 39 cars.
    Final Answer: Approximately 39 cars travelled between 25 and 45 mph.
    Active Recall Memory Test
    What is the difference between a histogram and a bar chart?
    Key Fact: A histogram is used for continuous data with no gaps between bars, and the area of each bar represents frequency. A bar chart is for discrete or categorical data with gaps between bars, and the height represents frequency.
    Define stratified sampling and state one advantage.
    Key Fact: Stratified sampling divides the population into homogeneous subgroups (strata) and randomly samples from each. An advantage is that it ensures representation of all subgroups, increasing precision.
    How do you calculate frequency density?
    Key Fact: Frequency density = frequency ÷ class width. It is used in histograms when class widths are unequal.
    What does the interquartile range measure?
    Key Fact: The interquartile range (IQR) measures the spread of the middle 50% of data. It is calculated as upper quartile minus lower quartile and is resistant to outliers.
    Frequently Asked Questions
    What is the difference between primary and secondary data?
    Primary data is collected firsthand by the researcher for a specific purpose, such as conducting a survey. Secondary data is data that already exists, collected by someone else, like government statistics. Primary data is more reliable for the specific research question but can be time-consuming to collect, while secondary data is quicker but may not perfectly match the researcher's needs.
    How do I choose the right sampling method for my investigation?
    The choice depends on the research question, available resources, and population characteristics. Simple random sampling is unbiased but requires a sampling frame. Systematic sampling is quick if a list exists. Stratified sampling ensures representation of subgroups. Quota sampling is quick but non-random and can be biased. Consider practicality and representativeness.
    Why do we use frequency density in histograms?
    Frequency density is used when class widths are unequal so that the area of each bar is proportional to the frequency. If you plotted frequency directly, bars with wider classes would appear larger even if they represent fewer data points, misleading the viewer. Frequency density ensures the visual representation is accurate.
    What is an outlier and how does it affect statistical measures?
    An outlier is a data point that is unusually far from the rest of the data. It can significantly affect the mean and range, pulling them towards its value, but has little effect on the median and interquartile range. Outliers should be investigated to determine if they are errors or genuine extreme values.
    How do I interpret a box plot?
    A box plot displays the median, lower quartile, upper quartile, and extremes (whiskers). The box represents the interquartile range (middle 50% of data). The median line shows the central value. Whiskers extend to the minimum and maximum values (or to a defined range). Comparing box plots helps identify differences in spread and skewness between distributions.
    What are the advantages of using a cumulative frequency diagram?
    A cumulative frequency diagram allows you to estimate the median, quartiles, and percentiles easily. It also shows the shape of the distribution and helps identify outliers. It is particularly useful for large data sets and for comparing distributions.