Skip to topic
    ← Back to course topics

    B1b — AQA GCSE Statistics

    Test yourself on B1b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Notes: students may be required to make decisions about appropriate class intervals given a data set.

    B1b exam tips

    Quick Revision Summary (Key Takeaway)

    B1b in AQA GCSE Statistics covers the collection, organisation and representation of data, including sampling methods, data types, and the construction and interpretation of tables, charts and graphs. Mastery of this unit is essential because every statistical analysis depends on accurate data collection and clear, appropriate data presentation.

    Topic Overview

    B1b is the second part of the data handling section in AQA GCSE Statistics. It focuses on how data is collected, organised and represented. You will learn about different sampling methods, the distinction between primary and secondary data, and how to choose and construct appropriate tables, charts and graphs for different types of data. This topic is fundamental because the quality of any statistical conclusion depends on the quality of the data collection and presentation.

    In the wider subject, B1b connects directly to later topics such as measures of central tendency and spread (B2), correlation and regression (B3), and probability (B4). Being able to collect and represent data accurately is a prerequisite for analysing it correctly. In the exam, this topic is assessed through questions that require you to describe sampling methods, criticise data collection techniques, and interpret or construct statistical diagrams.

    Key Concepts
    • →Primary data is collected first-hand for a specific purpose; secondary data is collected by someone else for a different purpose. Each has advantages and disadvantages depending on context.
    • →Sampling methods include simple random, systematic, stratified and quota sampling. Each method has a defined procedure and particular strengths and weaknesses.
    • →Data types include qualitative (categorical) and quantitative (discrete and continuous). The data type determines the appropriate graph or chart.
    • →Graphical representations include bar charts, pie charts, histograms, line graphs, scatter graphs and cumulative frequency graphs. Each is suited to specific data types and purposes.
    • →A histogram is used for continuous data with unequal class widths; the area of each bar, not the height, represents frequency.
    Examiner Tips
    • 💡When asked to describe a sampling method, always refer to the specific context. For example, say 'number the students 1 to 200 and use a random number generator to select 20' rather than 'pick students randomly'.
    • 💡When justifying a choice of graph, explicitly state the data type and why the graph is suitable. For example, 'A pie chart is appropriate because the data is categorical and shows proportions of a whole'.
    • 💡In questions about data collection, always consider possible sources of bias and how they could be reduced. Mentioning specific improvements, such as using a larger sample or random selection, will gain credit.
    Common Mistakes
    • Students often think a histogram is just a bar chart with no gaps. In fact, in a histogram the area of each bar represents frequency, and the vertical axis is frequency density, not frequency. This is crucial when class widths are unequal.
    • Students frequently confuse stratified sampling with quota sampling. Stratified sampling involves random selection within each stratum, while quota sampling involves selecting a specified number from each group non-randomly. Stratified sampling is more representative but requires a sampling frame.
    • Many students believe that a larger sample is always better. While larger samples generally reduce sampling error, a biased sampling method will produce unreliable results regardless of sample size. Representativeness is more important than size alone.
    Revision Plan
    1. 1Start by learning the definitions and differences between primary and secondary data, and the various sampling methods. Create a table summarising the advantages and disadvantages of each.
    2. 2Practice identifying data types and choosing appropriate graphs for each. Use past paper questions to test your ability to justify your choices.
    3. 3Learn how to construct and interpret histograms with unequal class widths. Focus on calculating frequency density and understanding that area represents frequency.
    4. 4Work through exam-style questions on sampling, including stratified sampling calculations. Check your answers against mark schemes to understand the level of detail required.
    5. 5Review common misconceptions and examiner tips, then complete a timed practice paper on B1b to consolidate your understanding.
    Exam Question Types
    • 📋Describe how to carry out a particular sampling method in context. Advice: be specific about the procedure, including how to select individuals and the sampling interval if applicable.
    • 📋Criticise a given data collection method or questionnaire. Advice: identify sources of bias, leading questions, or inappropriate sampling, and suggest improvements.
    • 📋Construct or complete a statistical diagram such as a histogram or pie chart. Advice: check scales, labels and accuracy; for histograms, calculate frequency density correctly.
    • 📋Interpret a graph or chart and comment on the trend or distribution. Advice: describe the overall pattern, any outliers, and relate back to the context.
    Command Word Expectations (AQA)
    Describe

    Give a detailed account of the steps or features. In AQA GCSE Statistics, this means stating what you would do in a logical order, often with specific reference to the context. For example, 'Describe how to take a stratified sample' requires you to outline the process of dividing the population into strata, calculating the number from each stratum, and randomly selecting within each.

    Explain

    Give reasons or justify why something is the case. This requires a cause-and-effect statement. For example, 'Explain why a histogram is more appropriate than a bar chart for this data' requires you to state that the data is continuous and grouped, and that the area represents frequency.

    Compare

    Identify similarities and differences between two things. For example, 'Compare primary and secondary data' requires you to state both similarities (both are sources of data) and differences (who collected it, purpose, reliability, cost).

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Confusing primary and secondary data, or failing to justify why one is more appropriate in context.
    ❌ Weak Answer (Loses Marks):Primary data is data you collect yourself and secondary data is data someone else collected. Primary data is better.
    Example improved answer:Primary data is collected first-hand by the researcher for the specific purpose of the investigation, whereas secondary data has been collected by someone else for a different purpose. Primary data is more reliable and tailored to the question but is time-consuming and expensive to collect. Secondary data is quicker and cheaper but may be outdated, biased or not exactly suited to the investigation. The choice depends on the context: for a school survey about lunch preferences, primary data is appropriate because the school needs current, specific information about its own students.
    Examiner Tip: Always link your justification to the specific context given in the question. Do not just state generic advantages and disadvantages; explain why one type is more suitable for that particular investigation.
    Pitfall: Choosing an inappropriate graph type for the data, such as a pie chart for continuous data or a histogram for discrete data.
    ❌ Weak Answer (Loses Marks):I would draw a bar chart because it shows the data clearly.
    Example improved answer:For discrete categorical data such as favourite subject, a bar chart or pie chart is appropriate because it compares frequencies across categories. For continuous data grouped into classes, a histogram should be used because the bars are drawn with no gaps and the area of each bar represents frequency. For time series data, a line graph is appropriate to show trends over time. The choice of graph must match the data type and the purpose of the representation.
    Examiner Tip: Before selecting a graph, identify whether the data is discrete, continuous, categorical or time-related. State the data type and then justify your graph choice in terms of what it shows effectively.
    Step-by-Step Worked Solutions

    Question: A student wants to investigate the average number of hours of sleep per night for students in her year group of 200 students. She decides to use a systematic sample of 20 students. Explain how she could carry out this sample and give one advantage and one disadvantage of using a systematic sample in this context.

    1. 1.Step 1: Identify the population size (200) and required sample size (20). Calculate the sampling interval by dividing the population size by the sample size: 200 / 20 = 10.
    2. 2.Step 2: Choose a random starting point between 1 and 10. For example, use a random number generator to select 7.
    3. 3.Step 3: Select every 10th student from the list, starting with the 7th student. So the sample would be students 7, 17, 27, 37, ... up to 197.
    4. 4.Step 4: State one advantage: systematic sampling is quick and easy to perform, and it ensures the sample is spread evenly across the population list.
    5. 5.Step 5: State one disadvantage: if the list has a repeating pattern that coincides with the sampling interval, the sample may be biased. For example, if every 10th student is absent on a particular day, the sample may not be representative.
    Final Answer: Use a sampling interval of 10, randomly select a starting point between 1 and 10, then select every 10th student. Advantage: quick and evenly spread. Disadvantage: may be biased if the list has a periodic pattern.

    Question: The table below shows the number of cars sold by a dealership over five months. Draw a suitable graph to represent the data and describe the trend. Month: Jan, Feb, Mar, Apr, May. Cars sold: 24, 30, 45, 38, 50.

    1. 1.Step 1: Identify the data type: time series data (months) with discrete numerical values (number of cars).
    2. 2.Step 2: Choose an appropriate graph: a line graph is suitable for showing a trend over time.
    3. 3.Step 3: Plot the months on the horizontal axis and the number of cars sold on the vertical axis. Use a sensible scale, for example 0 to 60 in increments of 10.
    4. 4.Step 4: Plot each point accurately and join them with straight line segments.
    5. 5.Step 5: Describe the trend: overall, car sales increased from January to May, with a peak in May (50 cars) and a slight dip in April (38 cars).
    Final Answer: A line graph with months on the x-axis and cars sold on the y-axis, showing an overall upward trend from 24 in January to 50 in May, with a small decrease in April.
    Active Recall Memory Test
    What is the difference between primary and secondary data?
    Key Fact: Primary data is collected first-hand by the researcher for the specific purpose of the investigation. Secondary data is collected by someone else for a different purpose and then used by the researcher.
    When should you use a histogram instead of a bar chart?
    Key Fact: Use a histogram for continuous data that is grouped into classes, especially when class widths are unequal. Use a bar chart for discrete or categorical data.
    What is the formula for calculating the number from each stratum in a stratified sample?
    Key Fact: Number from stratum = (stratum size / total population) * sample size.
    Give one advantage and one disadvantage of using secondary data.
    Key Fact: Advantage: it is often quicker and cheaper to obtain. Disadvantage: it may be outdated, biased, or not exactly suited to the investigation.
    Frequently Asked Questions
    What is the difference between stratified sampling and quota sampling?
    Stratified sampling involves dividing the population into strata (groups) based on a characteristic, then randomly selecting individuals from each stratum in proportion to its size. Quota sampling also divides the population into groups, but the researcher selects a specified number of individuals from each group non-randomly, often based on convenience. Stratified sampling is more representative because of the random selection, but it requires a sampling frame. Quota sampling is quicker but can introduce researcher bias.
    How do I know which graph to use for my data?
    The choice of graph depends on the data type and what you want to show. For categorical data, use a bar chart or pie chart. For discrete numerical data, a bar chart or line graph can be used. For continuous data, use a histogram or cumulative frequency graph. For time series data, use a line graph. For bivariate data (pairs of values), use a scatter graph. Always consider whether you are comparing categories, showing a trend over time, or showing the distribution of continuous data.
    What is frequency density and why is it used in histograms?
    Frequency density is calculated as frequency divided by class width. It is used in histograms because when class widths are unequal, the height of the bar cannot directly represent frequency. Instead, the area of each bar represents frequency. By plotting frequency density on the vertical axis, the area of each bar (frequency density * class width) equals the frequency. This allows for a fair comparison of frequencies across classes of different widths.
    What are the advantages and disadvantages of using a simple random sample?
    A simple random sample is unbiased because every member of the population has an equal chance of being selected. It is also easy to understand and explain. However, it requires a complete sampling frame (a list of the entire population), which may not be available. It can also be time-consuming and expensive to implement if the population is large and spread out. Additionally, it may not be representative if the sample size is small, as it could by chance under-represent certain groups.
    How do I avoid bias when designing a questionnaire?
    To avoid bias in a questionnaire, ensure questions are clear, concise and not leading. Avoid questions that suggest a particular answer, such as 'Don't you think that...?' Use neutral wording and provide balanced response options. Avoid double-barrelled questions that ask about two things at once. Also, consider the order of questions, as earlier questions can influence later responses. Pilot the questionnaire with a small group to identify any issues before distributing it widely.
    What is the difference between discrete and continuous data?
    Discrete data can only take specific values, often counts of items, such as the number of students in a class. It is usually represented by integers. Continuous data can take any value within a range, such as height, weight or time, and is often measured to a certain degree of accuracy. Continuous data is typically grouped into classes for representation in histograms or cumulative frequency graphs, while discrete data is often represented in bar charts or pie charts.