Using genome projects (A-level only)

    AQA
    A-Level

    The genome is the complete set of genes in a cell, or all the DNA in a cell, including non-coding DNA. Sequencing determines the order of bases along that DNA. The Human Genome Project mapped the vast majority of the human genome between 1990 and 2003. Genomes of bacteria, viruses, and model organisms like the fruit fly have also been sequenced, many concurrently or earlier. Comparing sequences between species measures how closely related organisms are; more similar sequences imply a more recent common ancestor. This evidence has led to species originally grouped by appearance being reclassified and renamed based on their evolutionary relationships.

    15
    Objectives
    14
    Exam Tips
    22
    Pitfalls
    30
    Key Terms
    24
    Mark Points

    Subtopics in this area

    Sequencing projects have read the genomes of a wide range of organisms, including humans.
    Determining the genome of simpler organisms allows the sequences of the proteins that derive from the genetic code (the proteome) of the organism to be determined.
    This may have many applications, including the identification of potential antigens for use in vaccine production.
    In more complex organisms, the presence of non-coding DNA and of regulatory genes means that knowledge of the genome cannot easily be translated into the proteome.
    Sequencing methods are continuously updated and have become automated.

    Using genome projects (A-level only) Revision Guide

    Learning Objectives

    What you need to know and understand

    • Define the genome precisely enough to avoid the species-level wording that mark schemes reject.
    • Explain how comparing base sequences between two organisms provides evidence about their evolutionary relationship.
    • Suggest why DNA sequencing has led to bacterial species being renamed when microscopy had grouped them together.
    • Define the proteome in wording that survives the mark scheme's restrictions on 'number of proteins' and on species-level answers.
    • Explain why the proteome of a bacterium can be predicted from its genome but that of a mammal cannot.
    • Suggest how knowing a pathogen's proteome leads to the identification of a candidate antigen for a vaccine.
    • Explain how a sequenced genome leads to a shortlist of candidate antigens for a vaccine.
    • Describe the immune response that follows injection of an antigen identified in this way, naming memory cells.
    • Suggest an advantage of finding antigens from sequence data rather than from cultures of the pathogen.
    • Explain why the base sequence of a eukaryotic gene cannot be translated directly into an amino acid sequence.
    • Describe how alternative splicing lets one gene code for more than one polypeptide.
    • Suggest why two cells from the same person, with identical genomes, contain different sets of proteins.
    • Describe how an automated sequencer handles a genome too long to read in one piece.
    • Explain two consequences of faster, cheaper sequencing for biology or medicine.
    • Suggest why species classified before automated sequencing have since been renamed.

    Marking Points

    Key points examiners look for in your answers

    • one mark for defining the genome as the complete set of genes in a cell, or all the DNA in a cell or organism
    • one mark for sequencing being the determination of the order of bases in the DNA
    • one mark for using base sequence comparison between species as evidence of how closely related they are
    • one mark for a named consequence of sequencing, such as species being reclassified or renamed on new sequence evidence
    • one mark for a named use of a sequenced genome within a species, such as identifying alleles linked to disease risk
    • The proteome is defined as the full range of proteins that a cell, genome or organism is able to produce.
    • In simple organisms like prokaryotes, the genome can be used to directly determine the proteome because their DNA does not contain introns.
    • The absence of non-coding DNA means the base sequence can be read in triplets using the genetic code to determine the exact amino acid sequence (primary structure).
    • Determining the proteome of pathogenic bacteria allows for the identification of surface proteins that can act as antigens for vaccine development.
    • one mark for sequencing the pathogen's genome and using it to determine the proteome
    • one mark for identifying proteins found on the surface of the pathogen as potential antigens
    • one mark for explaining that these antigens trigger an immune response, producing antibodies and memory cells
    • one mark for a stated advantage, such as identifying antigens without having to culture a dangerous pathogen, or doing so more quickly
    • one mark for producing the antigen in quantity using recombinant DNA technology before it is used in a vaccine
    • one mark for eukaryotic genes containing introns or non-coding DNA that is not translated
    • one mark for introns being removed from pre-mRNA by splicing, so the coding regions cannot be identified from the DNA sequence alone
    • one mark for one gene producing more than one polypeptide because pre-mRNA can be spliced in different combinations
    • one mark for regulatory genes or transcription factors determining which genes are expressed in a particular cell
    • one mark for epigenetic modification altering expression without altering the base sequence
    • one mark for sequencing now being automated, carried out by machines rather than by hand
    • one mark for the DNA being cut into many fragments that are sequenced at the same time and reassembled by computer
    • one mark for the consequence being that sequencing is faster and cheaper, with fewer errors
    • one mark for an application made possible, such as DNA or base sequencing evidence leading to species being reclassified or renamed
    • one mark for whole genomes of individuals being sequenced, allowing screening and personalised medicine

    Examiner Tips

    Expert advice for maximising your marks

    • 💡The definition of a genome is specific: the complete set of genes in a cell, or all the DNA in a cell.
    • 💡Say 'cell' or 'organism', never 'species' or 'population', in a genome definition.
    • 💡If asked why classification has changed, refer to specific techniques like DNA or RNA base sequencing providing new evidence.
    • 💡Always pair the definitions in your notes: the genome is the complete set of genes in a cell, whereas the proteome is the full range of proteins that the cell can produce.
    • 💡If asked why a bacterial proteome is easier to determine from its genome than a eukaryotic one, explicitly state the absence of introns and non-coding DNA.
    • 💡Keep antigen and antibody straight: the antigen is the pathogen's protein, the antibody is what the body makes against it.
    • 💡Write the chain in order - genome, proteome, surface protein, antigen, vaccine, memory cells.
    • 💡When asked for an advantage, make it specific: no need to culture the pathogen, or many candidates screened at once.
    • 💡Two separate ideas earn marks here - non-coding DNA and splicing, and regulation of expression - so give one of each.
    • 💡Say 'pre-mRNA' when you mention splicing; saying introns are cut from DNA loses the mark.
    • 💡Use the liver cell and neurone contrast to show that one genome supports many proteomes.
    • 💡Advantages need a comparison - faster than what, cheaper than what - so name the old method or the old timescale.
    • 💡Link automation to a biological consequence; the mark is rarely for the technology on its own.
    • 💡If a question asks why a classification has been revised, name the evidence precisely - base sequencing scores, and so does electron microscopy of greater resolution or improved staining, so do not cross a microscopy answer out.

    Common Mistakes

    Pitfalls to avoid in your exam answers

    • defining the genome as all the DNA in a species or a population, rather than per cell or per organism
    • defining the genome as the proteins an organism can produce, which confuses it with the proteome, or the expressed genes, which is the transcriptome
    • saying sequencing counts the genes, rather than reading the order of the bases
    • claiming the Human Genome Project identified all human proteins, when a genome does not give the proteome directly in a eukaryote
    • Defining the proteome simply as 'the number of proteins a cell produces'. Correction: State that it is the 'full range of proteins' a cell or organism can produce, as 'number' alone is insufficient.
    • Stating that the proteome is the proteins a cell is making at a specific moment. Correction: The proteome encompasses all the proteins a cell can produce, not just those currently being expressed.
    • Claiming that bacterial genomes must undergo splicing before the proteome can be determined. Correction: Remember that prokaryotic DNA does not contain introns, so no splicing occurs and the sequence translates directly.
    • saying the vaccine contains the pathogen's gene, rather than the antigen the gene codes for
    • saying the vaccine contains antibodies, which would describe passive immunity instead
    • writing that the vaccine 'kills the pathogen' with no account of the immune response it provokes
    • assuming any protein in the proteome would work as an antigen, when it must be exposed on the surface
    • leaving out memory cells, so the answer never explains long-term protection
    • saying that eukaryotic DNA is entirely non-coding, rather than mostly non-coding
    • saying introns are removed from the DNA, when they are removed from the pre-mRNA
    • assuming the number of genes equals the number of different proteins
    • confusing non-coding DNA with genes that happen not to be switched on in that cell
    • forgetting regulatory genes, so the answer only ever mentions introns
    • writing that modern methods sequence the protein, when sequencing reads the order of DNA bases
    • answering only 'it is automated' when the question asks for an advantage, so never stating faster, cheaper or more accurate
    • claiming a machine reads a whole chromosome in one pass, with no fragmentation or reassembly step
    • confusing DNA sequencing with genetic fingerprinting, which compares fragment lengths and never reads bases
    • saying improved technology changed the genetic code rather than our ability to read it