Skip to topic
    ← Back to course topics

    Compression, Encryption and Hashing — OCR A-Level Computer Science

    Test yourself on Compression, Encryption and Hashing with OCR A-Level practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Compression, Encryption and Hashing explained

    This topic covers the fundamental techniques used for data management and security in computer systems.

    Read the full explanation

    It explores the principles of data compression, the mechanisms of symmetric and asymmetric encryption, and the role of hashing in data integrity.

    What to demonstrate

    1. Distinction between lossy and lossless compression
    2. Application of run length encoding
    3. Application of dictionary coding
    Show all 5 objectives
    1. Understanding of symmetric encryption
    2. Understanding of asymmetric encryption

    Compression, Encryption and Hashing exam tips

    Topic Overview

    Compression, encryption, and hashing are fundamental techniques in computer science for managing data efficiently and securely. Compression reduces file sizes for storage and transmission, encryption ensures data confidentiality by converting plaintext into ciphertext, and hashing provides data integrity by generating fixed-size digests. These concepts are essential for understanding how modern systems handle data securely and efficiently, from streaming services to secure communications.

    In the OCR A-Level specification, you need to understand both lossy and lossless compression algorithms, symmetric and asymmetric encryption, and the properties of hash functions. These topics link to broader areas like network security, data storage, and error detection. Mastery of these concepts is crucial for exam questions that ask you to compare algorithms, explain their applications, or analyse their strengths and weaknesses.

    Real-world applications include JPEG compression for images, AES encryption for secure data transmission, and SHA-256 for verifying file integrity. Understanding these techniques prepares you for further study in cybersecurity, data science, and software engineering, and is directly assessed in Paper 1 and Paper 2 of the OCR A-Level.

    Key Concepts
    • →Lossless vs lossy compression: Lossless (e.g., Run-Length Encoding, Huffman coding) preserves all original data, while lossy (e.g., JPEG, MP3) discards some data to achieve higher compression ratios.
    • →Symmetric encryption uses the same key for encryption and decryption (e.g., AES), while asymmetric encryption uses a public/private key pair (e.g., RSA).
    • →Hash functions (e.g., SHA-256) produce a fixed-size output from any input, are one-way (cannot be reversed), and are collision-resistant (two different inputs should not produce the same hash).
    • →Encryption ensures confidentiality; hashing ensures integrity and authenticity (e.g., password storage, digital signatures).
    • →Compression ratio = original size / compressed size; higher ratio means more compression.
    Marking Points
    • Distinction between lossy and lossless compression
    • Application of run length encoding
    • Application of dictionary coding
    • Understanding of symmetric encryption
    • Understanding of asymmetric encryption
    Examiner Tips
    • 💡Ensure you can clearly define the difference between lossy and lossless compression
    • 💡Be prepared to perform run length encoding on a given string of data
    • 💡Understand the key difference between symmetric and asymmetric encryption regarding the use of keys
    • 💡When comparing compression algorithms, always mention the trade-off between file size and quality (lossy) or computational overhead (lossless). Use specific examples like ZIP vs JPEG.
    • 💡For encryption questions, clearly distinguish between symmetric and asymmetric, and state which is faster (symmetric) and which solves key distribution (asymmetric).
    • 💡In hashing questions, emphasise properties: deterministic, fast to compute, preimage resistant, collision resistant. Use real-world examples like password hashing (with salt) or file integrity checks.
    Common Mistakes
    • Confusing the mechanisms of symmetric and asymmetric encryption
    • Failing to explain the difference between lossy and lossless compression in terms of data recovery
    • Misapplying run length encoding to data that does not contain repeated sequences
    • Misconception: Encryption and hashing are the same. Correction: Encryption is reversible (with the key), while hashing is one-way and cannot be decrypted.
    • Misconception: Lossy compression always results in noticeable quality loss. Correction: Modern lossy algorithms (e.g., JPEG at high quality) can be nearly indistinguishable from the original while significantly reducing file size.
    • Misconception: A hash can be used to encrypt data. Correction: Hashes are not encryption; they are used for verification, not confidentiality.
    Frequently Asked Questions
    What is the difference between lossy and lossless compression?
    Lossless compression reduces file size without losing any data, so the original can be perfectly reconstructed. Examples include ZIP files and PNG images. Lossy compression permanently removes some data to achieve smaller file sizes, often used for media like JPEG images or MP3 audio, where some quality loss is acceptable.
    How does AES encryption work?
    AES (Advanced Encryption Standard) is a symmetric encryption algorithm that encrypts data in blocks of 128 bits using keys of 128, 192, or 256 bits. It uses multiple rounds of substitution, permutation, and mixing to transform plaintext into ciphertext. The same key is used for both encryption and decryption, making it fast and secure for bulk data.
    Why is hashing used for passwords instead of encryption?
    Hashing is one-way, so even if a database is breached, attackers cannot reverse hashes to get original passwords. Encryption is reversible, so if the encryption key is compromised, all passwords are exposed. Additionally, hashing with a salt (random data added to each password) prevents rainbow table attacks.
    What is a collision in hashing and why is it bad?
    A collision occurs when two different inputs produce the same hash output. This is bad because it undermines the integrity guarantee: if two files have the same hash, you cannot trust that they are different. Secure hash functions like SHA-256 are designed to be collision-resistant, making collisions extremely unlikely.
    How does Huffman coding achieve compression?
    Huffman coding is a lossless compression algorithm that assigns variable-length codes to characters based on their frequency. More frequent characters get shorter codes, less frequent get longer codes. It builds a binary tree (Huffman tree) from character frequencies, then traverses it to generate the codes. This reduces the total number of bits needed to represent the data.
    What is the difference between symmetric and asymmetric encryption?
    Symmetric encryption uses the same key for both encryption and decryption, making it fast but requiring secure key exchange. Asymmetric encryption uses a public key to encrypt and a private key to decrypt, solving key distribution but being slower. In practice, hybrid systems use asymmetric encryption to exchange a symmetric key, then use symmetric encryption for the actual data.