Skip to main content

Unit 2 · Topic 2.2

2.2 Data Compression

Compression shrinks the number of bits needed to store or send data. This topic compares lossless compression, which rebuilds the original exactly, with lossy compression, which gives up some detail for much smaller files, and how to choose between them.

Key terms

  • data compression
  • lossless compression
  • lossy compression
  • redundancy
  • trade-off

What compression does

Data compression reduces the number of bits used to store or transmit data. Smaller files take less storage space and travel faster over a network.

Fewer bits doesn't always mean less information. A compressed file can hold exactly the same information as the original, just written more efficiently.

How much a file shrinks depends on two things: how much redundancy the data has (repeated or predictable patterns) and which compression algorithm is used. A picture that's mostly blue sky compresses much more than one full of random static.

Lossless compression

Lossless compression reduces the number of bits while guaranteeing that the original data can be rebuilt exactly, bit for bit. Nothing is thrown away.

A simple example is run-length encoding, which replaces a run of repeated values with a count and the value. A row of 20 pixels in a black-and-white image, WWWWWWBBBWWWWWWWWWWB, can be stored as 6W 3B 10W 1B. Anyone with the rule can turn that back into the exact same 20 pixels.

Another lossless idea is to replace common, repeated chunks (like a word that appears many times in a document) with a short code, and keep a table of what each code stands for.

Lossless is used for text, program files, spreadsheets, medical scans and any data where one wrong bit could matter. ZIP and PNG files use lossless compression.

Lossy compression

Lossy compression removes some of the data for good, usually detail that people are unlikely to notice, like tiny color differences in a photo or sounds outside the range most people can hear. You can only rebuild an approximation of the original.

In return, lossy compression can usually shrink files much more than lossless can. That's why streaming video, JPEG photos and MP3 or AAC audio use it.

Compress with a lossy method too aggressively and the loss becomes visible or audible: blocky photos, blurry video, tinny music.

Choosing the right method

It's a trade-off between size and quality. Ask what matters most in the situation:

SituationBetter choiceWhy
Sending a legal contract or source codeLosslessEvery character must survive exactly
Archiving original photos for a museumLosslessQuality and exact reconstruction matter most
Live video calls on a slow connectionLossySmall size and speed matter more than perfect detail
Fitting thousands of songs on a phoneLossyStorage is limited and small losses are hard to hear

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1

    When run-length encoding helps and when it hurts

    A run-length encoding stores each run of identical pixels as a count and a color. Row A is WWWWWWBBBWWWWWWWWWWB (20 pixels). Row B is WBWBWBWBWB (10 pixels). For each row, how many count-and-color pairs does the encoding need, and does it make the row shorter?

    Show the solution
    1. Step 1: Row A has four runs: 6 W, 3 B, 10 W, 1 B. That's 4 pairs, or 8 values, instead of 20. It's shorter because the row has lots of repetition.
    2. Step 2: Row B never repeats a color twice in a row, so every pixel is its own run: 1W 1B 1W 1B ... That's 10 pairs, or 20 values, for 10 pixels.
    3. Step 3: For Row B the "compressed" version is twice as long as the original. Run-length encoding only helps when the data has redundancy in the form of long runs.

    Answer: Row A: 4 pairs, much shorter. Row B: 10 pairs, twice as long. Compression depends on the redundancy in the data.

  2. Example 2

    Picking lossless or lossy

    A hospital wants to send X-ray images to a specialist across the country. A video app wants to let users post short clips quickly over cell networks. Which kind of compression should each use?

    Show the solution
    1. Step 1: For the X-rays, a doctor may need to see tiny details, and the original must be recoverable exactly. That points to lossless.
    2. Step 2: For the video clips, upload speed and small file size matter more, and viewers won't notice small losses in detail. That points to lossy.

    Answer: The hospital should use lossless compression; the video app should use lossy compression.

Common mistakes

  • Saying lossy compression is "worse." It's the right choice when size or speed matters more than perfect quality.
  • Thinking lossless compression always makes a file much smaller. If the data has little redundancy, it may barely shrink, or even grow.
  • Believing a lossy file can be uncompressed back to the original. The removed detail is gone for good.

On the exam

  • Most questions give a scenario and ask which method fits, or ask what's true of a method. Look for words like "exact," "original" and "reconstruct" (lossless) versus "smallest size" or "fastest transmission" (lossy).
  • You may be shown a small made-up compression scheme and asked to decode it, or to say whether it's lossless. If you can always get the exact original back, it's lossless.

Connected topics

Videos

  • AP CS Principles Exam Review - Compression

    Flavio KupermanWatch on YouTube (opens in a new tab)

  • Text compression widget with Aloe Blacc

    CodeAIWatch on YouTube (opens in a new tab)

  • Compression, Lossy & Lossless (AP Computer Science Principles Unit 1: Digital Information)

    Professor CunninghamWatch on YouTube (opens in a new tab)

  • Compression: Crash Course Computer Science #21

    CrashCourseWatch on YouTube (opens in a new tab)

  • Lossy compression! Topic 2.2 and Code.org Unit 1.10 walkthrough. 8 mini-practice questions!

    Dr_WuWatch on YouTube (opens in a new tab)

  • Data Compression

    CodeHSWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 2.2 Data Compression. Pick an answer to see if you got it, and why.

Question 1 of 4

In which of the following situations is lossless compression the best choice?

Question 2 of 4

A video streaming service wants to send video to viewers with slow internet connections. Which approach best fits this goal?

Question 3 of 4

Which statement correctly compares lossless and lossy compression?

Question 4 of 4

A simple compression method replaces each run of repeated characters with the number of times it repeats, followed by the character. For example, AAAB becomes 3A1B. What is the result of compressing WWWWWWBBBWWWWBB?

0 of 4 answered