AP® Computer Science Principles review sheet from Aim for Five (aimforfive.com/csp/units/2/2-2)
Unit 2 · Topic 2.2
2.2 Data Compression
Compression shrinks the number of bits needed to store or send data. This topic compares lossless compression, which rebuilds the original exactly, with lossy compression, which gives up some detail for much smaller files, and how to choose between them.
Key terms
- data compression
- lossless compression
- lossy compression
- redundancy
- trade-off
What compression does
Data compression reduces the number of bits used to store or transmit data. Smaller files take less storage space and travel faster over a network.
Fewer bits doesn't always mean less information. A compressed file can hold exactly the same information as the original, just written more efficiently.
How much a file shrinks depends on two things: how much redundancy the data has (repeated or predictable patterns) and which compression algorithm is used. A picture that's mostly blue sky compresses much more than one full of random static.
Lossless compression
Lossless compression reduces the number of bits while guaranteeing that the original data can be rebuilt exactly, bit for bit. Nothing is thrown away.
A simple example is run-length encoding, which replaces a run of repeated values with a count and the value. A row of 20 pixels in a black-and-white image, WWWWWWBBBWWWWWWWWWWB, can be stored as 6W 3B 10W 1B. Anyone with the rule can turn that back into the exact same 20 pixels.
Another lossless idea is to replace common, repeated chunks (like a word that appears many times in a document) with a short code, and keep a table of what each code stands for.
Lossless is used for text, program files, spreadsheets, medical scans and any data where one wrong bit could matter. ZIP and PNG files use lossless compression.
Lossy compression
Lossy compression removes some of the data for good, usually detail that people are unlikely to notice, like tiny color differences in a photo or sounds outside the range most people can hear. You can only rebuild an approximation of the original.
In return, lossy compression can usually shrink files much more than lossless can. That's why streaming video, JPEG photos and MP3 or AAC audio use it.
Compress with a lossy method too aggressively and the loss becomes visible or audible: blocky photos, blurry video, tinny music.
Choosing the right method
It's a trade-off between size and quality. Ask what matters most in the situation:
| Situation | Better choice | Why |
|---|---|---|
| Sending a legal contract or source code | Lossless | Every character must survive exactly |
| Archiving original photos for a museum | Lossless | Quality and exact reconstruction matter most |
| Live video calls on a slow connection | Lossy | Small size and speed matter more than perfect detail |
| Fitting thousands of songs on a phone | Lossy | Storage is limited and small losses are hard to hear |
Worked examples
Try each one yourself first, then open the solution.
- Example 1
When run-length encoding helps and when it hurts
A run-length encoding stores each run of identical pixels as a count and a color. Row A is WWWWWWBBBWWWWWWWWWWB (20 pixels). Row B is WBWBWBWBWB (10 pixels). For each row, how many count-and-color pairs does the encoding need, and does it make the row shorter?
Show the solutionHide the solution
- Step 1: Row A has four runs: 6 W, 3 B, 10 W, 1 B. That's 4 pairs, or 8 values, instead of 20. It's shorter because the row has lots of repetition.
- Step 2: Row B never repeats a color twice in a row, so every pixel is its own run: 1W 1B 1W 1B ... That's 10 pairs, or 20 values, for 10 pixels.
- Step 3: For Row B the "compressed" version is twice as long as the original. Run-length encoding only helps when the data has redundancy in the form of long runs.
Answer: Row A: 4 pairs, much shorter. Row B: 10 pairs, twice as long. Compression depends on the redundancy in the data.
- Example 2
Picking lossless or lossy
A hospital wants to send X-ray images to a specialist across the country. A video app wants to let users post short clips quickly over cell networks. Which kind of compression should each use?
Show the solutionHide the solution
- Step 1: For the X-rays, a doctor may need to see tiny details, and the original must be recoverable exactly. That points to lossless.
- Step 2: For the video clips, upload speed and small file size matter more, and viewers won't notice small losses in detail. That points to lossy.
Answer: The hospital should use lossless compression; the video app should use lossy compression.
Common mistakes
- Saying lossy compression is "worse." It's the right choice when size or speed matters more than perfect quality.
- Thinking lossless compression always makes a file much smaller. If the data has little redundancy, it may barely shrink, or even grow.
- Believing a lossy file can be uncompressed back to the original. The removed detail is gone for good.
On the exam
- Most questions give a scenario and ask which method fits, or ask what's true of a method. Look for words like "exact," "original" and "reconstruct" (lossless) versus "smallest size" or "fastest transmission" (lossy).
- You may be shown a small made-up compression scheme and asked to decode it, or to say whether it's lossless. If you can always get the exact original back, it's lossless.
Connected topics
Videos
Check yourself
4 questions on 2.2 Data Compression. Pick an answer to see if you got it, and why.
In which of the following situations is lossless compression the best choice?
A video streaming service wants to send video to viewers with slow internet connections. Which approach best fits this goal?
Which statement correctly compares lossless and lossy compression?
A simple compression method replaces each run of repeated characters with the number of times it repeats, followed by the character. For example, AAAB becomes 3A1B. What is the result of compressing WWWWWWBBBWWWWBB?
0 of 4 answered