1.Lesson overview
- 1.3 Data storage and compression
- 3.3 Data storage and compression
- 3.3 Units of information
- 3.8 Data compression
- Understand how data storage is measured.
- Calculate the file size of an image file and a sound file, using information given.
- Understand the purpose of and need for data compression.
- Understand how files are compressed using lossy and lossless compression methods.
- Build and interpret a Huffman tree, and calculate the number of bits it saves compared with a fixed-length code.
Files are large and networks are finite, so almost everything you download has been compressed — made smaller before transmission and restored, or approximated, at the other end.
There are two fundamentally different ways to do it. Lossless compression finds a cleverer way to describe the same data, so nothing is discarded. Lossy compression throws data away permanently. Which is acceptable depends entirely on the file: a photograph can lose detail nobody will notice, a program cannot lose a single byte.
- Section 2 covers units of storage.
- Sections 3–4 cover file size calculations for images and sound.
- Section 5 covers why compression is needed.
- Sections 6–7 cover lossless and lossy compression.
- Section 8 covers choosing between them.
- Sections 9–11 consolidate with worked examples, misconceptions and a summary.
2.Measuring Data Storage
| Unit | Symbol | Equivalent |
|---|---|---|
| Bit | b | A single 1 or 0 — the smallest unit |
| Nibble | — | 4 bits |
| Byte | B | 8 bits |
| Kibibyte | KiB | 1024 bytes |
| Mebibyte | MiB | 1024 KiB |
| Gibibyte | GiB | 1024 MiB |
| Tebibyte | TiB | 1024 GiB |
| Pebibyte | PiB | 1024 TiB |
| Exbibyte | EiB | 1024 PiB |
Storage units step in 1024s rather than 1000s because computers work in binary, and — a round number in binary even though it looks arbitrary in denary.
The -bi- units (kibibyte, mebibyte) were introduced to make this explicit, since 'kilobyte' had been used ambiguously for both 1000 and 1024 bytes. In this syllabus, use the 1024-based units unless told otherwise.
- 1Converting down the scale - divide by 8, then by 1024 each time:
- 2bits --/8--> bytes --/1024--> KiB --/1024--> MiB --/1024--> GiB
- 3Converting up the scale - multiply instead.
- 4Worked: 20 000 000 bits
- 5/ 8 = 2 500 000 bytes
- 6/ 1024 = 2441.4 KiB
- 7/ 1024 = 2.38 MiB
This is also why a hard drive advertised as shows up as roughly once connected to a computer: drive manufacturers advertise using 1000-based units, while most operating systems display the 1024-based figure. No storage is actually missing — the same physical drive is simply being measured on two different scales.
Storage is also measured on a purely decimal scale, using the ordinary metric prefixes and powers of 1000 instead of 1024:
| Unit | Symbol | Equivalent |
|---|---|---|
| Kilobyte | kB | 1000 bytes () |
| Megabyte | MB | 1000 kB ( bytes) |
| Gigabyte | GB | 1000 MB ( bytes) |
| Terabyte | TB | 1000 GB ( bytes) |
These names look almost identical to the -bi- units but are a different, smaller value at every step, since . Some exam questions use kilobyte, megabyte, gigabyte and terabyte to mean the decimal (powers-of-1000) value directly, with no kibi/mebi/gibi/tebi form at all — so check whether a question gives you the definition it wants, and use exactly the multiplier it states.
3.Calculating Image File Size
- 1An image is 1024 x 768 pixels with a colour depth of 24 bits.
- 2Number of x 768 = 786 432
- 3File size in 432 x 24 = 18 874 368 bits
- 4File size in 874 368 / 8 = 2 359 296 bytes
- 5File size in 359 296 / 1024 = 2304 KiB
- 6File size in MiB
Work through the units in order and show each step. Method marks are available even if one division goes wrong, and skipping steps forfeits them.
- 1Sampling44 100 samples/s · × 16 bits/sample
- 2Duration× 10 seconds · × 2 channels
- 3Raw size14 112 000 bits · show the units
- 4Convert÷ 8 = 1 764 000 bytes then chosen prefix
Sample rate × resolution × duration × channels gives bits; divide by 8 only when converting bits to bytes.
4.Calculating Sound File Size
- 1A recording lasts 3 minutes at a sample rate of 44 100 Hz
- 2with a sample resolution of 16 bits.
- 3Time in seconds = 60 = 180 s
- 4File size in 100 x 16 x 180 = 127 008 000 bits
- 5File size in 008 000 / 8 = 15 876 000 bytes
- 6File size in 876 000 / 1024 = 15 503.9 KiB
- 7File size in 503.9 / 1024 = 15.1 MiB
The first line is the one most often forgotten: the formula needs seconds, so minutes must be converted before anything else.
- Both formulas give bits. Divide by 8 before doing anything with bytes.
- Convert time to seconds and check whether the question wants bits, bytes, KiB or MiB — then stop at that unit.
A rough sanity check helps too. A few minutes of CD-quality audio is tens of megabytes; a photograph is a few megabytes. An answer of bytes or gigabytes signals an error in the unit conversion.
- 1Repeated dataAAAAAAAABB · raw: 10 symbols
- 2RLE result8A 2B · 4 stored items
- 3Alternating dataABABABAB · raw: 8 symbols
- 4RLE result1A1B1A1B… can be larger
Lossless RLE reproduces the original exactly, but alternating data can create overhead instead of a saving.
5.Why Compression Is Needed
- 1Identify the file typeAsk whether the file still works if some data is missing.
- 2Program or document?If no — a program, spreadsheet or text file — choose lossless. Every byte is needed.
- 3Photo, music or video?If yes — choose lossy, since detail can be discarded without a noticeable difference.
- 4Justify the choiceState why the loss is or is not tolerable and what is gained — smaller file, faster transmission, less bandwidth.
Questions usually want the benefit tied to a consequence. 'It makes the file smaller' is the definition, not the reason.
Better: a smaller file takes less time to transmit over a network of a given bandwidth, and takes up less storage space so more files can be kept. Both name what is actually gained.
Putting real numbers on it makes the case concrete: uncompressed 4K video can require several hundred megabits per second to stream, far beyond what most home internet connections could sustain. Compression is not an optional extra for video streaming — it is what makes the service possible at all over an ordinary connection.
- 1LosslessOriginal recovered exactly · text, code, records
- 2LossySome detail discarded · photos, audio, video
- 3Comparesize saving · quality and purpose
- 4DecideCan any data be lost? · What quality is acceptable?
Use lossless when exact recovery matters; use lossy only when permanent information loss is acceptable.
6.Lossless Compression
Nothing is discarded. Instead the data is described more efficiently, by recording patterns and repetitions rather than every individual item.
The standard example is run-length encoding (RLE), which replaces a run of repeated values with the value and a count.
- 1Original pixel row (12 items):
- 2W W W W W B B W W W W W
- 3Run-length encoded (6 items):
- 45W 2B 5W
- 5The original can be rebuilt exactly:
- 65W -> W W W W W
- 72B -> B B
- 85W -> W W W W W
- 9Nothing has been lost.
The saving depends entirely on repetition. A picture with large areas of flat colour compresses dramatically; a detailed photograph where adjacent pixels differ may compress hardly at all — and RLE can even make such a file larger, since each single pixel now needs a count as well.
Another lossless technique replaces repeated words or phrases in a text file with a shorter code, storing an index that maps codes back to the original text.
This family of technique is called dictionary-based compression, and it underlies common general-purpose tools such as ZIP. Rather than working pixel by pixel like RLE, it scans for any repeated sequence of bytes — a common word, a repeated block of program code — and replaces later occurrences with a short reference back to the first one. This is why ordinary documents and program files, which contain far more repeated structure than a photograph does, often shrink substantially under lossless compression even though no pixel-style runs are involved.
RLE exploits repetition. A different lossless method, Huffman coding, exploits frequency instead: it gives the most common symbols the shortest binary codes and the rarest symbols the longest codes, instead of using a fixed number of bits for every symbol as ASCII does. A message dominated by one or two common characters can then be stored using far fewer bits per character on average.
The codes come from a binary tree, built from the bottom up:
- 1Count how many times each symbol occurs. Each symbol starts as its own one-node tree, labelled with its frequency.
- 2Take the two trees with the smallest frequency and join them under a new node, whose frequency is their sum.
- 3Repeat step 2, always combining the two lightest trees, until only one tree remains.
- 4Read off each symbol's code as the path from the root to its leaf: left is 0, right is 1.
Because every symbol sits at a leaf — a node with nothing below it — no symbol's code is ever the start of another symbol's code. That is what lets the decoder split a run of bits back into the right symbols with no separators, working left to right and stopping the instant a valid code is matched.
A message contains: A , B , C , D (8 symbols in total).
- 1Start (sorted by frequency): D:1 C:1 B:2 A:4
- 2Combine the two lightest, D and C: [CD]:2 B:2 A:4
- 3Combine the two lightest, B and [CD]: [BCD]:4 A:4
- 4Combine the last two, A and [BCD]: [ABCD]:8 <- root, one tree left
Reading the path from the root (, ), with A as the root's left child and [BCD] as its right child, B as [BCD]'s left child and [CD] as its right child, and C, D as [CD]'s left and right children:
| Symbol | Frequency | Code | Bits used |
|---|---|---|---|
| A | 4 | 0 | |
| B | 2 | 10 | |
| C | 1 | 110 | |
| D | 1 | 111 |
The most frequent symbol, A, gets the shortest code; the two rarest, C and D, get the longest. Notice that , , and — no code is the start of another, exactly as the tree structure guarantees.
Total Huffman size: . Stored in ASCII instead, each of the 8 symbols would need 8 bits: bits. Huffman coding saves on this message.
- Huffman coding is still lossless — a symbol's code always decodes back to exactly that symbol, so the original message is rebuilt with nothing lost.
- The tree itself must also be stored or sent alongside the coded data, since the decoder needs it to know what each code means. On a very short message this overhead can cancel out most of the saving — Huffman coding pays off on longer messages with skewed frequencies.
7.Lossy Compression
The data removed is chosen to be what humans are least likely to notice:
- In images, the colour depth may be reduced, or fine detail and subtle colour differences discarded.
- In sound, frequencies outside the range of human hearing are removed, along with quiet sounds masked by louder ones played at the same time.
- In video, both are applied, plus removing detail that does not change between frames.
These techniques are what give familiar file formats their purpose: JPEG applies lossy compression tuned for photographs, MP3 and AAC apply it to audio, and formats such as MPEG apply it to video. A camera or phone that offers a photo 'quality' setting is really letting the user choose how much lossy compression to apply — higher quality means less is thrown away, and a larger file results.
Lossy compression achieves much greater size reductions than lossless — often ten times smaller — which is why it is used for streaming music and video and for photographs on the web.
But the loss is permanent. Once the data is gone it cannot be recovered, and each further round of lossy compression degrades the file again. Repeatedly saving a photograph in a lossy format visibly damages it.
This is why lossy compression is never used for files where every byte matters: a compressed program would not run, a compressed spreadsheet would contain wrong numbers, and a compressed text document would lose words.
8.Choosing Between Them
| Lossless | Lossy | |
|---|---|---|
| Data removed? | No | Yes — permanently |
| Original recoverable? | Yes, exactly | No |
| Size reduction | Smaller reduction | Much greater reduction |
| Quality | Unchanged | Reduced |
| Used for | Text, spreadsheets, programs, archives | Streaming music and video, web images |
Ask one question: does the file still work if some data is missing?
- No → use lossless. A program with missing bytes will not run; a spreadsheet with altered numbers is wrong; a text document loses words.
- Yes → lossy is acceptable, and its much greater size reduction becomes the deciding advantage. A photograph or a song remains perfectly usable with detail removed that nobody notices.
That single test answers every 'which method should be used and why' question, and the justification should name both halves: why the loss is or is not tolerable, and what is gained.
9.Exam-Style Worked Examples
Question. An image is pixels with a colour depth of 24 bits. (a) Calculate the file size in bits. (b) Convert to MiB. (c) State one way to reduce the file size without compression, and its effect.
- 1(a) pixels; bits.
- 2(b) bytes; ; .
- 3(c) Reduce the resolution — fewer pixels means a smaller file, but less detail is captured.
- 4Alternatively reduce the colour depth — fewer bits per pixel, but fewer colours available, which may cause banding.
Marking. 1 mark per point. Show every conversion step — method marks survive an arithmetic slip.
Question. A 4-minute recording uses a sample rate of and a sample resolution of 8 bits. (a) Calculate the file size in bits. (b) Convert to KiB. (c) The sample resolution is doubled to 16 bits — state the effect on file size and quality.
- 1(a) Time ; bits.
- 2(b) bytes; (to the nearest KiB).
- 3(c) The file size would double, since each sample now needs twice as many bits.
- 4The quality would improve, because each amplitude is recorded more accurately.
Marking. 1 mark per point. Converting minutes to seconds is the first marking point.
Question. (a) Define lossless compression. (b) Describe how run-length encoding compresses the pixel row . (c) Explain why RLE would be a poor choice for a detailed photograph.
- 1(a) A method that reduces file size without permanently removing data, so the original can be reconstructed exactly.
- 2(b) Runs of repeated characters are replaced by the character and a count: — 14 items reduced to 6.
- 3(c) In a detailed photograph adjacent pixels usually differ, so there are few runs to compress.
- 4Each pixel would need a count storing alongside it, so the file could end up larger than the original.
Marking. 1 mark per point. Part (c) rewards understanding that RLE depends on repetition, not just on being lossless.
Question. A company sends (i) a software installer and (ii) a music track to customers. State which compression method should be used for each and justify your choice.
- 1(i) Software installer: lossless.
- 2Every byte of a program is needed — if any data were permanently removed the program would not run correctly. Lossless allows the file to be reconstructed exactly.
- 3(ii) Music track: lossy.
- 4Frequencies outside the range of human hearing can be removed without a noticeable difference, and lossy gives a much greater reduction in file size — so it downloads faster and uses less bandwidth.
Marking. 1 mark per point. Each justification must say why the loss is or is not tolerable and what is gained.
10.Exam Tips & Common Misconceptions
- Storage units step in 1024s, because .
- Both file size formulas give bits — divide by 8 for bytes.
- Convert time to seconds before using the sound formula.
- Show every conversion step — method marks are available.
- Justify compression by less storage and faster transmission using less bandwidth, not just 'smaller file'.
- Lossless: no data removed, original recoverable exactly, smaller reduction.
- Lossy: data permanently removed, original not recoverable, much greater reduction.
- For RLE, state that it depends on repetition and can enlarge a varied file.
- To choose a method, ask whether the file still works with data missing.
- Build a Huffman tree by repeatedly combining the two lightest trees; a code is the root-to-leaf path (, right = 1).
- Compare Huffman coding against ASCII (8 bits per symbol) unless told otherwise, and remember the code table itself must also be stored.
- Common misconception: a kilobyte is 1000 bytes. In this syllabus a kibibyte is 1024 bytes.
- Common misconception: the file size formulas give bytes. They give bits.
- Common misconception: lossless compression loses a little quality. It loses nothing — the original is recoverable exactly.
- Common misconception: lossy compression can be undone. The removed data is gone permanently.
- Common misconception: compression always makes a file smaller. RLE can enlarge a file with little repetition.
- Common misconception: lossy is simply worse. It gives a much greater reduction, which is why streaming depends on it.
- Common misconception: repeatedly compressing a lossy file is harmless. Each round degrades it further.
- Common misconception: lossy compression is fine for any file if the setting is high enough. A program or spreadsheet cannot lose any data at all.
- Common misconception: compression and reducing resolution are the same. Reducing resolution discards pixels; compression re-encodes what is there.
- Common misconception: Huffman coding shortens every symbol's code. It only shortens codes for frequent symbols — a rare symbol can end up with a code longer than a fixed-length one.
- Common misconception: a kilobyte always means 1024 bytes. Some questions define kilobyte/megabyte/gigabyte as exact powers of 1000 — use the definition the question gives.
11.Summary
- Storage is measured in bits, nibbles (4 bits), bytes (8 bits), then KiB, MiB, GiB, TiB and PiB, each 1024 times the last.
- .
- .
- Both give bits: divide by 8 for bytes, then by 1024 for each step up.
- Compression is needed to use less storage space, allow faster transmission, use less bandwidth and make pages load faster.
- Lossless compression removes no data, so the original can be reconstructed exactly. Run-length encoding replaces runs of repeated values with the value and a count, and depends on repetition.
- Lossy compression permanently removes data — reducing colour depth, or removing frequencies outside human hearing — so the original cannot be recovered, but the reduction is much greater.
- Use lossless for programs, text and spreadsheets, where missing data breaks the file; use lossy for streamed music, video and web images, where the loss is not noticeable.
- Huffman coding is a lossless method that gives frequent symbols short codes and rare symbols long codes. Build the tree by repeatedly combining the two lightest trees; a symbol's code is the root-to-leaf path. The code table must also be stored.
- I can name the storage units and explain why they step in 1024s.
- I can calculate an image file size and convert the units.
- I can calculate a sound file size, remembering to convert to seconds.
- I can give reasons for compression tied to real consequences.
- I can define lossless compression and describe run-length encoding.
- I can explain when RLE works well and when it fails.
- I can define lossy compression and describe what data is removed.
- I can choose the right method for a given file and justify it fully.
- I can build a Huffman tree from symbol frequencies, read off its codes and calculate the bits saved compared with ASCII.