1.Lesson overview

Syllabus focus
Cambridge IGCSE syllabus reference
  • 1.3 Data storage and compression
Edexcel IGCSE syllabus reference
  • 3.3 Data storage and compression
AQA IGCSE syllabus reference
  • 3.3 Units of information
  • 3.8 Data compression
By the end of this lesson you should be able to
  • Understand how data storage is measured.
  • Calculate the file size of an image file and a sound file, using information given.
  • Understand the purpose of and need for data compression.
  • Understand how files are compressed using lossy and lossless compression methods.
  • Build and interpret a Huffman tree, and calculate the number of bits it saves compared with a fixed-length code.

Files are large and networks are finite, so almost everything you download has been compressed — made smaller before transmission and restored, or approximated, at the other end.

There are two fundamentally different ways to do it. Lossless compression finds a cleverer way to describe the same data, so nothing is discarded. Lossy compression throws data away permanently. Which is acceptable depends entirely on the file: a photograph can lose detail nobody will notice, a program cannot lose a single byte.

How this chapter fits together
  • Section 2 covers units of storage.
  • Sections 3–4 cover file size calculations for images and sound.
  • Section 5 covers why compression is needed.
  • Sections 6–7 cover lossless and lossy compression.
  • Section 8 covers choosing between them.
  • Sections 9–11 consolidate with worked examples, misconceptions and a summary.

2.Measuring Data Storage

Units of data storage
UnitSymbolEquivalent
BitbA single 1 or 0 — the smallest unit
Nibble—4 bits
ByteB8 bits
KibibyteKiB1024 bytes
MebibyteMiB1024 KiB
GibibyteGiB1024 MiB
TebibyteTiB1024 GiB
PebibytePiB1024 TiB
ExbibyteEiB1024 PiB
Why 1024 and not 1000

Storage units step in 1024s rather than 1000s because computers work in binary, and — a round number in binary even though it looks arbitrary in denary.

The -bi- units (kibibyte, mebibyte) were introduced to make this explicit, since 'kilobyte' had been used ambiguously for both 1000 and 1024 bytes. In this syllabus, use the 1024-based units unless told otherwise.

Converting between units
  1. 1
    Converting down the scale - divide by 8, then by 1024 each time:
  2. 2
    bits --/8--> bytes --/1024--> KiB --/1024--> MiB --/1024--> GiB
  3. 3
    Converting up the scale - multiply instead.
  4. 4
    Worked: 20 000 000 bits
  5. 5
    / 8 = 2 500 000 bytes
  6. 6
    / 1024 = 2441.4 KiB
  7. 7
    / 1024 = 2.38 MiB

This is also why a hard drive advertised as shows up as roughly once connected to a computer: drive manufacturers advertise using 1000-based units, while most operating systems display the 1024-based figure. No storage is actually missing — the same physical drive is simply being measured on two different scales.

Decimal (SI) units: the other scale

Storage is also measured on a purely decimal scale, using the ordinary metric prefixes and powers of 1000 instead of 1024:

Decimal units of data storage
UnitSymbolEquivalent
KilobytekB1000 bytes ()
MegabyteMB1000 kB ( bytes)
GigabyteGB1000 MB ( bytes)
TerabyteTB1000 GB ( bytes)

These names look almost identical to the -bi- units but are a different, smaller value at every step, since . Some exam questions use kilobyte, megabyte, gigabyte and terabyte to mean the decimal (powers-of-1000) value directly, with no kibi/mebi/gibi/tebi form at all — so check whether a question gives you the definition it wants, and use exactly the multiplier it states.

3.Calculating Image File Size

Explore image resolution, colour depth and file sizeOpen full screen
Image file size
Working through one
Image file size calculation
  1. 1
    An image is 1024 x 768 pixels with a colour depth of 24 bits.
  2. 2
    Number of x 768 = 786 432
  3. 3
    File size in 432 x 24 = 18 874 368 bits
  4. 4
    File size in 874 368 / 8 = 2 359 296 bytes
  5. 5
    File size in 359 296 / 1024 = 2304 KiB
  6. 6
    File size in MiB

Work through the units in order and show each step. Method marks are available even if one division goes wrong, and skipping steps forfeits them.

Keep every factor and unit in a sound file-size calculation
  1. 1
    Sampling
    44 100 samples/s · × 16 bits/sample
  2. 2
    Duration
    × 10 seconds · × 2 channels
  3. 3
    Raw size
    14 112 000 bits · show the units
  4. 4
    Convert
    ÷ 8 = 1 764 000 bytes then chosen prefix

Sample rate × resolution × duration × channels gives bits; divide by 8 only when converting bits to bytes.

4.Calculating Sound File Size

Explore sample rate, bit depth and sound file sizeOpen full screen
Sound file size
Working through one
Sound file size calculation
  1. 1
    A recording lasts 3 minutes at a sample rate of 44 100 Hz
  2. 2
    with a sample resolution of 16 bits.
  3. 3
    Time in seconds = 60 = 180 s
  4. 4
    File size in 100 x 16 x 180 = 127 008 000 bits
  5. 5
    File size in 008 000 / 8 = 15 876 000 bytes
  6. 6
    File size in 876 000 / 1024 = 15 503.9 KiB
  7. 7
    File size in 503.9 / 1024 = 15.1 MiB

The first line is the one most often forgotten: the formula needs seconds, so minutes must be converted before anything else.

Two habits that prevent most errors
  • Both formulas give bits. Divide by 8 before doing anything with bytes.
  • Convert time to seconds and check whether the question wants bits, bytes, KiB or MiB — then stop at that unit.

A rough sanity check helps too. A few minutes of CD-quality audio is tens of megabytes; a photograph is a few megabytes. An answer of bytes or gigabytes signals an error in the unit conversion.

Run-length encoding helps only when runs are long enough
  1. 1
    Repeated data
    AAAAAAAABB · raw: 10 symbols
  2. 2
    RLE result
    8A 2B · 4 stored items
  3. 3
    Alternating data
    ABABABAB · raw: 8 symbols
  4. 4
    RLE result
    1A1B1A1B… can be larger

Lossless RLE reproduces the original exactly, but alternating data can create overhead instead of a saving.

5.Why Compression Is Needed

The reasons for compressing data
Less storage space
files take up less room
More files fit on a given drive, memory card or server, reducing the hardware needed.
Faster transmission
fewer bits to send
A smaller file takes less time to upload or download over a network of a given speed.
Less bandwidth used
cheaper and less congested
Streaming and downloading use less of the available bandwidth, so more users can be served at once.
Faster loading
better user experience
Web pages with compressed images load more quickly, and email attachments send faster.
Choosing a compression method
  1. 1
    Identify the file type
    Ask whether the file still works if some data is missing.
  2. 2
    Program or document?
    If no — a program, spreadsheet or text file — choose lossless. Every byte is needed.
  3. 3
    Photo, music or video?
    If yes — choose lossy, since detail can be discarded without a noticeable difference.
  4. 4
    Justify the choice
    State why the loss is or is not tolerable and what is gained — smaller file, faster transmission, less bandwidth.
The point that carries the mark

Questions usually want the benefit tied to a consequence. 'It makes the file smaller' is the definition, not the reason.

Better: a smaller file takes less time to transmit over a network of a given bandwidth, and takes up less storage space so more files can be kept. Both name what is actually gained.

Putting real numbers on it makes the case concrete: uncompressed 4K video can require several hundred megabits per second to stream, far beyond what most home internet connections could sustain. Compression is not an optional extra for video streaming — it is what makes the service possible at all over an ordinary connection.

Choose compression from the need to reconstruct the original
  1. 1
    Lossless
    Original recovered exactly · text, code, records
  2. 2
    Lossy
    Some detail discarded · photos, audio, video
  3. 3
    Compare
    size saving · quality and purpose
  4. 4
    Decide
    Can any data be lost? · What quality is acceptable?

Use lossless when exact recovery matters; use lossy only when permanent information loss is acceptable.

6.Lossless Compression

Sixteen coloured values grouped into four runs and replaced by four count-and-value pairs.
Figure 1: Run-length encoding reduces repeated values without losing information.
Build a lossless compression method and count its bitsOpen full screen
What it does
Lossless compression
A method that reduces file size without permanently removing any data, so the original file can be reconstructed exactly.

Nothing is discarded. Instead the data is described more efficiently, by recording patterns and repetitions rather than every individual item.

Run-length encoding

The standard example is run-length encoding (RLE), which replaces a run of repeated values with the value and a count.

Run-length encoding
  1. 1
    Original pixel row (12 items):
  2. 2
    W W W W W B B W W W W W
  3. 3
    Run-length encoded (6 items):
  4. 4
    5W 2B 5W
  5. 5
    The original can be rebuilt exactly:
  6. 6
    5W -> W W W W W
  7. 7
    2B -> B B
  8. 8
    5W -> W W W W W
  9. 9
    Nothing has been lost.

The saving depends entirely on repetition. A picture with large areas of flat colour compresses dramatically; a detailed photograph where adjacent pixels differ may compress hardly at all — and RLE can even make such a file larger, since each single pixel now needs a count as well.

Another lossless technique replaces repeated words or phrases in a text file with a shorter code, storing an index that maps codes back to the original text.

This family of technique is called dictionary-based compression, and it underlies common general-purpose tools such as ZIP. Rather than working pixel by pixel like RLE, it scans for any repeated sequence of bytes — a common word, a repeated block of program code — and replaces later occurrences with a short reference back to the first one. This is why ordinary documents and program files, which contain far more repeated structure than a photograph does, often shrink substantially under lossless compression even though no pixel-style runs are involved.

Huffman coding

RLE exploits repetition. A different lossless method, Huffman coding, exploits frequency instead: it gives the most common symbols the shortest binary codes and the rarest symbols the longest codes, instead of using a fixed number of bits for every symbol as ASCII does. A message dominated by one or two common characters can then be stored using far fewer bits per character on average.

Building a Huffman tree

The codes come from a binary tree, built from the bottom up:

  1. 1
    Count how many times each symbol occurs. Each symbol starts as its own one-node tree, labelled with its frequency.
  2. 2
    Take the two trees with the smallest frequency and join them under a new node, whose frequency is their sum.
  3. 3
    Repeat step 2, always combining the two lightest trees, until only one tree remains.
  4. 4
    Read off each symbol's code as the path from the root to its leaf: left is 0, right is 1.

Because every symbol sits at a leaf — a node with nothing below it — no symbol's code is ever the start of another symbol's code. That is what lets the decoder split a run of bits back into the right symbols with no separators, working left to right and stopping the instant a valid code is matched.

Worked example: building the tree

A message contains: A , B , C , D (8 symbols in total).

Combining the two lightest trees each time
  1. 1
    Start (sorted by frequency): D:1 C:1 B:2 A:4
  2. 2
    Combine the two lightest, D and C: [CD]:2 B:2 A:4
  3. 3
    Combine the two lightest, B and [CD]: [BCD]:4 A:4
  4. 4
    Combine the last two, A and [BCD]: [ABCD]:8 <- root, one tree left

Reading the path from the root (, ), with A as the root's left child and [BCD] as its right child, B as [BCD]'s left child and [CD] as its right child, and C, D as [CD]'s left and right children:

Codes read from the finished tree
SymbolFrequencyCodeBits used
A40
B210
C1110
D1111

The most frequent symbol, A, gets the shortest code; the two rarest, C and D, get the longest. Notice that , , and — no code is the start of another, exactly as the tree structure guarantees.

Total Huffman size: . Stored in ASCII instead, each of the 8 symbols would need 8 bits: bits. Huffman coding saves on this message.

Two things worth remembering
  • Huffman coding is still lossless — a symbol's code always decodes back to exactly that symbol, so the original message is rebuilt with nothing lost.
  • The tree itself must also be stored or sent alongside the coded data, since the decoder needs it to know what each code means. On a very short message this overhead can cancel out most of the saving — Huffman coding pays off on longer messages with skewed frequencies.
Five symbols with frequencies 12, 8, 6, 4 and 2 combined two at a time, smallest first, into one root. Left branches are 0 and right branches 1, so the commonest symbol gets the shortest code.
Figure 2: Depth in the tree is code length, so the commonest symbol sits nearest the root and the rarest sinks deepest.

7.Lossy Compression

Explore lossy compression and the error it introducesOpen full screen
What it does
Lossy compression
A method that reduces file size by permanently removing some data, so the original file cannot be reconstructed.

The data removed is chosen to be what humans are least likely to notice:

  • In images, the colour depth may be reduced, or fine detail and subtle colour differences discarded.
  • In sound, frequencies outside the range of human hearing are removed, along with quiet sounds masked by louder ones played at the same time.
  • In video, both are applied, plus removing detail that does not change between frames.

These techniques are what give familiar file formats their purpose: JPEG applies lossy compression tuned for photographs, MP3 and AAC apply it to audio, and formats such as MPEG apply it to video. A camera or phone that offers a photo 'quality' setting is really letting the user choose how much lossy compression to apply — higher quality means less is thrown away, and a larger file results.

The trade-off

Lossy compression achieves much greater size reductions than lossless — often ten times smaller — which is why it is used for streaming music and video and for photographs on the web.

But the loss is permanent. Once the data is gone it cannot be recovered, and each further round of lossy compression degrades the file again. Repeatedly saving a photograph in a lossy format visibly damages it.

This is why lossy compression is never used for files where every byte matters: a compressed program would not run, a compressed spreadsheet would contain wrong numbers, and a compressed text document would lose words.

8.Choosing Between Them

Lossless compared with lossy
LosslessLossy
Data removed?NoYes — permanently
Original recoverable?Yes, exactlyNo
Size reductionSmaller reductionMuch greater reduction
QualityUnchangedReduced
Used forText, spreadsheets, programs, archivesStreaming music and video, web images
How to decide

Ask one question: does the file still work if some data is missing?

  • No → use lossless. A program with missing bytes will not run; a spreadsheet with altered numbers is wrong; a text document loses words.
  • Yes → lossy is acceptable, and its much greater size reduction becomes the deciding advantage. A photograph or a song remains perfectly usable with detail removed that nobody notices.

That single test answers every 'which method should be used and why' question, and the justification should name both halves: why the loss is or is not tolerable, and what is gained.

9.Exam-Style Worked Examples

Worked example 1 — image file size (4 marks)

Question. An image is pixels with a colour depth of 24 bits. (a) Calculate the file size in bits. (b) Convert to MiB. (c) State one way to reduce the file size without compression, and its effect.

  1. 1
    (a) pixels; bits.
  2. 2
    (b) bytes; ; .
  3. 3
    (c) Reduce the resolution — fewer pixels means a smaller file, but less detail is captured.
  4. 4
    Alternatively reduce the colour depth — fewer bits per pixel, but fewer colours available, which may cause banding.

Marking. 1 mark per point. Show every conversion step — method marks survive an arithmetic slip.

Worked example 2 — sound file size (4 marks)

Question. A 4-minute recording uses a sample rate of and a sample resolution of 8 bits. (a) Calculate the file size in bits. (b) Convert to KiB. (c) The sample resolution is doubled to 16 bits — state the effect on file size and quality.

  1. 1
    (a) Time ; bits.
  2. 2
    (b) bytes; (to the nearest KiB).
  3. 3
    (c) The file size would double, since each sample now needs twice as many bits.
  4. 4
    The quality would improve, because each amplitude is recorded more accurately.

Marking. 1 mark per point. Converting minutes to seconds is the first marking point.

Worked example 3 — lossless compression (4 marks)

Question. (a) Define lossless compression. (b) Describe how run-length encoding compresses the pixel row . (c) Explain why RLE would be a poor choice for a detailed photograph.

  1. 1
    (a) A method that reduces file size without permanently removing data, so the original can be reconstructed exactly.
  2. 2
    (b) Runs of repeated characters are replaced by the character and a count: — 14 items reduced to 6.
  3. 3
    (c) In a detailed photograph adjacent pixels usually differ, so there are few runs to compress.
  4. 4
    Each pixel would need a count storing alongside it, so the file could end up larger than the original.

Marking. 1 mark per point. Part (c) rewards understanding that RLE depends on repetition, not just on being lossless.

Worked example 4 — choosing a method (4 marks)

Question. A company sends (i) a software installer and (ii) a music track to customers. State which compression method should be used for each and justify your choice.

  1. 1
    (i) Software installer: lossless.
  2. 2
    Every byte of a program is needed — if any data were permanently removed the program would not run correctly. Lossless allows the file to be reconstructed exactly.
  3. 3
    (ii) Music track: lossy.
  4. 4
    Frequencies outside the range of human hearing can be removed without a noticeable difference, and lossy gives a much greater reduction in file size — so it downloads faster and uses less bandwidth.

Marking. 1 mark per point. Each justification must say why the loss is or is not tolerable and what is gained.

10.Exam Tips & Common Misconceptions

Exam tips
  • Storage units step in 1024s, because .
  • Both file size formulas give bits — divide by 8 for bytes.
  • Convert time to seconds before using the sound formula.
  • Show every conversion step — method marks are available.
  • Justify compression by less storage and faster transmission using less bandwidth, not just 'smaller file'.
  • Lossless: no data removed, original recoverable exactly, smaller reduction.
  • Lossy: data permanently removed, original not recoverable, much greater reduction.
  • For RLE, state that it depends on repetition and can enlarge a varied file.
  • To choose a method, ask whether the file still works with data missing.
  • Build a Huffman tree by repeatedly combining the two lightest trees; a code is the root-to-leaf path (, right = 1).
  • Compare Huffman coding against ASCII (8 bits per symbol) unless told otherwise, and remember the code table itself must also be stored.
Common misconceptions
  • Common misconception: a kilobyte is 1000 bytes. In this syllabus a kibibyte is 1024 bytes.
  • Common misconception: the file size formulas give bytes. They give bits.
  • Common misconception: lossless compression loses a little quality. It loses nothing — the original is recoverable exactly.
  • Common misconception: lossy compression can be undone. The removed data is gone permanently.
  • Common misconception: compression always makes a file smaller. RLE can enlarge a file with little repetition.
  • Common misconception: lossy is simply worse. It gives a much greater reduction, which is why streaming depends on it.
  • Common misconception: repeatedly compressing a lossy file is harmless. Each round degrades it further.
  • Common misconception: lossy compression is fine for any file if the setting is high enough. A program or spreadsheet cannot lose any data at all.
  • Common misconception: compression and reducing resolution are the same. Reducing resolution discards pixels; compression re-encodes what is there.
  • Common misconception: Huffman coding shortens every symbol's code. It only shortens codes for frequent symbols — a rare symbol can end up with a code longer than a fixed-length one.
  • Common misconception: a kilobyte always means 1024 bytes. Some questions define kilobyte/megabyte/gigabyte as exact powers of 1000 — use the definition the question gives.

11.Summary

Chapter summary
  • Storage is measured in bits, nibbles (4 bits), bytes (8 bits), then KiB, MiB, GiB, TiB and PiB, each 1024 times the last.
  • .
  • .
  • Both give bits: divide by 8 for bytes, then by 1024 for each step up.
  • Compression is needed to use less storage space, allow faster transmission, use less bandwidth and make pages load faster.
  • Lossless compression removes no data, so the original can be reconstructed exactly. Run-length encoding replaces runs of repeated values with the value and a count, and depends on repetition.
  • Lossy compression permanently removes data — reducing colour depth, or removing frequencies outside human hearing — so the original cannot be recovered, but the reduction is much greater.
  • Use lossless for programs, text and spreadsheets, where missing data breaks the file; use lossy for streamed music, video and web images, where the loss is not noticeable.
  • Huffman coding is a lossless method that gives frequent symbols short codes and rare symbols long codes. Build the tree by repeatedly combining the two lightest trees; a symbol's code is the root-to-leaf path. The code table must also be stored.
Check your understanding
  • I can name the storage units and explain why they step in 1024s.
  • I can calculate an image file size and convert the units.
  • I can calculate a sound file size, remembering to convert to seconds.
  • I can give reasons for compression tied to real consequences.
  • I can define lossless compression and describe run-length encoding.
  • I can explain when RLE works well and when it fails.
  • I can define lossy compression and describe what data is removed.
  • I can choose the right method for a given file and justify it fully.
  • I can build a Huffman tree from symbol frequencies, read off its codes and calculate the bits saved compared with ASCII.