1.Lesson overview
- 1.2 Text, sound and images
- 3.2 Data representation
- 3.5 Character encoding
- 3.6 Representing images
- 3.7 Representing sound
- Understand how and why a computer represents text, and the use of character sets including ASCII and Unicode.
- Understand how and why a computer represents sound, including the effects of sample rate and sample resolution.
- Understand how and why a computer represents an image, including the effects of resolution and colour depth.
A computer can store only binary, yet it handles text, music and photographs. Each of those is made storable by the same two-step trick: agree on a code that assigns numbers to things, then store those numbers in binary.
For text the code is a character set. For images and sound the situation is harder, because both are naturally continuous and binary is not — so they must be sampled, chopping something smooth into discrete measurements. That sampling is always an approximation, and the whole topic turns on the trade-off between quality and file size.
- Sections 2–3 cover text and character sets.
- Sections 4–5 cover sound and its two quality settings.
- Sections 6–7 cover images and their two quality settings.
- Section 8 draws out the pattern common to all three.
- Sections 9–11 consolidate with worked examples, misconceptions and a summary.
2.Representing Text
Every character a computer stores — letters, digits, punctuation, spaces — is held as a number, and that number is stored in binary. The character set is what decides which number means which character.
The agreement is essential. If one computer stored for 'A' and another used for something else, text would be unreadable when transferred. A shared standard is what makes text portable between systems.
This was not always the case. In the early decades of computing, manufacturers each used their own, incompatible codes — IBM mainframes, for example, commonly used a different scheme called EBCDIC. Exchanging a document between two different makes of computer could turn readable text into garbled symbols, because each side attached a different meaning to the same numbers. ASCII, standardised in the 1960s, was the industry's agreement to stop this from happening.
- 1The word CAT in ASCII:
- 2C -> 67 -> 01000011
- 3A -> 65 -> 01000001
- 4T -> 84 -> 01010100
- 5Stored as: 01000011 01000001 01010100 (3 bytes)
Two details are worth noticing. Codes for consecutive letters are consecutive numbers, so 'A' is 65, 'B' is 66, 'C' is 67 — which makes sorting alphabetically simple arithmetic.
And upper and lower case are different characters with different codes: 'A' is 65 but 'a' is 97. That is why a program must convert case explicitly before comparing text.
This numeric relationship is genuinely useful in programs, not just trivia: because 'a' is 97 and 'z' is 122, a program can test whether a character is a lower-case letter simply by checking that its code lies in that range, and can convert a lower-case letter to upper case by subtracting 32 from its code — the constant gap between 'a' (97) and 'A' (65), and between every other matching pair.
- 1Analogue sourceContinuous waveform · changes over time
- 2Sample rateMeasurements per second · horizontal spacing
- 3Sample resolutionBits per sample · vertical quantisation
- 4Stored soundrate × bits × seconds · × channels
Higher rate captures time detail; higher resolution gives more amplitude levels. Both increase file size.
3.ASCII and Unicode
| ASCII | Unicode | |
|---|---|---|
| Bits per character | 7 (often stored in 8) | 16 or more |
| Number of characters | 128 | Over a million |
| Covers | English letters, digits, punctuation, control codes | Most of the world's writing systems, plus symbols and emoji |
| File size | Smaller | Larger — more bits per character |
| Suitable for | English-only text | International text |
ASCII uses 7 bits, so it can represent different characters. That is enough for the English alphabet in both cases, the digits, punctuation and some control codes — and nothing else.
It cannot represent Greek, Arabic, Chinese, Hindi or any accented character. As computers became international this was untenable, so Unicode was developed using 16 bits or more, giving over a million codes — enough for most of the world's writing systems.
The cost is file size: Unicode uses at least twice as many bits per character as ASCII, so the same document stored in Unicode is at least twice as large. That is the trade-off the syllabus expects you to state.
Usefully, the first 128 Unicode codes are the same as ASCII, so ASCII text is already valid Unicode and older files still work.
A concrete example of the gap Unicode closes: the code point represents the Chinese character 中 (meaning 'middle'), and represents 😀, a grinning-face emoji. Neither could ever exist inside ASCII's 128 codes — there simply is not room — which is exactly the problem Unicode was built to solve.
4.Representing Sound
Sound is a continuously varying wave — an analogue quantity that takes any value at any instant. Binary can only store discrete numbers, so the wave must be converted.
This is done by sampling: measuring the amplitude of the wave at regular intervals and storing each measurement as a binary number.
The same trick already appears in cinema: a film is really a rapid sequence of still photographs — traditionally 24 per second — and the eye perceives continuous motion because the gaps between frames are too brief to notice. Digital audio does the same job in the time dimension: individual snapshots of amplitude, taken close enough together that the ear perceives an unbroken sound.
- 1The analogue waveThe microphone produces a continuously varying electrical signal matching the sound wave.
- 2Sample at intervalsThe amplitude is measured at regular time intervals — how often is the sample rate.
- 3Record each measurementEach amplitude is recorded as a binary number — how many bits are used is the sample resolution.
- 4Store the sequenceThe sequence of binary numbers is the digital sound file.
- 5Play backThe numbers are converted back into a continuous wave to drive a loudspeaker — an approximation of the original.
Between two samples the computer stores nothing — the original wave's behaviour in that gap is simply lost. Reconstruction joins the recorded points, producing a stepped approximation rather than the true smooth curve.
Taking samples more often and measuring each one more precisely both bring the approximation closer to the original — and both make the file larger. That is the whole trade-off in one sentence.
- 1Textcharacter → code point → encoded bits
- 2Soundwave → samples → quantised values
- 3Imagescene → pixel grid → colour codes
The controlling choices are character encoding, sample rate/resolution, and image resolution/colour depth.
5.Sample Rate and Sample Resolution
These are frequently confused, and questions often ask about one specifically:
- Sample rate is about time — how frequently the wave is measured.
- Sample resolution is about accuracy — how finely each individual measurement is recorded.
An analogy: photographing a moving object. The rate is how many photographs per second; the resolution is how much detail is in each one. Increasing either improves the record, and both cost storage.
Both increase file size and both improve quality, so the answer to 'what happens if you increase X' is always the same pair — but the reason differs, and that is where the mark is.
6.Representing Images
An image is stored as a grid of pixels. Each pixel's colour is recorded as a binary number, and the file is the sequence of those numbers.
- 1A simple 2-colour (1-bit) image, 8 x 8 pixels:
- 20 0 1 1 1 1 0 0 0 = white
- 30 1 0 0 0 0 1 0 1 = black
- 41 0 0 0 0 0 0 1
- 51 0 1 0 0 1 0 1
- 61 0 0 0 0 0 0 1
- 71 0 1 1 1 1 0 1
- 80 1 0 0 0 0 1 0
- 90 0 1 1 1 1 0 0
- 10Stored as 64 bytes
A real scene, like a real sound, is continuous — there is no smallest piece of it. Dividing it into a grid of pixels is the visual equivalent of sampling: the image is chopped into discrete units, and everything within one pixel is recorded as a single colour.
So detail smaller than one pixel is lost, which is exactly why enlarging a low-resolution image produces visible blocks rather than more detail. The detail was never stored.
This is why pixel density matters in practice: a printed photograph viewed from a normal reading distance can look sharp at around pixels per inch, while the same image viewed on a phone screen held much closer to the eye needs a higher density to look equally sharp. In both cases, detail finer than a single pixel's width simply cannot be shown, no matter how good the display is.
7.Resolution and Colour Depth
| Colour depth | Number of colours | Calculation |
|---|---|---|
| 1 bit | 2 | |
| 2 bits | 4 | |
| 8 bits | 256 | |
| 24 bits | 16.7 million |
A colour depth of bits gives different colours. This is directly examinable in both directions:
- Given the colour depth, the number of colours is .
- Given the number of colours needed, the colour depth is the smallest for which is at least that many.
So an image needing 100 distinct colours requires 7 bits, since is too few and is enough.
24-bit colour is the common standard for photographs: 8 bits each for red, green and blue, giving over 16 million combinations — more than the human eye can distinguish.
8.The Pattern Behind All Three
| Text | Sound | Images | |
|---|---|---|---|
| Unit stored | Character | Sample | Pixel |
| How many | Number of characters | Sample rate time | Resolution (width height) |
| Bits per unit | Character set (7 or 16+) | Sample resolution | Colour depth |
| File size | characters bits | rate resolution time | width height depth |
All three representations work identically: count the number of units, multiply by the bits per unit. Every file size calculation in the course is that single multiplication.
The two quality settings for sound and images map onto each other exactly:
- Sample rate and image resolution both control how many units are captured.
- Sample resolution and colour depth both control how many bits per unit.
Recognising this reduces four separate facts to one idea — and it explains why increasing any of the four improves quality and increases file size.
9.Exam-Style Worked Examples
Question. A company stores documents in ASCII but is expanding internationally. (a) State how many characters ASCII can represent and why. (b) Explain why Unicode is needed. (c) State one disadvantage of changing to Unicode.
- 1(a) ASCII uses 7 bits, so it can represent characters.
- 2(b) ASCII covers only English letters, digits and punctuation; it cannot represent other writing systems such as Greek, Arabic or Chinese.
- 3Unicode uses 16 bits or more, giving over a million codes — enough for most of the world's writing systems.
- 4(c) Unicode uses more bits per character, so the same document produces a larger file.
Marking. 1 mark per point. Part (a) needs the calculation , not just the number 128.
Question. A sound is recorded at a sample rate of with a sample resolution of 16 bits, for 30 seconds. (a) Calculate the file size in bits. (b) Convert to megabytes. (c) State what happens to the file size and quality if the sample rate is doubled.
- 1(a) bits.
- 2(b) Divide by 8 for bytes: bytes. Divide by twice: .
- 3(c) The file size would double, because twice as many samples are stored.
- 4The quality would improve, because the wave is measured more often, so the stored version is a closer approximation to the original.
Marking. 1 mark per point. The formula gives bits — divide by 8 before converting to bytes, kilobytes or megabytes.
Question. An image is pixels wide and pixels high with a colour depth of 8 bits. (a) Calculate the file size in bits. (b) State how many colours can be used. (c) The colour depth is increased to 24 bits — state the effect on the number of colours and the file size.
- 1(a) bits.
- 2(b) colours.
- 3(c) The number of colours rises to , about 16.7 million.
- 4The file size triples, to bits, since each pixel now needs three times as many bits.
Marking. 1 mark per point. Colour depth affects file size proportionally — tripling the bits per pixel triples the total.
Question. A photographer reduces both the resolution and the colour depth of an image. (a) Define each term. (b) Explain the effect of each reduction on the image. (c) State why the photographer might do this.
- 1(a) Resolution is the number of pixels in the image; colour depth is the number of bits used per pixel.
- 2(b) Lower resolution means fewer pixels, so less detail is captured and the image may look blocky when enlarged.
- 3Lower colour depth means fewer colours available, so shades that were distinct may be stored as the same colour, causing banding.
- 4(c) To produce a smaller file — for faster transmission, less storage, or quicker loading on a web page.
Marking. 1 mark per point. Each effect must be tied to its own setting — resolution to detail, colour depth to colours.
10.Exam Tips & Common Misconceptions
- A character set is an agreed list — say so, because the agreement is what makes text portable.
- ASCII: 7 bits, 128 characters. Unicode: 16+ bits, over a million.
- State Unicode's disadvantage as larger file size.
- Sample rate is about time; sample resolution is about accuracy per sample.
- Resolution is the number of pixels; colour depth is bits per pixel.
- A colour depth of bits gives colours.
- File size formulas give bits — divide by 8 for bytes.
- Increasing any quality setting improves quality and increases file size — but give the reason specific to that setting.
- Both sound and images are approximations, because sampling discards what lies between the samples.
- Common misconception: the computer stores the letter itself. It stores a binary number from an agreed character set.
- Common misconception: 'A' and 'a' have the same code. They are different characters with different codes — 65 and 97.
- Common misconception: Unicode replaced ASCII entirely. The first 128 Unicode codes are the same as ASCII, so ASCII text is still valid.
- Common misconception: sample rate and sample resolution are the same. One is how often, the other how precisely.
- Common misconception: a digital recording is identical to the original. It is an approximation — everything between samples is lost.
- Common misconception: enlarging an image adds detail. Detail smaller than a pixel was never stored.
- Common misconception: colour depth is the number of colours. It is the number of bits; the colours are .
- Common misconception: 8-bit colour means 8 colours. It means .
- Common misconception: file size formulas give bytes. They give bits.
11.Summary
- All data is stored as binary by agreeing a code and storing the resulting numbers.
- A character set is an agreed list of characters, each with a unique binary code — the agreement is what lets text move between systems.
- ASCII uses 7 bits, giving 128 characters — enough for English only. Unicode uses 16 bits or more, giving over a million codes for most of the world's writing systems, at the cost of larger files.
- Sound is continuous, so it is sampled: the amplitude is measured at regular intervals and stored as binary numbers.
- Sample rate is the number of samples per second; sample resolution is the number of bits per sample. Increasing either improves quality and increases file size.
- .
- An image is a grid of pixels, each stored as a binary colour value.
- Resolution is the number of pixels; colour depth is the bits per pixel, giving colours.
- .
- All three follow one pattern: number of units bits per unit. Sound and image files are always approximations, because sampling discards everything between the samples.
- I can explain what a character set is and why it must be agreed.
- I can compare ASCII and Unicode on bits, characters and file size.
- I can describe how sound is sampled and why playback is an approximation.
- I can define sample rate and sample resolution and keep them distinct.
- I can describe how an image is stored as pixels.
- I can define resolution and colour depth and calculate the number of colours.
- I can calculate sound and image file sizes and convert the units.
- I can explain the effect of changing any quality setting, with the correct reason.