1.Lesson overview

Syllabus focus
Cambridge IAL syllabus reference
  • 13.3 Floating-point numbers, representation and manipulation
AQA IAL syllabus reference
  • 5.3 The binary number system
AP Computer Science Principles syllabus reference
  • 2.1 Binary Numbers
By the end of this lesson you should be able to
  1. 1
    Describe binary floating-point format using two's-complement mantissa and exponent.
  2. 2
    Convert binary floating-point values to denary and denary values to binary floating-point.
  3. 3
    Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.
  4. 4
    Normalise a floating-point value and explain why normalisation is used.
  5. 5
    Explain the range and precision trade-off when bits are allocated between mantissa and exponent.
  6. 6
    Explain rounding error, underflow and overflow in floating-point calculations.
  7. 7
    Calculate the absolute and relative error of a stored approximation, and explain why relative error is usually the more useful measure.
Floating-point representation stores real numbers as a signed mantissa multiplied by a power of two. It trades exactness for a large range of magnitudes, which is powerful but means many apparently simple decimal values cannot be stored exactly.
How this lesson fits together
  • Start with the language. Establish mantissa and exponent before attempting a trace or design.

  • Then explain the mechanism. Follow the lesson's core model through encode sign and scale → normalise value.

  • Apply it to a system. Use the case study to connect technical choices to a constraint, risk and consequence.

  • Finish with exam reasoning. Show the working or trace, state assumptions and justify a choice against the stated requirement.

2.Why integers and fixed point are not enough

An INTEGER stores whole values within a fixed range. Fixed-point representation can store fractions with a fixed binary-point position, but its range and precision are tied to that position. Scientific and engineering applications often need both very small and very large values.
Representation trade-off
RepresentationStrengthLimitation
IntegerExact whole-number arithmetic in range.No fractional values.
Fixed pointPredictable fractional precision.Limited range when the point position is fixed.
Floating pointRepresents a wide range of magnitudes.Precision is limited and many values are approximations.
The core form is
. The exact bit allocation and binary-point convention are stated by the representation being used.
Technical checkpoint

Floating point represents a real value as a signed mantissa multiplied by a power of the base. The exponent gives range; the mantissa gives precision, so a fixed total number of bits forces a trade-off. Normalisation shifts the binary point and adjusts the exponent so the mantissa uses its available significant bits consistently. Many decimal fractions have no finite binary expansion, which means a stored value may be an approximation and arithmetic can accumulate rounding error.

3.Fixed-point representation

A fixed-point representation stores a value as a signed whole-number bit pattern with the binary point fixed at an agreed position, rather than at the right-hand end. Every value in the format uses exactly the same number of whole-number bits and the same number of fractional bits, so the range and precision are fixed once that position is chosen — they cannot trade off against each other as they can in floating point.
Converting a fixed-point pattern to denary
Treat the bits to the left of the fixed binary point as an ordinary two's-complement whole number, and the bits to the right as fractional place values that continue the pattern ½, ¼, ⅛, ... after the point.
Convert 0101.101₂ (4 whole bits, 3 fractional bits) to denary
  1. 1
    Whole-number part

    0101 before the point is 0 × 8 + 1 × 4 + 0 × 2 + 1 × 1 = 5.

  2. 2
    Fractional part

    101 after the point is 1 × ½ + 0 × ¼ + 1 × ⅛ = 0.5 + 0.125 = 0.625.

  3. 3
    Combine

    0101.101₂ represents 5 + 0.625 = 5.625 in denary.

Converting a denary value to fixed point
Convert the whole-number part using ordinary denary-to-binary conversion. Convert the fractional part by repeatedly multiplying by 2 and recording whether the result passes 1: record a 1 and subtract 1 whenever it does, record a 0 otherwise, and stop once the stated number of fractional bits has been produced.
Convert 3.25 to fixed point with 3 fractional bits
  1. 1
    Whole part

    3 in binary is 011.

  2. 2
    Fractional part, bit 1

    0.25 × 2 = 0.5, which does not pass 1: write 0.

  3. 3
    Fractional part, bit 2

    0.5 × 2 = 1.0, which passes 1: write 1, remainder 0.

  4. 4
    Fractional part, bit 3

    Remainder 0 × 2 = 0: write 0. The fractional bits are 010.

  5. 5
    Combine

    3.25 is stored as 011.010₂ in this format.

State the point position explicitly

A fixed-point bit pattern is meaningless without knowing exactly where its binary point sits and how many bits lie on each side. Always state or use the format given in the question rather than assuming the point is at the right-hand end.

Fixed point compared with floating point
Range, precision and speed
PropertyFixed pointFloating point
RangeFixed by the whole-number bit allocation; cannot represent very large or very small magnitudes in the same format.Wide, because the exponent scales the mantissa across many magnitudes.
PrecisionConstant absolute precision across the whole range, since the fractional-bit count never changes.Precision varies with magnitude, because the same number of mantissa bits covers a larger gap at a larger exponent.
Speed of calculationFaster: arithmetic uses ordinary integer circuits with no separate exponent alignment step.Slower: hardware must align exponents before adding or subtracting, and re-normalise the result.

Fixed point is often chosen for financial or embedded applications where values stay within a known range and fast, predictable arithmetic matters; floating point is chosen when a program must handle a much wider range of magnitudes and can accept variable precision.

4.Finite precision creates approximation

A floating-point format stores sign, significand and exponent, giving a wide dynamic range but only finite precision. Many simple decimal fractions have repeating binary expansions and cannot be represented exactly. Rounding error may accumulate, while subtracting nearly equal large values can lose significant digits through cancellation. Normalisation increases effective precision within a format but does not eliminate error. Compare values with a justified tolerance rather than assuming exact equality of computed decimals.

5.Mantissa, exponent and binary point

A binary floating-point value stores two signed quantities. The mantissa carries significant binary digits and sign; the exponent scales the mantissa by a power of two. Cambridge questions use two's-complement representation for both.
Read the format before calculating
Identify how many bits belong to the mantissa and exponent.
Identify the fixed binary-point position for the mantissa, often immediately after its sign bit in a stated format.
Interpret the exponent as a two's-complement integer.
Apply the exponent as a power of two, moving the effective binary point right for positive exponent and left for negative exponent.
Do not assume a universal bit layout
Use the bit lengths and binary-point convention supplied in the question. A correct technique with the wrong format produces the wrong value.
Trace the constraint

Keep the representation width and binary-point convention visible on every conversion line. A correct decimal value with the wrong mantissa/exponent allocation is not a correct floating-point answer.

6.Two's-complement mantissa and exponent

Two's complement gives a signed representation. For a fixed binary point, the leading bit has negative weight and the remaining bits have positive fractional weights. The exponent is also interpreted as a signed two's-complement integer.
Useful check
For a normalised non-zero two's-complement mantissa, the first two bits differ: positive values begin 01 and negative values begin 10. Repeated leading sign bits indicate that the mantissa can be shifted to improve precision, subject to its representation rules.
Keep the mantissa's fixed binary point in mind. Treating its bits as an ordinary unsigned integer is a common source of incorrect conversion.
Design decision

Increase exponent bits when the application must represent much larger or smaller magnitudes; increase mantissa bits when it needs finer precision. Neither change solves every problem because finite representation still has limits.

7.Convert a floating-point value to denary

Method
  1. 1
    Separate the mantissa and exponent bit fields.
  2. 2
    Convert the signed mantissa using the specified fixed binary point.
  3. 3
    Convert the signed two's-complement exponent to denary.
  4. 4
    Multiply the mantissa by two raised to that exponent.
  5. 5
    Check the approximate magnitude: a positive exponent must enlarge the magnitude and a negative exponent must reduce it.
Illustrative calculation
If the decoded mantissa is and the decoded exponent is , then
. The stored mantissa is scaled; it is not the final denary value by itself.
Failure-mode check

More exponent bits expand range but leave fewer bits for significant digits; more mantissa bits do the opposite.

8.Convert a denary value to binary floating point

Method
  1. 1
    Convert the magnitude's integer and fractional parts to binary to the precision needed.
  2. 2
    Choose a binary scaling so the mantissa fits the stated field and can be normalised.
  3. 3
    Record the number of binary-point shifts as a signed exponent.
  4. 4
    Convert a negative mantissa and negative exponent to two's complement where required.
  5. 5
    Round or truncate according to the specified rule if the value does not fit exactly.
Always verify by converting your result back to denary. This catches an exponent sign error or a binary-point placement error before it propagates.

9.Normalisation

Normalisation shifts the mantissa and adjusts the exponent in the opposite direction so that the represented value stays unchanged while the available significant bits are used efficiently.
A 16-bit floating-point number: mantissa 0.1011000 (+0.6875) and exponent 00000011 (+3) give 5.5. Below, an unnormalised mantissa is shifted two places left and the exponent reduced by two, storing the same value.
Figure 1: Shifting the mantissa left and lowering the exponent by the same amount keeps the value but gives it more significant bits.
Direction of a normalising shift
Mantissa changeExponent changeWhy value is preserved
Shift mantissa left by one binary placeDecrease exponent by oneMantissa doubles while the power-of-two factor halves.
Shift mantissa right by one binary placeIncrease exponent by oneMantissa halves while the power-of-two factor doubles.
For a non-zero two's-complement mantissa, normalise until the first two bits differ if the stated representation uses that convention. This avoids wasting leading bits that carry no new precision.

10.Range and precision: allocation of bits

The number of mantissa bits controls precision: more bits distinguish more nearby values. The number of exponent bits controls range: more exponent values permit larger and smaller magnitudes.
Sixteen bits split three ways. 8 mantissa and 8 exponent bits: resolution 1 in 128, exponents -128 to 127. 12 and 4: 1 in 2048, exponents -8 to 7. 4 and 12: 1 in 8, exponents -2048 to 2047.
Figure 2: Every bit moved to the mantissa buys precision and costs range, and every bit moved to the exponent does the opposite.
The unavoidable trade-off
More mantissa bits
Improves precision and reduces rounding error for representable values, but leaves fewer bits for exponent range.
More exponent bits
Extends the range of magnitudes, but leaves fewer significant bits and therefore reduces precision.
There is no universally best allocation. A scientific system modelling enormous ranges may favour exponent bits; a financial or measurement task may favour precision, although fixed-point decimal approaches are often used for money.

11.Rounding error, overflow and underflow

Many denary fractions have infinite recurring binary expansions, so a finite mantissa stores an approximation. Rounding selects a nearby representable value; truncation discards remaining bits. Repeated calculations can accumulate or amplify the resulting small errors.
Limits of a finite format
ConditionMeaningTypical response
OverflowThe result's magnitude is too large for the maximum exponent or mantissa range.Signal an error, use infinity or a defined exceptional result depending on the system.
UnderflowThe non-zero magnitude is too small for the minimum representable range.Round to zero or a smallest representable non-zero value depending on the format.
Rounding errorThe exact value lies between representable values.Store the nearest or otherwise specified approximation.
Equality tests need care
After floating-point calculations, two values intended to be mathematically equal may differ by a tiny rounding error. Real programs often compare whether their difference is within an acceptable tolerance rather than using exact equality.
Failure-mode check

A finite binary format cannot represent every decimal fraction exactly.

Absolute and relative error
When a stored value only approximates the true value, the size of that approximation can be reported in two ways.
Worked example

A sensor's true reading is 200.0, but a floating-point format stores it as 200.5. The absolute error is |200.5 − 200.0| = 0.5. The relative error is 0.5 ÷ 200.0 = 0.0025, or 0.25%.

A second sensor's true reading is 2.0, stored as 2.5: the absolute error is again 0.5, but the relative error is 0.5 ÷ 2.0 = 0.25, or 25%. The same absolute error is a far more serious problem for the smaller true value.

Why relative error is usually the more useful measure

Absolute error alone does not show how serious an error is: the same absolute error can be negligible against a very large true value and severe against a very small one. Relative error scales the error against the size of the true value, so it allows a fair comparison of approximation quality between measurements of very different magnitudes.

12.Worked example: normalise without changing value

Question
A positive binary mantissa has unnecessary leading sign bits and is shifted one place left during normalisation. Explain how the exponent must change and why.
Step-by-step answer
  1. 1
    A one-place left shift doubles the mantissa's represented value.
  2. 2
    To preserve the overall floating-point value, decrease the exponent by one.
  3. 3
    Decreasing the exponent divides the power-of-two scale factor by two.
  4. 4
    The doubling of the mantissa and halving of the scale factor cancel, so the represented real value remains unchanged.
  5. 5
    Repeat only while the mantissa can be normalised within the available bit field.
The key idea is compensation: a mantissa shift always requires an opposite exponent adjustment.

13.Extended worked example: reason through the method

Problem

Explain why a loop that adds 0.1 ten times may not compare exactly equal to 1.0 in binary floating point.

Solution with reasoning
  1. 1

    The decimal fraction 0.1 has an infinite repeating binary expansion, so a finite significand stores a nearby value.

  2. 2

    Each addition rounds to the format's precision; the ten-step result may differ slightly from the separately stored representation of 1.0.

  3. 3

    Use a tolerance such as |sum-1| < 10^-10 when that tolerance suits the application's required accuracy.

Interpret and check

A tolerance must be chosen from the scale and error requirements; a single fixed tolerance is not correct for all computations.

14.Exam tips and common misconceptions

Common errors
  • Use the binary-point position supplied in the question; do not assume the mantissa is an ordinary integer.
  • Interpret both mantissa and exponent as signed two's-complement values when the format states this.
  • When shifting the mantissa left, decrease the exponent; when shifting it right, increase the exponent.
  • Normalisation improves use of available precision but does not create infinite precision.
  • Distinguish overflow, underflow and rounding error by whether the issue is range or finite precision.

15.Evidence and limits: Why integers and fixed point are not enough

The lesson gives this specific detail: An INTEGER stores whole values within a fixed range. Fixed-point representation can store fractions with a fixed binary-point position, but its range and precision are tied to that position. Scientific and engineering applications often need both very small and very large values. table Representation trade-off header-row Representation Strength Limitation Integer Exact whole-number arithmetic in range. No fractional values. Fixed point Predictable fractional precision. Limited range when the point position is fixed. Floating point Represents a wide range of magnitudes. Precision is limited and many values are…

Boundary check
State what the evidence, representation, calculation, or model does establish—and what extra condition or information would be needed before extending the conclusion. Keep the focus on “Describe binary floating-point format using two's-complement mantissa and exponent.”.

This final check makes an answer more rigorous: it connects the conclusion back to the exact case rather than relying on a memorised sentence.

16.Method checkpoint: Fixed-point representation

This lesson-specific route is useful when working with Fixed-point representation. Keep each stage visible so that a reader can check the reasoning rather than only the final claim.

  1. 1

    Whole-number part

  2. 2

    Fractional part

  3. 3

    Combine

Why the order matters
The method is tied to this lesson’s aim: Convert binary floating-point values to denary and denary values to binary floating-point.. A skipped stage can change the interpretation or invalidate the conclusion.

17.Reasoning through Mantissa, exponent and binary point

Core explanation. A binary floating-point value stores two signed quantities. The mantissa carries significant binary digits and sign; the exponent scales the mantissa by a power of two. Cambridge questions use two's-complement representation for both. container Read the format before calculating sky Identify how many bits belong to the mantissa and exponent. Identify the fixed binary-point position for the mantissa, often immediately after its sign bit in a stated format. Interpret the exponent as a two's-complement integer. Apply the exponent as a power of two, moving the effective binary point right for positive exponent and…

Read this as a chain: identify the object or evidence first, connect it to the relevant principle, then make a conclusion that is no broader than the evidence allows.

Explain, do not just name

Use this detail to support the learning target “Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.”. State why the displayed relationship leads to the outcome, rather than listing isolated facts or steps.

18.Compare the cases: Two's-complement mantissa and exponent and Convert a floating-point value to denary

Two's-complement mantissa and exponent

Two's complement gives a signed representation. For a fixed binary point, the leading bit has negative weight and the remaining bits have positive fractional weights. The exponent is also interpreted as a signed two's-complement integer. container Useful check yellow For a normalised non-zero two's-complement mantissa, the first two bits differ: positive values begin 01 and negative values begin 10. Repeated leading sign bits indicate that the mantissa can be shifted to improve precision, subject to its representation rules. Keep the mantissa's fixed binary point in mind. Treating its bits as an ordinary…

Convert a floating-point value to denary

list Method number Separate the mantissa and exponent bit fields. Convert the signed mantissa using the specified fixed binary point. Convert the signed two's-complement exponent to denary. Multiply the mantissa by two raised to that exponent. Check the approximate magnitude: a positive exponent must enlarge the magnitude and a negative exponent must reduce it. container Illustrative calculation green If the decoded mantissa is and the decoded exponent is , then

. The stored mantissa is scaled; it is not the final denary value by itself. container Failure-mode check yellow More…

A strong comparison identifies one shared idea, one important difference, and the condition that tells you which case or method applies. This comparison supports “Normalise a floating-point value and explain why normalisation is used.”.

19.Summary and self-check

Key takeaways
  • A floating-point number uses a signed mantissa scaled by a power-of-two exponent.
  • Normalisation shifts mantissa and compensates with an opposite exponent change.
  • Mantissa bits improve precision; exponent bits improve magnitude range.
  • Finite representation causes rounding error and can produce overflow or underflow.
  • Fixed point uses a constant binary-point position: predictable, fast arithmetic but fixed range and precision. Floating point trades calculation speed for a wide range of magnitudes.
  • Relative error (absolute error divided by the true value) compares approximation quality fairly across measurements of different sizes; absolute error alone does not.
Self-check
  1. 1
    What happens to the exponent when a mantissa is shifted left once?
  2. 2
    Why can a binary floating-point value approximate rather than equal a denary fraction?
  3. 3
    Which part of the format would you enlarge to represent a wider range of magnitudes?
  4. 4
    Convert 6.5 to fixed point with 3 fractional bits, then convert your answer back to denary to check it.
  5. 5
    A true value is 50.0 and a stored approximation is 50.4. Calculate the absolute error and the relative error, and explain which is more informative here.

20.Detailed revision focus — normalisation, precision and representation limits

What this topic requires you to connect

This lesson is about normalisation, precision and representation limits. In a strong answer, name the relevant representation or mechanism, apply it to the stated evidence, then give a conclusion that fits the conditions of the question.

Syllabus-aligned checkpoints
Checkpoint 1
Secure this before moving on

Describe binary floating-point format using two's-complement mantissa and exponent.

Checkpoint 2
Secure this before moving on

Convert binary floating-point values to denary and denary values to binary floating-point.

Checkpoint 3
Secure this before moving on

Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.

21.Worked Example 3 — normalisation, precision and representation limits

Question

A binary floating-point format stores a normalised mantissa with four fractional bits. Explain why the denary value 0.1 may be stored approximately rather than exactly, and identify a consequence for repeated calculations.

Method
  1. 1
    Express the value in binary conceptually: some denary fractions recur indefinitely in base 2.
  2. 2
    Fit the recurring expansion into the finite mantissa, so a rounding decision is necessary.
  3. 3
    Remember that normalisation changes where the binary point is placed; it does not create extra mantissa bits.
  4. 4
    Link the small initial approximation to accumulation or comparison problems in later calculations.
Answer and why it earns credit

0.1 usually has no finite binary expansion, so a finite mantissa stores a nearby value. Repeated addition can accumulate error, and direct equality tests can fail even when two displayed values look identical.

22.High-value distinction — range and precision

Use the precise term
TermMeaningWhy the distinction matters
rangethe spread of magnitudes reachable by the exponentUse range only for its specific role; it is not interchangeable with precision.
precisionthe fineness of values distinguishable at a chosen magnitude, largely controlled by mantissa bitsUse precision when this is the mechanism, condition or property the question actually describes.
Exam wording

When comparing these ideas, state one difference in purpose or mechanism before giving an example. A pair of definitions with no comparison does not fully answer a “compare” question.

23.Mark-ready route — normalisation, precision and representation limits

Reasoning sequence
  1. 1
    Identify the rule or representation
    Step 1

    Express the value in binary conceptually: some denary fractions recur indefinitely in base 2.

  2. 2
    Apply it to the evidence
    Step 2

    Fit the recurring expansion into the finite mantissa, so a rounding decision is necessary.

  3. 3
    Keep the condition visible
    Step 3

    Remember that normalisation changes where the binary point is placed; it does not create extra mantissa bits.

  4. 4
    Check the conclusion
    Step 4

    Link the small initial approximation to accumulation or comparison problems in later calculations.

Quality check

Before finalising, check the command word, any stated width, unit, order or condition, and whether your conclusion answers the exact scenario rather than a similar one.

24.Targeted correction and transfer — normalisation, precision and representation limits

Common trap

Saying normalisation makes a value exact, or confusing floating-point overflow with ordinary integer overflow.

Independent transfer

New situation: A format gains two exponent bits but loses two mantissa bits. Explain the trade-off for a scientific measurement containing both very large and very close values.

Retrieval prompt

Without notes, explain the difference between range and precision, then outline the method from the worked example in four or fewer steps.

25.Worked Example 4 — rounding a normalised binary value

Question

Normalise 1101.101₂. If only five fractional mantissa bits are retained, give the rounded stored value and the absolute error in denary.

Method
  1. 1
    Move the binary point four places left: 1101.101₂ = 0.1101101₂ × 2⁴.
  2. 2
    Keep five mantissa fractional bits: 0.11011₂.
  3. 3
    The next discarded bit is 0, so the retained mantissa is unchanged when rounding.
  4. 4
    Convert the stored value back to denary before calculating the error.
Answer and check

The stored value is 0.11011₂ × 2⁴, which is 13.5₁₀. The original value is 13.625₁₀, so the absolute error is 0.125. The loss is caused by finite mantissa precision, not by normalisation itself.