1.Lesson overview
- 13.3 Floating-point numbers, representation and manipulation
- 5.3 The binary number system
- 2.1 Binary Numbers
- 1Describe binary floating-point format using two's-complement mantissa and exponent.
- 2Convert binary floating-point values to denary and denary values to binary floating-point.
- 3Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.
- 4Normalise a floating-point value and explain why normalisation is used.
- 5Explain the range and precision trade-off when bits are allocated between mantissa and exponent.
- 6Explain rounding error, underflow and overflow in floating-point calculations.
- 7Calculate the absolute and relative error of a stored approximation, and explain why relative error is usually the more useful measure.
Start with the language. Establish mantissa and exponent before attempting a trace or design.
Then explain the mechanism. Follow the lesson's core model through encode sign and scale → normalise value.
Apply it to a system. Use the case study to connect technical choices to a constraint, risk and consequence.
Finish with exam reasoning. Show the working or trace, state assumptions and justify a choice against the stated requirement.
2.Why integers and fixed point are not enough
| Representation | Strength | Limitation |
|---|---|---|
| Integer | Exact whole-number arithmetic in range. | No fractional values. |
| Fixed point | Predictable fractional precision. | Limited range when the point position is fixed. |
| Floating point | Represents a wide range of magnitudes. | Precision is limited and many values are approximations. |
Floating point represents a real value as a signed mantissa multiplied by a power of the base. The exponent gives range; the mantissa gives precision, so a fixed total number of bits forces a trade-off. Normalisation shifts the binary point and adjusts the exponent so the mantissa uses its available significant bits consistently. Many decimal fractions have no finite binary expansion, which means a stored value may be an approximation and arithmetic can accumulate rounding error.
3.Fixed-point representation
- 1Whole-number part
0101 before the point is 0 × 8 + 1 × 4 + 0 × 2 + 1 × 1 = 5.
- 2Fractional part
101 after the point is 1 × ½ + 0 × ¼ + 1 × ⅛ = 0.5 + 0.125 = 0.625.
- 3Combine
0101.101₂ represents 5 + 0.625 = 5.625 in denary.
- 1Whole part
3 in binary is 011.
- 2Fractional part, bit 1
0.25 × 2 = 0.5, which does not pass 1: write 0.
- 3Fractional part, bit 2
0.5 × 2 = 1.0, which passes 1: write 1, remainder 0.
- 4Fractional part, bit 3
Remainder 0 × 2 = 0: write 0. The fractional bits are 010.
- 5Combine
3.25 is stored as 011.010₂ in this format.
A fixed-point bit pattern is meaningless without knowing exactly where its binary point sits and how many bits lie on each side. Always state or use the format given in the question rather than assuming the point is at the right-hand end.
| Property | Fixed point | Floating point |
|---|---|---|
| Range | Fixed by the whole-number bit allocation; cannot represent very large or very small magnitudes in the same format. | Wide, because the exponent scales the mantissa across many magnitudes. |
| Precision | Constant absolute precision across the whole range, since the fractional-bit count never changes. | Precision varies with magnitude, because the same number of mantissa bits covers a larger gap at a larger exponent. |
| Speed of calculation | Faster: arithmetic uses ordinary integer circuits with no separate exponent alignment step. | Slower: hardware must align exponents before adding or subtracting, and re-normalise the result. |
Fixed point is often chosen for financial or embedded applications where values stay within a known range and fast, predictable arithmetic matters; floating point is chosen when a program must handle a much wider range of magnitudes and can accept variable precision.
4.Finite precision creates approximation
A floating-point format stores sign, significand and exponent, giving a wide dynamic range but only finite precision. Many simple decimal fractions have repeating binary expansions and cannot be represented exactly. Rounding error may accumulate, while subtracting nearly equal large values can lose significant digits through cancellation. Normalisation increases effective precision within a format but does not eliminate error. Compare values with a justified tolerance rather than assuming exact equality of computed decimals.
5.Mantissa, exponent and binary point
Keep the representation width and binary-point convention visible on every conversion line. A correct decimal value with the wrong mantissa/exponent allocation is not a correct floating-point answer.
6.Two's-complement mantissa and exponent
Increase exponent bits when the application must represent much larger or smaller magnitudes; increase mantissa bits when it needs finer precision. Neither change solves every problem because finite representation still has limits.
7.Convert a floating-point value to denary
- 1Separate the mantissa and exponent bit fields.
- 2Convert the signed mantissa using the specified fixed binary point.
- 3Convert the signed two's-complement exponent to denary.
- 4Multiply the mantissa by two raised to that exponent.
- 5Check the approximate magnitude: a positive exponent must enlarge the magnitude and a negative exponent must reduce it.
More exponent bits expand range but leave fewer bits for significant digits; more mantissa bits do the opposite.
8.Convert a denary value to binary floating point
- 1Convert the magnitude's integer and fractional parts to binary to the precision needed.
- 2Choose a binary scaling so the mantissa fits the stated field and can be normalised.
- 3Record the number of binary-point shifts as a signed exponent.
- 4Convert a negative mantissa and negative exponent to two's complement where required.
- 5Round or truncate according to the specified rule if the value does not fit exactly.
9.Normalisation
| Mantissa change | Exponent change | Why value is preserved |
|---|---|---|
| Shift mantissa left by one binary place | Decrease exponent by one | Mantissa doubles while the power-of-two factor halves. |
| Shift mantissa right by one binary place | Increase exponent by one | Mantissa halves while the power-of-two factor doubles. |
10.Range and precision: allocation of bits
11.Rounding error, overflow and underflow
| Condition | Meaning | Typical response |
|---|---|---|
| Overflow | The result's magnitude is too large for the maximum exponent or mantissa range. | Signal an error, use infinity or a defined exceptional result depending on the system. |
| Underflow | The non-zero magnitude is too small for the minimum representable range. | Round to zero or a smallest representable non-zero value depending on the format. |
| Rounding error | The exact value lies between representable values. | Store the nearest or otherwise specified approximation. |
A finite binary format cannot represent every decimal fraction exactly.
A sensor's true reading is 200.0, but a floating-point format stores it as 200.5. The absolute error is |200.5 − 200.0| = 0.5. The relative error is 0.5 ÷ 200.0 = 0.0025, or 0.25%.
A second sensor's true reading is 2.0, stored as 2.5: the absolute error is again 0.5, but the relative error is 0.5 ÷ 2.0 = 0.25, or 25%. The same absolute error is a far more serious problem for the smaller true value.
Absolute error alone does not show how serious an error is: the same absolute error can be negligible against a very large true value and severe against a very small one. Relative error scales the error against the size of the true value, so it allows a fair comparison of approximation quality between measurements of very different magnitudes.
12.Worked example: normalise without changing value
- 1A one-place left shift doubles the mantissa's represented value.
- 2To preserve the overall floating-point value, decrease the exponent by one.
- 3Decreasing the exponent divides the power-of-two scale factor by two.
- 4The doubling of the mantissa and halving of the scale factor cancel, so the represented real value remains unchanged.
- 5Repeat only while the mantissa can be normalised within the available bit field.
13.Extended worked example: reason through the method
Explain why a loop that adds 0.1 ten times may not compare exactly equal to 1.0 in binary floating point.
- 1
The decimal fraction 0.1 has an infinite repeating binary expansion, so a finite significand stores a nearby value.
- 2
Each addition rounds to the format's precision; the ten-step result may differ slightly from the separately stored representation of 1.0.
- 3
Use a tolerance such as |sum-1| < 10^-10 when that tolerance suits the application's required accuracy.
A tolerance must be chosen from the scale and error requirements; a single fixed tolerance is not correct for all computations.
14.Exam tips and common misconceptions
- Use the binary-point position supplied in the question; do not assume the mantissa is an ordinary integer.
- Interpret both mantissa and exponent as signed two's-complement values when the format states this.
- When shifting the mantissa left, decrease the exponent; when shifting it right, increase the exponent.
- Normalisation improves use of available precision but does not create infinite precision.
- Distinguish overflow, underflow and rounding error by whether the issue is range or finite precision.
15.Evidence and limits: Why integers and fixed point are not enough
The lesson gives this specific detail: An INTEGER stores whole values within a fixed range. Fixed-point representation can store fractions with a fixed binary-point position, but its range and precision are tied to that position. Scientific and engineering applications often need both very small and very large values. table Representation trade-off header-row Representation Strength Limitation Integer Exact whole-number arithmetic in range. No fractional values. Fixed point Predictable fractional precision. Limited range when the point position is fixed. Floating point Represents a wide range of magnitudes. Precision is limited and many values are…
This final check makes an answer more rigorous: it connects the conclusion back to the exact case rather than relying on a memorised sentence.
16.Method checkpoint: Fixed-point representation
This lesson-specific route is useful when working with Fixed-point representation. Keep each stage visible so that a reader can check the reasoning rather than only the final claim.
- 1
Whole-number part
- 2
Fractional part
- 3
Combine
17.Reasoning through Mantissa, exponent and binary point
Core explanation. A binary floating-point value stores two signed quantities. The mantissa carries significant binary digits and sign; the exponent scales the mantissa by a power of two. Cambridge questions use two's-complement representation for both. container Read the format before calculating sky Identify how many bits belong to the mantissa and exponent. Identify the fixed binary-point position for the mantissa, often immediately after its sign bit in a stated format. Interpret the exponent as a two's-complement integer. Apply the exponent as a power of two, moving the effective binary point right for positive exponent and…
Read this as a chain: identify the object or evidence first, connect it to the relevant principle, then make a conclusion that is no broader than the evidence allows.
Use this detail to support the learning target “Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.”. State why the displayed relationship leads to the outcome, rather than listing isolated facts or steps.
18.Compare the cases: Two's-complement mantissa and exponent and Convert a floating-point value to denary
Two's complement gives a signed representation. For a fixed binary point, the leading bit has negative weight and the remaining bits have positive fractional weights. The exponent is also interpreted as a signed two's-complement integer. container Useful check yellow For a normalised non-zero two's-complement mantissa, the first two bits differ: positive values begin 01 and negative values begin 10. Repeated leading sign bits indicate that the mantissa can be shifted to improve precision, subject to its representation rules. Keep the mantissa's fixed binary point in mind. Treating its bits as an ordinary…
list Method number Separate the mantissa and exponent bit fields. Convert the signed mantissa using the specified fixed binary point. Convert the signed two's-complement exponent to denary. Multiply the mantissa by two raised to that exponent. Check the approximate magnitude: a positive exponent must enlarge the magnitude and a negative exponent must reduce it. container Illustrative calculation green If the decoded mantissa is and the decoded exponent is , then
A strong comparison identifies one shared idea, one important difference, and the condition that tells you which case or method applies. This comparison supports “Normalise a floating-point value and explain why normalisation is used.”.
19.Summary and self-check
- A floating-point number uses a signed mantissa scaled by a power-of-two exponent.
- Normalisation shifts mantissa and compensates with an opposite exponent change.
- Mantissa bits improve precision; exponent bits improve magnitude range.
- Finite representation causes rounding error and can produce overflow or underflow.
- Fixed point uses a constant binary-point position: predictable, fast arithmetic but fixed range and precision. Floating point trades calculation speed for a wide range of magnitudes.
- Relative error (absolute error divided by the true value) compares approximation quality fairly across measurements of different sizes; absolute error alone does not.
- 1What happens to the exponent when a mantissa is shifted left once?
- 2Why can a binary floating-point value approximate rather than equal a denary fraction?
- 3Which part of the format would you enlarge to represent a wider range of magnitudes?
- 4Convert 6.5 to fixed point with 3 fractional bits, then convert your answer back to denary to check it.
- 5A true value is 50.0 and a stored approximation is 50.4. Calculate the absolute error and the relative error, and explain which is more informative here.
20.Detailed revision focus — normalisation, precision and representation limits
This lesson is about normalisation, precision and representation limits. In a strong answer, name the relevant representation or mechanism, apply it to the stated evidence, then give a conclusion that fits the conditions of the question.
Describe binary floating-point format using two's-complement mantissa and exponent.
Convert binary floating-point values to denary and denary values to binary floating-point.
Convert between denary and binary fixed-point representation for a stated number of whole and fractional bits, and compare fixed-point with floating-point range, precision and calculation speed.
21.Worked Example 3 — normalisation, precision and representation limits
A binary floating-point format stores a normalised mantissa with four fractional bits. Explain why the denary value 0.1 may be stored approximately rather than exactly, and identify a consequence for repeated calculations.
- 1Express the value in binary conceptually: some denary fractions recur indefinitely in base 2.
- 2Fit the recurring expansion into the finite mantissa, so a rounding decision is necessary.
- 3Remember that normalisation changes where the binary point is placed; it does not create extra mantissa bits.
- 4Link the small initial approximation to accumulation or comparison problems in later calculations.
0.1 usually has no finite binary expansion, so a finite mantissa stores a nearby value. Repeated addition can accumulate error, and direct equality tests can fail even when two displayed values look identical.
22.High-value distinction — range and precision
| Term | Meaning | Why the distinction matters |
|---|---|---|
| range | the spread of magnitudes reachable by the exponent | Use range only for its specific role; it is not interchangeable with precision. |
| precision | the fineness of values distinguishable at a chosen magnitude, largely controlled by mantissa bits | Use precision when this is the mechanism, condition or property the question actually describes. |
When comparing these ideas, state one difference in purpose or mechanism before giving an example. A pair of definitions with no comparison does not fully answer a “compare” question.
23.Mark-ready route — normalisation, precision and representation limits
- 1Identify the rule or representationStep 1
Express the value in binary conceptually: some denary fractions recur indefinitely in base 2.
- 2Apply it to the evidenceStep 2
Fit the recurring expansion into the finite mantissa, so a rounding decision is necessary.
- 3Keep the condition visibleStep 3
Remember that normalisation changes where the binary point is placed; it does not create extra mantissa bits.
- 4Check the conclusionStep 4
Link the small initial approximation to accumulation or comparison problems in later calculations.
Before finalising, check the command word, any stated width, unit, order or condition, and whether your conclusion answers the exact scenario rather than a similar one.
24.Targeted correction and transfer — normalisation, precision and representation limits
Saying normalisation makes a value exact, or confusing floating-point overflow with ordinary integer overflow.
New situation: A format gains two exponent bits but loses two mantissa bits. Explain the trade-off for a scientific measurement containing both very large and very close values.
Without notes, explain the difference between range and precision, then outline the method from the worked example in four or fewer steps.
25.Worked Example 4 — rounding a normalised binary value
Normalise 1101.101₂. If only five fractional mantissa bits are retained, give the rounded stored value and the absolute error in denary.
- 1Move the binary point four places left:
1101.101₂ = 0.1101101₂ × 2⁴. - 2Keep five mantissa fractional bits:
0.11011₂. - 3The next discarded bit is 0, so the retained mantissa is unchanged when rounding.
- 4Convert the stored value back to denary before calculating the error.
The stored value is 0.11011₂ × 2⁴, which is 13.5₁₀. The original value is 13.625₁₀, so the absolute error is 0.125. The loss is caused by finite mantissa precision, not by normalisation itself.