How IEEE 754 Floating Point Works: Formats, Rounding, and Why 0.1 Is Not Exact
A statement like 0.1 + 0.2 == 0.3 fails in Python, JavaScript, Java, Go, and C for the same reason. None of those languages stored the decimals you wrote. They stored the nearest values that fit a fixed binary format, then added those values. The surprise is not a language bug. It is the encoding.
IEEE 754 is the standard that defines that encoding, the rounding rules, and the special values. This article is about the binary formats almost every general-purpose language uses: how a finite number is packed into bits, why most decimal fractions have no exact representation, and which operations throw information away even when each input looked exact. Instruction execution itself is a separate topic, covered in How CPUs Execute Instructions.
What a binary floating-point number actually stores
A finite binary floating-point number is a sign, a power of two, and a significand. The significand is a fixed-width binary fraction. The width is the precision. The power of two is the exponent. Moving the exponent is a shift, not a change of the significant bits, which is why the format can represent both tiny and huge magnitudes with the same number of digits.
The two formats developers meet constantly are binary32 and binary64. C calls them float and double. JavaScript has only binary64 for Number. Python float is binary64. Their layouts:
- binary32: 1 sign bit, 8 exponent bits, 23 explicit significand bits. Exponent bias 127.
- binary64: 1 sign bit, 11 exponent bits, 52 explicit significand bits. Exponent bias 1023.
The stored significand field does not hold the leading bit of a normal number. That bit is implied to be 1. A normal value decodes as
(−1)sign × 2stored exponent − bias × (1 + fraction / 2p)
where p is 23 or 52. The implicit 1 is why people say binary64 has 53 bits of precision even though only 52 bits are written down. The all-zero and all-ones exponent patterns are reserved, so the largest finite binary64 exponent is 1023, not 1024, and the smallest normal exponent is −1022.
Four encodings, not one formula
The exponent field selects the kind of value. Using binary64:
- Exponent 1 through 2046: normal number. Implicit leading 1. Unbiased exponent is the field minus 1023.
- Exponent 0, fraction 0: zero. The sign bit still exists, so +0 and −0 are different bit patterns. They compare equal, but the sign can survive in a reciprocal:
1 / -0.0is −infinity in IEEE arithmetic. - Exponent 0, fraction nonzero: subnormal. There is no implicit 1. The exponent is fixed at −1022, and the significand is
0.fraction. Subnormals fill the gap between zero and the smallest normal, at the cost of trailing precision. Without them, subtracting two close small numbers could flush to zero abruptly. - Exponent 2047, fraction 0: infinity, signed. Overflow with the default rounding mode produces infinity rather than wrapping.
- Exponent 2047, fraction nonzero: NaN. A quiet NaN has the top fraction bit set; a signaling NaN does not. Ordinary operations propagate a quiet NaN and do not trap. A signaling NaN is an operand that should raise the invalid-operation exception if the implementation is trapping.
NaN is not a single value. Any nonzero fraction with an all-ones exponent is a NaN, and payloads in the remaining bits are allowed. The comparison rule is the one that bites in application code: every comparison involving a NaN is false, including NaN == NaN. Sorting and map keys that use numeric equality will not treat NaN as a duplicate of itself. The predicate that detects it is x != x, or a language helper such as math.isnan.
Why 0.1 has no home in this format
A binary significand can represent fractions whose denominator is a power of two. 0.5 is 2−1. 0.25 is 2−2. 0.125 is 2−3. 0.1 is 1/10. Ten is 2 × 5, and the factor of 5 never cancels, so the binary expansion repeats:
0.110 = 0.0001100110011001100…2 = 1.1001100110011…2 × 2−4
The stored exponent for binary64 is −4 + 1023 = 1019, which is 01111111011. The explicit field keeps the first 52 bits after the leading 1. The next bits are 1001…, which is more than half an ulp, so round-ties-to-even rounds the last kept bit upward. The bit pattern programmers can check is 0x3fb999999999999a. Decoded back to decimal, that value is 0.1000000000000000055511151231257827021181583404541015625, not 0.1.
0.2 and 0.3 are the same kind of repeating expansion, rounded independently. Adding the rounded 0.1 to the rounded 0.2 does not pass through real arithmetic. It adds the two already-rounded significands, then rounds the sum again. The result is a binary64 value one ulp away from the rounding of 0.3, so equality fails. Printing often hides this, because the shortest round-trip decimal for both the stored 0.1 and the real 0.1 is the string 0.1.
The same literal can be rounded at different times. A C or Rust const double is usually converted by the compiler. A Python float literal is converted when the parser builds the constant. The required rounding for a decimal-to-binary conversion is still round-ties-to-even, so the bit pattern of 0.1 agrees across compliant implementations. Where languages diverge is decimal formatting, not the stored value.
Rounding is a rule, not a vibe
Every operation that cannot represent its exact real result must round to a representable value. IEEE 754 defines five rounding attributes that matter in practice:
- roundTiesToEven, the default. Round to nearest, and if the exact value sits halfway between two candidates, pick the one whose least significant bit is zero.
- roundTowardPositive and roundTowardNegative, used by interval arithmetic.
- roundTowardZero, truncation, which is what many integer casts do rather than what floating-point add does.
- roundTiesToAway, less common in binary hardware.
Ties-to-even is why a halfway case does not always round away from zero. It is also why repeating a round-to-nearest step does not systematically drift in one direction. The unit in the last place, the ulp, is the gap between a value and the next representable value at that magnitude. Near 1, binary64 ulps are about 2−52, roughly 2.22 × 10−16. Near 253, the gap is 1, so not every integer above 253 is representable. Above 253, adding 1 can be a no-op. That is a format limit, not an overflow.
Where exact inputs still lose digits
Representation error is only the first cut. Two further operations discard information even when every input was a dyadic rational.
Cancellation happens when two close numbers are subtracted. The leading bits match and disappear. The result is exact in the floating-point sense — subtraction of nearby values is correctly rounded and often lands on a representable number — but the surviving bits may be the ones that were already noise in the inputs. A classic shape is b*b - 4*a*c when the discriminant is small relative to b*b. The mathematical identity is fine. The evaluated form amplifies the relative error already present in the product.
Absorption happens at the other end. Adding a small value to a large one shifts the small significand so far right that it falls off the end. 1e16 + 1.0 in binary64 is 1e16. Loops that accumulate a tiny increment into a large total silently stop advancing. Compensated summation, such as Kahan summation, keeps a running correction term so those lost low bits are reintroduced on a later step.
A fused multiply-add computes a * b + c with one rounding instead of two. Hardware instructions such as fma do this. It is more accurate for the fused expression, and it is a different operation from a multiply followed by an add. Code that was tuned against the two-rounding result can change when a compiler contracts the expression. Contraction is an optimization that is allowed to alter bit patterns, which is why numerical tests should not assume it is off.
What to do about it
Equality against a decimal literal is the wrong test for a computed binary result. Compare with a tolerance scaled to the magnitude, or compare ulps, and do not use an absolute epsilon for values that can be both tiny and huge. For quantities that are decimal by definition — money, tax rates, invoice quantities — use an integer minor unit or a decimal type. IEEE 754 also defines decimal formats, but the float and double in mainstream languages are binary. Python's decimal.Decimal, Java's BigDecimal, and fixed-point integers are the tools that round in base 10.
Binary floating point is the right tool when the problem is scientific or geometric and relative error is acceptable: coordinates, physical simulation, machine learning activations. It is the wrong tool when the specification names decimal places. The format did not fail in the first case. The type choice failed in the second.
Takeaways
A binary64 value is a sign, a biased power of two, and a 53-bit significand with the leading 1 implied for normal numbers. Subnormals, infinities, and NaNs reuse the reserved exponent patterns rather than sitting outside the format. Most decimal fractions are infinite in base 2, so a literal such as 0.1 is stored as the nearest representable value, 0x3fb999999999999a, and later operations round again. The useful check is not whether floating point is "precise." It is whether the quantity you have is a binary approximation or a decimal fact.