29.200 007 514 493 3 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 29.200 007 514 493 3(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
29.200 007 514 493 3(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 29.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

29(10) =


1 1101(2)


3. Convert to binary (base 2) the fractional part: 0.200 007 514 493 3.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.200 007 514 493 3 × 2 = 0 + 0.400 015 028 986 6;
  • 2) 0.400 015 028 986 6 × 2 = 0 + 0.800 030 057 973 2;
  • 3) 0.800 030 057 973 2 × 2 = 1 + 0.600 060 115 946 4;
  • 4) 0.600 060 115 946 4 × 2 = 1 + 0.200 120 231 892 8;
  • 5) 0.200 120 231 892 8 × 2 = 0 + 0.400 240 463 785 6;
  • 6) 0.400 240 463 785 6 × 2 = 0 + 0.800 480 927 571 2;
  • 7) 0.800 480 927 571 2 × 2 = 1 + 0.600 961 855 142 4;
  • 8) 0.600 961 855 142 4 × 2 = 1 + 0.201 923 710 284 8;
  • 9) 0.201 923 710 284 8 × 2 = 0 + 0.403 847 420 569 6;
  • 10) 0.403 847 420 569 6 × 2 = 0 + 0.807 694 841 139 2;
  • 11) 0.807 694 841 139 2 × 2 = 1 + 0.615 389 682 278 4;
  • 12) 0.615 389 682 278 4 × 2 = 1 + 0.230 779 364 556 8;
  • 13) 0.230 779 364 556 8 × 2 = 0 + 0.461 558 729 113 6;
  • 14) 0.461 558 729 113 6 × 2 = 0 + 0.923 117 458 227 2;
  • 15) 0.923 117 458 227 2 × 2 = 1 + 0.846 234 916 454 4;
  • 16) 0.846 234 916 454 4 × 2 = 1 + 0.692 469 832 908 8;
  • 17) 0.692 469 832 908 8 × 2 = 1 + 0.384 939 665 817 6;
  • 18) 0.384 939 665 817 6 × 2 = 0 + 0.769 879 331 635 2;
  • 19) 0.769 879 331 635 2 × 2 = 1 + 0.539 758 663 270 4;
  • 20) 0.539 758 663 270 4 × 2 = 1 + 0.079 517 326 540 8;
  • 21) 0.079 517 326 540 8 × 2 = 0 + 0.159 034 653 081 6;
  • 22) 0.159 034 653 081 6 × 2 = 0 + 0.318 069 306 163 2;
  • 23) 0.318 069 306 163 2 × 2 = 0 + 0.636 138 612 326 4;
  • 24) 0.636 138 612 326 4 × 2 = 1 + 0.272 277 224 652 8;
  • 25) 0.272 277 224 652 8 × 2 = 0 + 0.544 554 449 305 6;
  • 26) 0.544 554 449 305 6 × 2 = 1 + 0.089 108 898 611 2;
  • 27) 0.089 108 898 611 2 × 2 = 0 + 0.178 217 797 222 4;
  • 28) 0.178 217 797 222 4 × 2 = 0 + 0.356 435 594 444 8;
  • 29) 0.356 435 594 444 8 × 2 = 0 + 0.712 871 188 889 6;
  • 30) 0.712 871 188 889 6 × 2 = 1 + 0.425 742 377 779 2;
  • 31) 0.425 742 377 779 2 × 2 = 0 + 0.851 484 755 558 4;
  • 32) 0.851 484 755 558 4 × 2 = 1 + 0.702 969 511 116 8;
  • 33) 0.702 969 511 116 8 × 2 = 1 + 0.405 939 022 233 6;
  • 34) 0.405 939 022 233 6 × 2 = 0 + 0.811 878 044 467 2;
  • 35) 0.811 878 044 467 2 × 2 = 1 + 0.623 756 088 934 4;
  • 36) 0.623 756 088 934 4 × 2 = 1 + 0.247 512 177 868 8;
  • 37) 0.247 512 177 868 8 × 2 = 0 + 0.495 024 355 737 6;
  • 38) 0.495 024 355 737 6 × 2 = 0 + 0.990 048 711 475 2;
  • 39) 0.990 048 711 475 2 × 2 = 1 + 0.980 097 422 950 4;
  • 40) 0.980 097 422 950 4 × 2 = 1 + 0.960 194 845 900 8;
  • 41) 0.960 194 845 900 8 × 2 = 1 + 0.920 389 691 801 6;
  • 42) 0.920 389 691 801 6 × 2 = 1 + 0.840 779 383 603 2;
  • 43) 0.840 779 383 603 2 × 2 = 1 + 0.681 558 767 206 4;
  • 44) 0.681 558 767 206 4 × 2 = 1 + 0.363 117 534 412 8;
  • 45) 0.363 117 534 412 8 × 2 = 0 + 0.726 235 068 825 6;
  • 46) 0.726 235 068 825 6 × 2 = 1 + 0.452 470 137 651 2;
  • 47) 0.452 470 137 651 2 × 2 = 0 + 0.904 940 275 302 4;
  • 48) 0.904 940 275 302 4 × 2 = 1 + 0.809 880 550 604 8;
  • 49) 0.809 880 550 604 8 × 2 = 1 + 0.619 761 101 209 6;
  • 50) 0.619 761 101 209 6 × 2 = 1 + 0.239 522 202 419 2;
  • 51) 0.239 522 202 419 2 × 2 = 0 + 0.479 044 404 838 4;
  • 52) 0.479 044 404 838 4 × 2 = 0 + 0.958 088 809 676 8;
  • 53) 0.958 088 809 676 8 × 2 = 1 + 0.916 177 619 353 6;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.200 007 514 493 3(10) =


0.0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1(2)

5. Positive number before normalization:

29.200 007 514 493 3(10) =


1 1101.0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 4 positions to the left, so that only one non zero digit remains to the left of it:


29.200 007 514 493 3(10) =


1 1101.0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1(2) =


1 1101.0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1(2) × 20 =


1.1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1(2) × 24


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 4


Mantissa (not normalized):
1.1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1100 1


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


4 + 2(11-1) - 1 =


(4 + 1 023)(10) =


1 027(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 027 ÷ 2 = 513 + 1;
  • 513 ÷ 2 = 256 + 1;
  • 256 ÷ 2 = 128 + 0;
  • 128 ÷ 2 = 64 + 0;
  • 64 ÷ 2 = 32 + 0;
  • 32 ÷ 2 = 16 + 0;
  • 16 ÷ 2 = 8 + 0;
  • 8 ÷ 2 = 4 + 0;
  • 4 ÷ 2 = 2 + 0;
  • 2 ÷ 2 = 1 + 0;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1027(10) =


100 0000 0011(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101 1 1001 =


1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
100 0000 0011


Mantissa (52 bits) =
1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101


Decimal number 29.200 007 514 493 3 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 100 0000 0011 - 1101 0011 0011 0011 0011 1011 0001 0100 0101 1011 0011 1111 0101


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100