-0.000 084 993 381 976 778 52 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal -0.000 084 993 381 976 778 52(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
-0.000 084 993 381 976 778 52(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. Start with the positive version of the number:

|-0.000 084 993 381 976 778 52| = 0.000 084 993 381 976 778 52


2. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

3. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


4. Convert to binary (base 2) the fractional part: 0.000 084 993 381 976 778 52.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 084 993 381 976 778 52 × 2 = 0 + 0.000 169 986 763 953 557 04;
  • 2) 0.000 169 986 763 953 557 04 × 2 = 0 + 0.000 339 973 527 907 114 08;
  • 3) 0.000 339 973 527 907 114 08 × 2 = 0 + 0.000 679 947 055 814 228 16;
  • 4) 0.000 679 947 055 814 228 16 × 2 = 0 + 0.001 359 894 111 628 456 32;
  • 5) 0.001 359 894 111 628 456 32 × 2 = 0 + 0.002 719 788 223 256 912 64;
  • 6) 0.002 719 788 223 256 912 64 × 2 = 0 + 0.005 439 576 446 513 825 28;
  • 7) 0.005 439 576 446 513 825 28 × 2 = 0 + 0.010 879 152 893 027 650 56;
  • 8) 0.010 879 152 893 027 650 56 × 2 = 0 + 0.021 758 305 786 055 301 12;
  • 9) 0.021 758 305 786 055 301 12 × 2 = 0 + 0.043 516 611 572 110 602 24;
  • 10) 0.043 516 611 572 110 602 24 × 2 = 0 + 0.087 033 223 144 221 204 48;
  • 11) 0.087 033 223 144 221 204 48 × 2 = 0 + 0.174 066 446 288 442 408 96;
  • 12) 0.174 066 446 288 442 408 96 × 2 = 0 + 0.348 132 892 576 884 817 92;
  • 13) 0.348 132 892 576 884 817 92 × 2 = 0 + 0.696 265 785 153 769 635 84;
  • 14) 0.696 265 785 153 769 635 84 × 2 = 1 + 0.392 531 570 307 539 271 68;
  • 15) 0.392 531 570 307 539 271 68 × 2 = 0 + 0.785 063 140 615 078 543 36;
  • 16) 0.785 063 140 615 078 543 36 × 2 = 1 + 0.570 126 281 230 157 086 72;
  • 17) 0.570 126 281 230 157 086 72 × 2 = 1 + 0.140 252 562 460 314 173 44;
  • 18) 0.140 252 562 460 314 173 44 × 2 = 0 + 0.280 505 124 920 628 346 88;
  • 19) 0.280 505 124 920 628 346 88 × 2 = 0 + 0.561 010 249 841 256 693 76;
  • 20) 0.561 010 249 841 256 693 76 × 2 = 1 + 0.122 020 499 682 513 387 52;
  • 21) 0.122 020 499 682 513 387 52 × 2 = 0 + 0.244 040 999 365 026 775 04;
  • 22) 0.244 040 999 365 026 775 04 × 2 = 0 + 0.488 081 998 730 053 550 08;
  • 23) 0.488 081 998 730 053 550 08 × 2 = 0 + 0.976 163 997 460 107 100 16;
  • 24) 0.976 163 997 460 107 100 16 × 2 = 1 + 0.952 327 994 920 214 200 32;
  • 25) 0.952 327 994 920 214 200 32 × 2 = 1 + 0.904 655 989 840 428 400 64;
  • 26) 0.904 655 989 840 428 400 64 × 2 = 1 + 0.809 311 979 680 856 801 28;
  • 27) 0.809 311 979 680 856 801 28 × 2 = 1 + 0.618 623 959 361 713 602 56;
  • 28) 0.618 623 959 361 713 602 56 × 2 = 1 + 0.237 247 918 723 427 205 12;
  • 29) 0.237 247 918 723 427 205 12 × 2 = 0 + 0.474 495 837 446 854 410 24;
  • 30) 0.474 495 837 446 854 410 24 × 2 = 0 + 0.948 991 674 893 708 820 48;
  • 31) 0.948 991 674 893 708 820 48 × 2 = 1 + 0.897 983 349 787 417 640 96;
  • 32) 0.897 983 349 787 417 640 96 × 2 = 1 + 0.795 966 699 574 835 281 92;
  • 33) 0.795 966 699 574 835 281 92 × 2 = 1 + 0.591 933 399 149 670 563 84;
  • 34) 0.591 933 399 149 670 563 84 × 2 = 1 + 0.183 866 798 299 341 127 68;
  • 35) 0.183 866 798 299 341 127 68 × 2 = 0 + 0.367 733 596 598 682 255 36;
  • 36) 0.367 733 596 598 682 255 36 × 2 = 0 + 0.735 467 193 197 364 510 72;
  • 37) 0.735 467 193 197 364 510 72 × 2 = 1 + 0.470 934 386 394 729 021 44;
  • 38) 0.470 934 386 394 729 021 44 × 2 = 0 + 0.941 868 772 789 458 042 88;
  • 39) 0.941 868 772 789 458 042 88 × 2 = 1 + 0.883 737 545 578 916 085 76;
  • 40) 0.883 737 545 578 916 085 76 × 2 = 1 + 0.767 475 091 157 832 171 52;
  • 41) 0.767 475 091 157 832 171 52 × 2 = 1 + 0.534 950 182 315 664 343 04;
  • 42) 0.534 950 182 315 664 343 04 × 2 = 1 + 0.069 900 364 631 328 686 08;
  • 43) 0.069 900 364 631 328 686 08 × 2 = 0 + 0.139 800 729 262 657 372 16;
  • 44) 0.139 800 729 262 657 372 16 × 2 = 0 + 0.279 601 458 525 314 744 32;
  • 45) 0.279 601 458 525 314 744 32 × 2 = 0 + 0.559 202 917 050 629 488 64;
  • 46) 0.559 202 917 050 629 488 64 × 2 = 1 + 0.118 405 834 101 258 977 28;
  • 47) 0.118 405 834 101 258 977 28 × 2 = 0 + 0.236 811 668 202 517 954 56;
  • 48) 0.236 811 668 202 517 954 56 × 2 = 0 + 0.473 623 336 405 035 909 12;
  • 49) 0.473 623 336 405 035 909 12 × 2 = 0 + 0.947 246 672 810 071 818 24;
  • 50) 0.947 246 672 810 071 818 24 × 2 = 1 + 0.894 493 345 620 143 636 48;
  • 51) 0.894 493 345 620 143 636 48 × 2 = 1 + 0.788 986 691 240 287 272 96;
  • 52) 0.788 986 691 240 287 272 96 × 2 = 1 + 0.577 973 382 480 574 545 92;
  • 53) 0.577 973 382 480 574 545 92 × 2 = 1 + 0.155 946 764 961 149 091 84;
  • 54) 0.155 946 764 961 149 091 84 × 2 = 0 + 0.311 893 529 922 298 183 68;
  • 55) 0.311 893 529 922 298 183 68 × 2 = 0 + 0.623 787 059 844 596 367 36;
  • 56) 0.623 787 059 844 596 367 36 × 2 = 1 + 0.247 574 119 689 192 734 72;
  • 57) 0.247 574 119 689 192 734 72 × 2 = 0 + 0.495 148 239 378 385 469 44;
  • 58) 0.495 148 239 378 385 469 44 × 2 = 0 + 0.990 296 478 756 770 938 88;
  • 59) 0.990 296 478 756 770 938 88 × 2 = 1 + 0.980 592 957 513 541 877 76;
  • 60) 0.980 592 957 513 541 877 76 × 2 = 1 + 0.961 185 915 027 083 755 52;
  • 61) 0.961 185 915 027 083 755 52 × 2 = 1 + 0.922 371 830 054 167 511 04;
  • 62) 0.922 371 830 054 167 511 04 × 2 = 1 + 0.844 743 660 108 335 022 08;
  • 63) 0.844 743 660 108 335 022 08 × 2 = 1 + 0.689 487 320 216 670 044 16;
  • 64) 0.689 487 320 216 670 044 16 × 2 = 1 + 0.378 974 640 433 340 088 32;
  • 65) 0.378 974 640 433 340 088 32 × 2 = 0 + 0.757 949 280 866 680 176 64;
  • 66) 0.757 949 280 866 680 176 64 × 2 = 1 + 0.515 898 561 733 360 353 28;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


5. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 084 993 381 976 778 52(10) =


0.0000 0000 0000 0101 1001 0001 1111 0011 1100 1011 1100 0100 0111 1001 0011 1111 01(2)

6. Positive number before normalization:

0.000 084 993 381 976 778 52(10) =


0.0000 0000 0000 0101 1001 0001 1111 0011 1100 1011 1100 0100 0111 1001 0011 1111 01(2)

7. Normalize the binary representation of the number.

Shift the decimal mark 14 positions to the right, so that only one non zero digit remains to the left of it:


0.000 084 993 381 976 778 52(10) =


0.0000 0000 0000 0101 1001 0001 1111 0011 1100 1011 1100 0100 0111 1001 0011 1111 01(2) =


0.0000 0000 0000 0101 1001 0001 1111 0011 1100 1011 1100 0100 0111 1001 0011 1111 01(2) × 20 =


1.0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101(2) × 2-14


8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 1 (a negative number)


Exponent (unadjusted): -14


Mantissa (not normalized):
1.0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101


9. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-14 + 2(11-1) - 1 =


(-14 + 1 023)(10) =


1 009(10)


10. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 009 ÷ 2 = 504 + 1;
  • 504 ÷ 2 = 252 + 0;
  • 252 ÷ 2 = 126 + 0;
  • 126 ÷ 2 = 63 + 0;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

11. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1009(10) =


011 1111 0001(2)


12. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101 =


0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101


13. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
1 (a negative number)


Exponent (11 bits) =
011 1111 0001


Mantissa (52 bits) =
0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101


Decimal number -0.000 084 993 381 976 778 52 converted to 64 bit double precision IEEE 754 binary floating point representation:

1 - 011 1111 0001 - 0110 0100 0111 1100 1111 0010 1111 0001 0001 1110 0100 1111 1101


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100