-0.000 025 693 776 151 095 2 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal -0.000 025 693 776 151 095 2(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
-0.000 025 693 776 151 095 2(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. Start with the positive version of the number:

|-0.000 025 693 776 151 095 2| = 0.000 025 693 776 151 095 2


2. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

3. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


4. Convert to binary (base 2) the fractional part: 0.000 025 693 776 151 095 2.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 025 693 776 151 095 2 × 2 = 0 + 0.000 051 387 552 302 190 4;
  • 2) 0.000 051 387 552 302 190 4 × 2 = 0 + 0.000 102 775 104 604 380 8;
  • 3) 0.000 102 775 104 604 380 8 × 2 = 0 + 0.000 205 550 209 208 761 6;
  • 4) 0.000 205 550 209 208 761 6 × 2 = 0 + 0.000 411 100 418 417 523 2;
  • 5) 0.000 411 100 418 417 523 2 × 2 = 0 + 0.000 822 200 836 835 046 4;
  • 6) 0.000 822 200 836 835 046 4 × 2 = 0 + 0.001 644 401 673 670 092 8;
  • 7) 0.001 644 401 673 670 092 8 × 2 = 0 + 0.003 288 803 347 340 185 6;
  • 8) 0.003 288 803 347 340 185 6 × 2 = 0 + 0.006 577 606 694 680 371 2;
  • 9) 0.006 577 606 694 680 371 2 × 2 = 0 + 0.013 155 213 389 360 742 4;
  • 10) 0.013 155 213 389 360 742 4 × 2 = 0 + 0.026 310 426 778 721 484 8;
  • 11) 0.026 310 426 778 721 484 8 × 2 = 0 + 0.052 620 853 557 442 969 6;
  • 12) 0.052 620 853 557 442 969 6 × 2 = 0 + 0.105 241 707 114 885 939 2;
  • 13) 0.105 241 707 114 885 939 2 × 2 = 0 + 0.210 483 414 229 771 878 4;
  • 14) 0.210 483 414 229 771 878 4 × 2 = 0 + 0.420 966 828 459 543 756 8;
  • 15) 0.420 966 828 459 543 756 8 × 2 = 0 + 0.841 933 656 919 087 513 6;
  • 16) 0.841 933 656 919 087 513 6 × 2 = 1 + 0.683 867 313 838 175 027 2;
  • 17) 0.683 867 313 838 175 027 2 × 2 = 1 + 0.367 734 627 676 350 054 4;
  • 18) 0.367 734 627 676 350 054 4 × 2 = 0 + 0.735 469 255 352 700 108 8;
  • 19) 0.735 469 255 352 700 108 8 × 2 = 1 + 0.470 938 510 705 400 217 6;
  • 20) 0.470 938 510 705 400 217 6 × 2 = 0 + 0.941 877 021 410 800 435 2;
  • 21) 0.941 877 021 410 800 435 2 × 2 = 1 + 0.883 754 042 821 600 870 4;
  • 22) 0.883 754 042 821 600 870 4 × 2 = 1 + 0.767 508 085 643 201 740 8;
  • 23) 0.767 508 085 643 201 740 8 × 2 = 1 + 0.535 016 171 286 403 481 6;
  • 24) 0.535 016 171 286 403 481 6 × 2 = 1 + 0.070 032 342 572 806 963 2;
  • 25) 0.070 032 342 572 806 963 2 × 2 = 0 + 0.140 064 685 145 613 926 4;
  • 26) 0.140 064 685 145 613 926 4 × 2 = 0 + 0.280 129 370 291 227 852 8;
  • 27) 0.280 129 370 291 227 852 8 × 2 = 0 + 0.560 258 740 582 455 705 6;
  • 28) 0.560 258 740 582 455 705 6 × 2 = 1 + 0.120 517 481 164 911 411 2;
  • 29) 0.120 517 481 164 911 411 2 × 2 = 0 + 0.241 034 962 329 822 822 4;
  • 30) 0.241 034 962 329 822 822 4 × 2 = 0 + 0.482 069 924 659 645 644 8;
  • 31) 0.482 069 924 659 645 644 8 × 2 = 0 + 0.964 139 849 319 291 289 6;
  • 32) 0.964 139 849 319 291 289 6 × 2 = 1 + 0.928 279 698 638 582 579 2;
  • 33) 0.928 279 698 638 582 579 2 × 2 = 1 + 0.856 559 397 277 165 158 4;
  • 34) 0.856 559 397 277 165 158 4 × 2 = 1 + 0.713 118 794 554 330 316 8;
  • 35) 0.713 118 794 554 330 316 8 × 2 = 1 + 0.426 237 589 108 660 633 6;
  • 36) 0.426 237 589 108 660 633 6 × 2 = 0 + 0.852 475 178 217 321 267 2;
  • 37) 0.852 475 178 217 321 267 2 × 2 = 1 + 0.704 950 356 434 642 534 4;
  • 38) 0.704 950 356 434 642 534 4 × 2 = 1 + 0.409 900 712 869 285 068 8;
  • 39) 0.409 900 712 869 285 068 8 × 2 = 0 + 0.819 801 425 738 570 137 6;
  • 40) 0.819 801 425 738 570 137 6 × 2 = 1 + 0.639 602 851 477 140 275 2;
  • 41) 0.639 602 851 477 140 275 2 × 2 = 1 + 0.279 205 702 954 280 550 4;
  • 42) 0.279 205 702 954 280 550 4 × 2 = 0 + 0.558 411 405 908 561 100 8;
  • 43) 0.558 411 405 908 561 100 8 × 2 = 1 + 0.116 822 811 817 122 201 6;
  • 44) 0.116 822 811 817 122 201 6 × 2 = 0 + 0.233 645 623 634 244 403 2;
  • 45) 0.233 645 623 634 244 403 2 × 2 = 0 + 0.467 291 247 268 488 806 4;
  • 46) 0.467 291 247 268 488 806 4 × 2 = 0 + 0.934 582 494 536 977 612 8;
  • 47) 0.934 582 494 536 977 612 8 × 2 = 1 + 0.869 164 989 073 955 225 6;
  • 48) 0.869 164 989 073 955 225 6 × 2 = 1 + 0.738 329 978 147 910 451 2;
  • 49) 0.738 329 978 147 910 451 2 × 2 = 1 + 0.476 659 956 295 820 902 4;
  • 50) 0.476 659 956 295 820 902 4 × 2 = 0 + 0.953 319 912 591 641 804 8;
  • 51) 0.953 319 912 591 641 804 8 × 2 = 1 + 0.906 639 825 183 283 609 6;
  • 52) 0.906 639 825 183 283 609 6 × 2 = 1 + 0.813 279 650 366 567 219 2;
  • 53) 0.813 279 650 366 567 219 2 × 2 = 1 + 0.626 559 300 733 134 438 4;
  • 54) 0.626 559 300 733 134 438 4 × 2 = 1 + 0.253 118 601 466 268 876 8;
  • 55) 0.253 118 601 466 268 876 8 × 2 = 0 + 0.506 237 202 932 537 753 6;
  • 56) 0.506 237 202 932 537 753 6 × 2 = 1 + 0.012 474 405 865 075 507 2;
  • 57) 0.012 474 405 865 075 507 2 × 2 = 0 + 0.024 948 811 730 151 014 4;
  • 58) 0.024 948 811 730 151 014 4 × 2 = 0 + 0.049 897 623 460 302 028 8;
  • 59) 0.049 897 623 460 302 028 8 × 2 = 0 + 0.099 795 246 920 604 057 6;
  • 60) 0.099 795 246 920 604 057 6 × 2 = 0 + 0.199 590 493 841 208 115 2;
  • 61) 0.199 590 493 841 208 115 2 × 2 = 0 + 0.399 180 987 682 416 230 4;
  • 62) 0.399 180 987 682 416 230 4 × 2 = 0 + 0.798 361 975 364 832 460 8;
  • 63) 0.798 361 975 364 832 460 8 × 2 = 1 + 0.596 723 950 729 664 921 6;
  • 64) 0.596 723 950 729 664 921 6 × 2 = 1 + 0.193 447 901 459 329 843 2;
  • 65) 0.193 447 901 459 329 843 2 × 2 = 0 + 0.386 895 802 918 659 686 4;
  • 66) 0.386 895 802 918 659 686 4 × 2 = 0 + 0.773 791 605 837 319 372 8;
  • 67) 0.773 791 605 837 319 372 8 × 2 = 1 + 0.547 583 211 674 638 745 6;
  • 68) 0.547 583 211 674 638 745 6 × 2 = 1 + 0.095 166 423 349 277 491 2;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


5. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 025 693 776 151 095 2(10) =


0.0000 0000 0000 0001 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011(2)

6. Positive number before normalization:

0.000 025 693 776 151 095 2(10) =


0.0000 0000 0000 0001 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011(2)

7. Normalize the binary representation of the number.

Shift the decimal mark 16 positions to the right, so that only one non zero digit remains to the left of it:


0.000 025 693 776 151 095 2(10) =


0.0000 0000 0000 0001 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011(2) =


0.0000 0000 0000 0001 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011(2) × 20 =


1.1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011(2) × 2-16


8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 1 (a negative number)


Exponent (unadjusted): -16


Mantissa (not normalized):
1.1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011


9. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-16 + 2(11-1) - 1 =


(-16 + 1 023)(10) =


1 007(10)


10. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 007 ÷ 2 = 503 + 1;
  • 503 ÷ 2 = 251 + 1;
  • 251 ÷ 2 = 125 + 1;
  • 125 ÷ 2 = 62 + 1;
  • 62 ÷ 2 = 31 + 0;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

11. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1007(10) =


011 1110 1111(2)


12. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011 =


1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011


13. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
1 (a negative number)


Exponent (11 bits) =
011 1110 1111


Mantissa (52 bits) =
1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011


Decimal number -0.000 025 693 776 151 095 2 converted to 64 bit double precision IEEE 754 binary floating point representation:

1 - 011 1110 1111 - 1010 1111 0001 0001 1110 1101 1010 0011 1011 1101 0000 0011 0011


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100