0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996 × 2 = 0 + 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 992;
  • 2) 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 992 × 2 = 0 + 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 984;
  • 3) 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 984 × 2 = 1 + 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 968;
  • 4) 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 968 × 2 = 0 + 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 207 936;
  • 5) 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 207 936 × 2 = 0 + 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 415 872;
  • 6) 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 415 872 × 2 = 1 + 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 831 744;
  • 7) 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 831 744 × 2 = 0 + 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 663 488;
  • 8) 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 663 488 × 2 = 1 + 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 326 976;
  • 9) 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 326 976 × 2 = 0 + 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 653 952;
  • 10) 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 653 952 × 2 = 1 + 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 307 904;
  • 11) 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 307 904 × 2 = 1 + 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 615 808;
  • 12) 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 615 808 × 2 = 1 + 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 005 231 616;
  • 13) 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 005 231 616 × 2 = 1 + 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 010 463 232;
  • 14) 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 010 463 232 × 2 = 1 + 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 020 926 464;
  • 15) 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 020 926 464 × 2 = 0 + 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 041 852 928;
  • 16) 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 041 852 928 × 2 = 1 + 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 083 705 856;
  • 17) 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 083 705 856 × 2 = 1 + 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 167 411 712;
  • 18) 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 167 411 712 × 2 = 0 + 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 334 823 424;
  • 19) 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 334 823 424 × 2 = 0 + 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 669 646 848;
  • 20) 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 669 646 848 × 2 = 0 + 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 339 293 696;
  • 21) 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 339 293 696 × 2 = 0 + 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 678 587 392;
  • 22) 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 678 587 392 × 2 = 1 + 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 357 174 784;
  • 23) 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 357 174 784 × 2 = 1 + 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 714 349 568;
  • 24) 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 714 349 568 × 2 = 0 + 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 965 428 699 136;
  • 25) 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 965 428 699 136 × 2 = 0 + 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 930 857 398 272;
  • 26) 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 930 857 398 272 × 2 = 1 + 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 861 714 796 544;
  • 27) 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 861 714 796 544 × 2 = 1 + 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 723 429 593 088;
  • 28) 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 723 429 593 088 × 2 = 0 + 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 446 859 186 176;
  • 29) 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 446 859 186 176 × 2 = 0 + 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 893 718 372 352;
  • 30) 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 893 718 372 352 × 2 = 1 + 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 787 436 744 704;
  • 31) 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 787 436 744 704 × 2 = 1 + 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 574 873 489 408;
  • 32) 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 574 873 489 408 × 2 = 0 + 0.011 975 153 774 305 734 986 845 826 627 305 209 451 255 149 746 978 816;
  • 33) 0.011 975 153 774 305 734 986 845 826 627 305 209 451 255 149 746 978 816 × 2 = 0 + 0.023 950 307 548 611 469 973 691 653 254 610 418 902 510 299 493 957 632;
  • 34) 0.023 950 307 548 611 469 973 691 653 254 610 418 902 510 299 493 957 632 × 2 = 0 + 0.047 900 615 097 222 939 947 383 306 509 220 837 805 020 598 987 915 264;
  • 35) 0.047 900 615 097 222 939 947 383 306 509 220 837 805 020 598 987 915 264 × 2 = 0 + 0.095 801 230 194 445 879 894 766 613 018 441 675 610 041 197 975 830 528;
  • 36) 0.095 801 230 194 445 879 894 766 613 018 441 675 610 041 197 975 830 528 × 2 = 0 + 0.191 602 460 388 891 759 789 533 226 036 883 351 220 082 395 951 661 056;
  • 37) 0.191 602 460 388 891 759 789 533 226 036 883 351 220 082 395 951 661 056 × 2 = 0 + 0.383 204 920 777 783 519 579 066 452 073 766 702 440 164 791 903 322 112;
  • 38) 0.383 204 920 777 783 519 579 066 452 073 766 702 440 164 791 903 322 112 × 2 = 0 + 0.766 409 841 555 567 039 158 132 904 147 533 404 880 329 583 806 644 224;
  • 39) 0.766 409 841 555 567 039 158 132 904 147 533 404 880 329 583 806 644 224 × 2 = 1 + 0.532 819 683 111 134 078 316 265 808 295 066 809 760 659 167 613 288 448;
  • 40) 0.532 819 683 111 134 078 316 265 808 295 066 809 760 659 167 613 288 448 × 2 = 1 + 0.065 639 366 222 268 156 632 531 616 590 133 619 521 318 335 226 576 896;
  • 41) 0.065 639 366 222 268 156 632 531 616 590 133 619 521 318 335 226 576 896 × 2 = 0 + 0.131 278 732 444 536 313 265 063 233 180 267 239 042 636 670 453 153 792;
  • 42) 0.131 278 732 444 536 313 265 063 233 180 267 239 042 636 670 453 153 792 × 2 = 0 + 0.262 557 464 889 072 626 530 126 466 360 534 478 085 273 340 906 307 584;
  • 43) 0.262 557 464 889 072 626 530 126 466 360 534 478 085 273 340 906 307 584 × 2 = 0 + 0.525 114 929 778 145 253 060 252 932 721 068 956 170 546 681 812 615 168;
  • 44) 0.525 114 929 778 145 253 060 252 932 721 068 956 170 546 681 812 615 168 × 2 = 1 + 0.050 229 859 556 290 506 120 505 865 442 137 912 341 093 363 625 230 336;
  • 45) 0.050 229 859 556 290 506 120 505 865 442 137 912 341 093 363 625 230 336 × 2 = 0 + 0.100 459 719 112 581 012 241 011 730 884 275 824 682 186 727 250 460 672;
  • 46) 0.100 459 719 112 581 012 241 011 730 884 275 824 682 186 727 250 460 672 × 2 = 0 + 0.200 919 438 225 162 024 482 023 461 768 551 649 364 373 454 500 921 344;
  • 47) 0.200 919 438 225 162 024 482 023 461 768 551 649 364 373 454 500 921 344 × 2 = 0 + 0.401 838 876 450 324 048 964 046 923 537 103 298 728 746 909 001 842 688;
  • 48) 0.401 838 876 450 324 048 964 046 923 537 103 298 728 746 909 001 842 688 × 2 = 0 + 0.803 677 752 900 648 097 928 093 847 074 206 597 457 493 818 003 685 376;
  • 49) 0.803 677 752 900 648 097 928 093 847 074 206 597 457 493 818 003 685 376 × 2 = 1 + 0.607 355 505 801 296 195 856 187 694 148 413 194 914 987 636 007 370 752;
  • 50) 0.607 355 505 801 296 195 856 187 694 148 413 194 914 987 636 007 370 752 × 2 = 1 + 0.214 711 011 602 592 391 712 375 388 296 826 389 829 975 272 014 741 504;
  • 51) 0.214 711 011 602 592 391 712 375 388 296 826 389 829 975 272 014 741 504 × 2 = 0 + 0.429 422 023 205 184 783 424 750 776 593 652 779 659 950 544 029 483 008;
  • 52) 0.429 422 023 205 184 783 424 750 776 593 652 779 659 950 544 029 483 008 × 2 = 0 + 0.858 844 046 410 369 566 849 501 553 187 305 559 319 901 088 058 966 016;
  • 53) 0.858 844 046 410 369 566 849 501 553 187 305 559 319 901 088 058 966 016 × 2 = 1 + 0.717 688 092 820 739 133 699 003 106 374 611 118 639 802 176 117 932 032;
  • 54) 0.717 688 092 820 739 133 699 003 106 374 611 118 639 802 176 117 932 032 × 2 = 1 + 0.435 376 185 641 478 267 398 006 212 749 222 237 279 604 352 235 864 064;
  • 55) 0.435 376 185 641 478 267 398 006 212 749 222 237 279 604 352 235 864 064 × 2 = 0 + 0.870 752 371 282 956 534 796 012 425 498 444 474 559 208 704 471 728 128;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

5. Positive number before normalization:

0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 3 positions to the right, so that only one non zero digit remains to the left of it:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) × 20 =


1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110(2) × 2-3


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -3


Mantissa (not normalized):
1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-3 + 2(11-1) - 1 =


(-3 + 1 023)(10) =


1 020(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 020 ÷ 2 = 510 + 0;
  • 510 ÷ 2 = 255 + 0;
  • 255 ÷ 2 = 127 + 1;
  • 127 ÷ 2 = 63 + 1;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1020(10) =


011 1111 1100(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110 =


0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1111 1100


Mantissa (52 bits) =
0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


Decimal number 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 996 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1111 1100 - 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100