0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926 × 2 = 0 + 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 852;
  • 2) 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 852 × 2 = 0 + 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 704;
  • 3) 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 704 × 2 = 1 + 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 408;
  • 4) 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 408 × 2 = 0 + 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 206 816;
  • 5) 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 206 816 × 2 = 0 + 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 413 632;
  • 6) 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 413 632 × 2 = 1 + 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 827 264;
  • 7) 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 827 264 × 2 = 0 + 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 654 528;
  • 8) 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 654 528 × 2 = 1 + 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 309 056;
  • 9) 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 309 056 × 2 = 0 + 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 618 112;
  • 10) 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 618 112 × 2 = 1 + 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 236 224;
  • 11) 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 236 224 × 2 = 1 + 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 472 448;
  • 12) 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 472 448 × 2 = 1 + 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 004 944 896;
  • 13) 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 004 944 896 × 2 = 1 + 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 009 889 792;
  • 14) 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 009 889 792 × 2 = 1 + 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 019 779 584;
  • 15) 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 019 779 584 × 2 = 0 + 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 039 559 168;
  • 16) 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 039 559 168 × 2 = 1 + 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 079 118 336;
  • 17) 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 079 118 336 × 2 = 1 + 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 158 236 672;
  • 18) 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 158 236 672 × 2 = 0 + 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 316 473 344;
  • 19) 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 316 473 344 × 2 = 0 + 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 632 946 688;
  • 20) 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 632 946 688 × 2 = 0 + 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 265 893 376;
  • 21) 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 265 893 376 × 2 = 0 + 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 531 786 752;
  • 22) 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 531 786 752 × 2 = 1 + 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 063 573 504;
  • 23) 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 063 573 504 × 2 = 1 + 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 127 147 008;
  • 24) 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 127 147 008 × 2 = 0 + 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 964 254 294 016;
  • 25) 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 964 254 294 016 × 2 = 0 + 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 928 508 588 032;
  • 26) 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 928 508 588 032 × 2 = 1 + 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 857 017 176 064;
  • 27) 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 857 017 176 064 × 2 = 1 + 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 714 034 352 128;
  • 28) 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 714 034 352 128 × 2 = 0 + 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 428 068 704 256;
  • 29) 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 428 068 704 256 × 2 = 0 + 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 856 137 408 512;
  • 30) 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 856 137 408 512 × 2 = 1 + 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 712 274 817 024;
  • 31) 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 712 274 817 024 × 2 = 1 + 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 424 549 634 048;
  • 32) 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 424 549 634 048 × 2 = 0 + 0.011 975 153 774 305 734 986 845 826 627 305 209 451 254 849 099 268 096;
  • 33) 0.011 975 153 774 305 734 986 845 826 627 305 209 451 254 849 099 268 096 × 2 = 0 + 0.023 950 307 548 611 469 973 691 653 254 610 418 902 509 698 198 536 192;
  • 34) 0.023 950 307 548 611 469 973 691 653 254 610 418 902 509 698 198 536 192 × 2 = 0 + 0.047 900 615 097 222 939 947 383 306 509 220 837 805 019 396 397 072 384;
  • 35) 0.047 900 615 097 222 939 947 383 306 509 220 837 805 019 396 397 072 384 × 2 = 0 + 0.095 801 230 194 445 879 894 766 613 018 441 675 610 038 792 794 144 768;
  • 36) 0.095 801 230 194 445 879 894 766 613 018 441 675 610 038 792 794 144 768 × 2 = 0 + 0.191 602 460 388 891 759 789 533 226 036 883 351 220 077 585 588 289 536;
  • 37) 0.191 602 460 388 891 759 789 533 226 036 883 351 220 077 585 588 289 536 × 2 = 0 + 0.383 204 920 777 783 519 579 066 452 073 766 702 440 155 171 176 579 072;
  • 38) 0.383 204 920 777 783 519 579 066 452 073 766 702 440 155 171 176 579 072 × 2 = 0 + 0.766 409 841 555 567 039 158 132 904 147 533 404 880 310 342 353 158 144;
  • 39) 0.766 409 841 555 567 039 158 132 904 147 533 404 880 310 342 353 158 144 × 2 = 1 + 0.532 819 683 111 134 078 316 265 808 295 066 809 760 620 684 706 316 288;
  • 40) 0.532 819 683 111 134 078 316 265 808 295 066 809 760 620 684 706 316 288 × 2 = 1 + 0.065 639 366 222 268 156 632 531 616 590 133 619 521 241 369 412 632 576;
  • 41) 0.065 639 366 222 268 156 632 531 616 590 133 619 521 241 369 412 632 576 × 2 = 0 + 0.131 278 732 444 536 313 265 063 233 180 267 239 042 482 738 825 265 152;
  • 42) 0.131 278 732 444 536 313 265 063 233 180 267 239 042 482 738 825 265 152 × 2 = 0 + 0.262 557 464 889 072 626 530 126 466 360 534 478 084 965 477 650 530 304;
  • 43) 0.262 557 464 889 072 626 530 126 466 360 534 478 084 965 477 650 530 304 × 2 = 0 + 0.525 114 929 778 145 253 060 252 932 721 068 956 169 930 955 301 060 608;
  • 44) 0.525 114 929 778 145 253 060 252 932 721 068 956 169 930 955 301 060 608 × 2 = 1 + 0.050 229 859 556 290 506 120 505 865 442 137 912 339 861 910 602 121 216;
  • 45) 0.050 229 859 556 290 506 120 505 865 442 137 912 339 861 910 602 121 216 × 2 = 0 + 0.100 459 719 112 581 012 241 011 730 884 275 824 679 723 821 204 242 432;
  • 46) 0.100 459 719 112 581 012 241 011 730 884 275 824 679 723 821 204 242 432 × 2 = 0 + 0.200 919 438 225 162 024 482 023 461 768 551 649 359 447 642 408 484 864;
  • 47) 0.200 919 438 225 162 024 482 023 461 768 551 649 359 447 642 408 484 864 × 2 = 0 + 0.401 838 876 450 324 048 964 046 923 537 103 298 718 895 284 816 969 728;
  • 48) 0.401 838 876 450 324 048 964 046 923 537 103 298 718 895 284 816 969 728 × 2 = 0 + 0.803 677 752 900 648 097 928 093 847 074 206 597 437 790 569 633 939 456;
  • 49) 0.803 677 752 900 648 097 928 093 847 074 206 597 437 790 569 633 939 456 × 2 = 1 + 0.607 355 505 801 296 195 856 187 694 148 413 194 875 581 139 267 878 912;
  • 50) 0.607 355 505 801 296 195 856 187 694 148 413 194 875 581 139 267 878 912 × 2 = 1 + 0.214 711 011 602 592 391 712 375 388 296 826 389 751 162 278 535 757 824;
  • 51) 0.214 711 011 602 592 391 712 375 388 296 826 389 751 162 278 535 757 824 × 2 = 0 + 0.429 422 023 205 184 783 424 750 776 593 652 779 502 324 557 071 515 648;
  • 52) 0.429 422 023 205 184 783 424 750 776 593 652 779 502 324 557 071 515 648 × 2 = 0 + 0.858 844 046 410 369 566 849 501 553 187 305 559 004 649 114 143 031 296;
  • 53) 0.858 844 046 410 369 566 849 501 553 187 305 559 004 649 114 143 031 296 × 2 = 1 + 0.717 688 092 820 739 133 699 003 106 374 611 118 009 298 228 286 062 592;
  • 54) 0.717 688 092 820 739 133 699 003 106 374 611 118 009 298 228 286 062 592 × 2 = 1 + 0.435 376 185 641 478 267 398 006 212 749 222 236 018 596 456 572 125 184;
  • 55) 0.435 376 185 641 478 267 398 006 212 749 222 236 018 596 456 572 125 184 × 2 = 0 + 0.870 752 371 282 956 534 796 012 425 498 444 472 037 192 913 144 250 368;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

5. Positive number before normalization:

0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 3 positions to the right, so that only one non zero digit remains to the left of it:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) × 20 =


1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110(2) × 2-3


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -3


Mantissa (not normalized):
1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-3 + 2(11-1) - 1 =


(-3 + 1 023)(10) =


1 020(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 020 ÷ 2 = 510 + 0;
  • 510 ÷ 2 = 255 + 0;
  • 255 ÷ 2 = 127 + 1;
  • 127 ÷ 2 = 63 + 1;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1020(10) =


011 1111 1100(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110 =


0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1111 1100


Mantissa (52 bits) =
0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


Decimal number 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 926 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1111 1100 - 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110

How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100