0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918 × 2 = 0 + 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 836;
  • 2) 0.292 893 218 813 452 475 599 155 637 895 150 960 715 164 062 311 525 836 × 2 = 0 + 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 672;
  • 3) 0.585 786 437 626 904 951 198 311 275 790 301 921 430 328 124 623 051 672 × 2 = 1 + 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 344;
  • 4) 0.171 572 875 253 809 902 396 622 551 580 603 842 860 656 249 246 103 344 × 2 = 0 + 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 206 688;
  • 5) 0.343 145 750 507 619 804 793 245 103 161 207 685 721 312 498 492 206 688 × 2 = 0 + 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 413 376;
  • 6) 0.686 291 501 015 239 609 586 490 206 322 415 371 442 624 996 984 413 376 × 2 = 1 + 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 826 752;
  • 7) 0.372 583 002 030 479 219 172 980 412 644 830 742 885 249 993 968 826 752 × 2 = 0 + 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 653 504;
  • 8) 0.745 166 004 060 958 438 345 960 825 289 661 485 770 499 987 937 653 504 × 2 = 1 + 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 307 008;
  • 9) 0.490 332 008 121 916 876 691 921 650 579 322 971 540 999 975 875 307 008 × 2 = 0 + 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 614 016;
  • 10) 0.980 664 016 243 833 753 383 843 301 158 645 943 081 999 951 750 614 016 × 2 = 1 + 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 228 032;
  • 11) 0.961 328 032 487 667 506 767 686 602 317 291 886 163 999 903 501 228 032 × 2 = 1 + 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 456 064;
  • 12) 0.922 656 064 975 335 013 535 373 204 634 583 772 327 999 807 002 456 064 × 2 = 1 + 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 004 912 128;
  • 13) 0.845 312 129 950 670 027 070 746 409 269 167 544 655 999 614 004 912 128 × 2 = 1 + 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 009 824 256;
  • 14) 0.690 624 259 901 340 054 141 492 818 538 335 089 311 999 228 009 824 256 × 2 = 1 + 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 019 648 512;
  • 15) 0.381 248 519 802 680 108 282 985 637 076 670 178 623 998 456 019 648 512 × 2 = 0 + 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 039 297 024;
  • 16) 0.762 497 039 605 360 216 565 971 274 153 340 357 247 996 912 039 297 024 × 2 = 1 + 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 078 594 048;
  • 17) 0.524 994 079 210 720 433 131 942 548 306 680 714 495 993 824 078 594 048 × 2 = 1 + 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 157 188 096;
  • 18) 0.049 988 158 421 440 866 263 885 096 613 361 428 991 987 648 157 188 096 × 2 = 0 + 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 314 376 192;
  • 19) 0.099 976 316 842 881 732 527 770 193 226 722 857 983 975 296 314 376 192 × 2 = 0 + 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 628 752 384;
  • 20) 0.199 952 633 685 763 465 055 540 386 453 445 715 967 950 592 628 752 384 × 2 = 0 + 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 257 504 768;
  • 21) 0.399 905 267 371 526 930 111 080 772 906 891 431 935 901 185 257 504 768 × 2 = 0 + 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 515 009 536;
  • 22) 0.799 810 534 743 053 860 222 161 545 813 782 863 871 802 370 515 009 536 × 2 = 1 + 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 030 019 072;
  • 23) 0.599 621 069 486 107 720 444 323 091 627 565 727 743 604 741 030 019 072 × 2 = 1 + 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 060 038 144;
  • 24) 0.199 242 138 972 215 440 888 646 183 255 131 455 487 209 482 060 038 144 × 2 = 0 + 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 964 120 076 288;
  • 25) 0.398 484 277 944 430 881 777 292 366 510 262 910 974 418 964 120 076 288 × 2 = 0 + 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 928 240 152 576;
  • 26) 0.796 968 555 888 861 763 554 584 733 020 525 821 948 837 928 240 152 576 × 2 = 1 + 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 856 480 305 152;
  • 27) 0.593 937 111 777 723 527 109 169 466 041 051 643 897 675 856 480 305 152 × 2 = 1 + 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 712 960 610 304;
  • 28) 0.187 874 223 555 447 054 218 338 932 082 103 287 795 351 712 960 610 304 × 2 = 0 + 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 425 921 220 608;
  • 29) 0.375 748 447 110 894 108 436 677 864 164 206 575 590 703 425 921 220 608 × 2 = 0 + 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 851 842 441 216;
  • 30) 0.751 496 894 221 788 216 873 355 728 328 413 151 181 406 851 842 441 216 × 2 = 1 + 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 703 684 882 432;
  • 31) 0.502 993 788 443 576 433 746 711 456 656 826 302 362 813 703 684 882 432 × 2 = 1 + 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 407 369 764 864;
  • 32) 0.005 987 576 887 152 867 493 422 913 313 652 604 725 627 407 369 764 864 × 2 = 0 + 0.011 975 153 774 305 734 986 845 826 627 305 209 451 254 814 739 529 728;
  • 33) 0.011 975 153 774 305 734 986 845 826 627 305 209 451 254 814 739 529 728 × 2 = 0 + 0.023 950 307 548 611 469 973 691 653 254 610 418 902 509 629 479 059 456;
  • 34) 0.023 950 307 548 611 469 973 691 653 254 610 418 902 509 629 479 059 456 × 2 = 0 + 0.047 900 615 097 222 939 947 383 306 509 220 837 805 019 258 958 118 912;
  • 35) 0.047 900 615 097 222 939 947 383 306 509 220 837 805 019 258 958 118 912 × 2 = 0 + 0.095 801 230 194 445 879 894 766 613 018 441 675 610 038 517 916 237 824;
  • 36) 0.095 801 230 194 445 879 894 766 613 018 441 675 610 038 517 916 237 824 × 2 = 0 + 0.191 602 460 388 891 759 789 533 226 036 883 351 220 077 035 832 475 648;
  • 37) 0.191 602 460 388 891 759 789 533 226 036 883 351 220 077 035 832 475 648 × 2 = 0 + 0.383 204 920 777 783 519 579 066 452 073 766 702 440 154 071 664 951 296;
  • 38) 0.383 204 920 777 783 519 579 066 452 073 766 702 440 154 071 664 951 296 × 2 = 0 + 0.766 409 841 555 567 039 158 132 904 147 533 404 880 308 143 329 902 592;
  • 39) 0.766 409 841 555 567 039 158 132 904 147 533 404 880 308 143 329 902 592 × 2 = 1 + 0.532 819 683 111 134 078 316 265 808 295 066 809 760 616 286 659 805 184;
  • 40) 0.532 819 683 111 134 078 316 265 808 295 066 809 760 616 286 659 805 184 × 2 = 1 + 0.065 639 366 222 268 156 632 531 616 590 133 619 521 232 573 319 610 368;
  • 41) 0.065 639 366 222 268 156 632 531 616 590 133 619 521 232 573 319 610 368 × 2 = 0 + 0.131 278 732 444 536 313 265 063 233 180 267 239 042 465 146 639 220 736;
  • 42) 0.131 278 732 444 536 313 265 063 233 180 267 239 042 465 146 639 220 736 × 2 = 0 + 0.262 557 464 889 072 626 530 126 466 360 534 478 084 930 293 278 441 472;
  • 43) 0.262 557 464 889 072 626 530 126 466 360 534 478 084 930 293 278 441 472 × 2 = 0 + 0.525 114 929 778 145 253 060 252 932 721 068 956 169 860 586 556 882 944;
  • 44) 0.525 114 929 778 145 253 060 252 932 721 068 956 169 860 586 556 882 944 × 2 = 1 + 0.050 229 859 556 290 506 120 505 865 442 137 912 339 721 173 113 765 888;
  • 45) 0.050 229 859 556 290 506 120 505 865 442 137 912 339 721 173 113 765 888 × 2 = 0 + 0.100 459 719 112 581 012 241 011 730 884 275 824 679 442 346 227 531 776;
  • 46) 0.100 459 719 112 581 012 241 011 730 884 275 824 679 442 346 227 531 776 × 2 = 0 + 0.200 919 438 225 162 024 482 023 461 768 551 649 358 884 692 455 063 552;
  • 47) 0.200 919 438 225 162 024 482 023 461 768 551 649 358 884 692 455 063 552 × 2 = 0 + 0.401 838 876 450 324 048 964 046 923 537 103 298 717 769 384 910 127 104;
  • 48) 0.401 838 876 450 324 048 964 046 923 537 103 298 717 769 384 910 127 104 × 2 = 0 + 0.803 677 752 900 648 097 928 093 847 074 206 597 435 538 769 820 254 208;
  • 49) 0.803 677 752 900 648 097 928 093 847 074 206 597 435 538 769 820 254 208 × 2 = 1 + 0.607 355 505 801 296 195 856 187 694 148 413 194 871 077 539 640 508 416;
  • 50) 0.607 355 505 801 296 195 856 187 694 148 413 194 871 077 539 640 508 416 × 2 = 1 + 0.214 711 011 602 592 391 712 375 388 296 826 389 742 155 079 281 016 832;
  • 51) 0.214 711 011 602 592 391 712 375 388 296 826 389 742 155 079 281 016 832 × 2 = 0 + 0.429 422 023 205 184 783 424 750 776 593 652 779 484 310 158 562 033 664;
  • 52) 0.429 422 023 205 184 783 424 750 776 593 652 779 484 310 158 562 033 664 × 2 = 0 + 0.858 844 046 410 369 566 849 501 553 187 305 558 968 620 317 124 067 328;
  • 53) 0.858 844 046 410 369 566 849 501 553 187 305 558 968 620 317 124 067 328 × 2 = 1 + 0.717 688 092 820 739 133 699 003 106 374 611 117 937 240 634 248 134 656;
  • 54) 0.717 688 092 820 739 133 699 003 106 374 611 117 937 240 634 248 134 656 × 2 = 1 + 0.435 376 185 641 478 267 398 006 212 749 222 235 874 481 268 496 269 312;
  • 55) 0.435 376 185 641 478 267 398 006 212 749 222 235 874 481 268 496 269 312 × 2 = 0 + 0.870 752 371 282 956 534 796 012 425 498 444 471 748 962 536 992 538 624;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

5. Positive number before normalization:

0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 3 positions to the right, so that only one non zero digit remains to the left of it:


0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918(10) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) =


0.0010 0101 0111 1101 1000 0110 0110 0110 0000 0011 0001 0000 1100 110(2) × 20 =


1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110(2) × 2-3


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -3


Mantissa (not normalized):
1.0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-3 + 2(11-1) - 1 =


(-3 + 1 023)(10) =


1 020(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 020 ÷ 2 = 510 + 0;
  • 510 ÷ 2 = 255 + 0;
  • 255 ÷ 2 = 127 + 1;
  • 127 ÷ 2 = 63 + 1;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1020(10) =


011 1111 1100(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110 =


0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1111 1100


Mantissa (52 bits) =
0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110


Decimal number 0.146 446 609 406 726 237 799 577 818 947 575 480 357 582 031 155 762 918 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1111 1100 - 0010 1011 1110 1100 0011 0011 0011 0000 0001 1000 1000 0110 0110

How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100