0.000 000 000 000 000 000 054 248 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.000 000 000 000 000 000 054 248(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.000 000 000 000 000 000 054 248(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.000 000 000 000 000 000 054 248.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 000 000 000 000 000 054 248 × 2 = 0 + 0.000 000 000 000 000 000 108 496;
  • 2) 0.000 000 000 000 000 000 108 496 × 2 = 0 + 0.000 000 000 000 000 000 216 992;
  • 3) 0.000 000 000 000 000 000 216 992 × 2 = 0 + 0.000 000 000 000 000 000 433 984;
  • 4) 0.000 000 000 000 000 000 433 984 × 2 = 0 + 0.000 000 000 000 000 000 867 968;
  • 5) 0.000 000 000 000 000 000 867 968 × 2 = 0 + 0.000 000 000 000 000 001 735 936;
  • 6) 0.000 000 000 000 000 001 735 936 × 2 = 0 + 0.000 000 000 000 000 003 471 872;
  • 7) 0.000 000 000 000 000 003 471 872 × 2 = 0 + 0.000 000 000 000 000 006 943 744;
  • 8) 0.000 000 000 000 000 006 943 744 × 2 = 0 + 0.000 000 000 000 000 013 887 488;
  • 9) 0.000 000 000 000 000 013 887 488 × 2 = 0 + 0.000 000 000 000 000 027 774 976;
  • 10) 0.000 000 000 000 000 027 774 976 × 2 = 0 + 0.000 000 000 000 000 055 549 952;
  • 11) 0.000 000 000 000 000 055 549 952 × 2 = 0 + 0.000 000 000 000 000 111 099 904;
  • 12) 0.000 000 000 000 000 111 099 904 × 2 = 0 + 0.000 000 000 000 000 222 199 808;
  • 13) 0.000 000 000 000 000 222 199 808 × 2 = 0 + 0.000 000 000 000 000 444 399 616;
  • 14) 0.000 000 000 000 000 444 399 616 × 2 = 0 + 0.000 000 000 000 000 888 799 232;
  • 15) 0.000 000 000 000 000 888 799 232 × 2 = 0 + 0.000 000 000 000 001 777 598 464;
  • 16) 0.000 000 000 000 001 777 598 464 × 2 = 0 + 0.000 000 000 000 003 555 196 928;
  • 17) 0.000 000 000 000 003 555 196 928 × 2 = 0 + 0.000 000 000 000 007 110 393 856;
  • 18) 0.000 000 000 000 007 110 393 856 × 2 = 0 + 0.000 000 000 000 014 220 787 712;
  • 19) 0.000 000 000 000 014 220 787 712 × 2 = 0 + 0.000 000 000 000 028 441 575 424;
  • 20) 0.000 000 000 000 028 441 575 424 × 2 = 0 + 0.000 000 000 000 056 883 150 848;
  • 21) 0.000 000 000 000 056 883 150 848 × 2 = 0 + 0.000 000 000 000 113 766 301 696;
  • 22) 0.000 000 000 000 113 766 301 696 × 2 = 0 + 0.000 000 000 000 227 532 603 392;
  • 23) 0.000 000 000 000 227 532 603 392 × 2 = 0 + 0.000 000 000 000 455 065 206 784;
  • 24) 0.000 000 000 000 455 065 206 784 × 2 = 0 + 0.000 000 000 000 910 130 413 568;
  • 25) 0.000 000 000 000 910 130 413 568 × 2 = 0 + 0.000 000 000 001 820 260 827 136;
  • 26) 0.000 000 000 001 820 260 827 136 × 2 = 0 + 0.000 000 000 003 640 521 654 272;
  • 27) 0.000 000 000 003 640 521 654 272 × 2 = 0 + 0.000 000 000 007 281 043 308 544;
  • 28) 0.000 000 000 007 281 043 308 544 × 2 = 0 + 0.000 000 000 014 562 086 617 088;
  • 29) 0.000 000 000 014 562 086 617 088 × 2 = 0 + 0.000 000 000 029 124 173 234 176;
  • 30) 0.000 000 000 029 124 173 234 176 × 2 = 0 + 0.000 000 000 058 248 346 468 352;
  • 31) 0.000 000 000 058 248 346 468 352 × 2 = 0 + 0.000 000 000 116 496 692 936 704;
  • 32) 0.000 000 000 116 496 692 936 704 × 2 = 0 + 0.000 000 000 232 993 385 873 408;
  • 33) 0.000 000 000 232 993 385 873 408 × 2 = 0 + 0.000 000 000 465 986 771 746 816;
  • 34) 0.000 000 000 465 986 771 746 816 × 2 = 0 + 0.000 000 000 931 973 543 493 632;
  • 35) 0.000 000 000 931 973 543 493 632 × 2 = 0 + 0.000 000 001 863 947 086 987 264;
  • 36) 0.000 000 001 863 947 086 987 264 × 2 = 0 + 0.000 000 003 727 894 173 974 528;
  • 37) 0.000 000 003 727 894 173 974 528 × 2 = 0 + 0.000 000 007 455 788 347 949 056;
  • 38) 0.000 000 007 455 788 347 949 056 × 2 = 0 + 0.000 000 014 911 576 695 898 112;
  • 39) 0.000 000 014 911 576 695 898 112 × 2 = 0 + 0.000 000 029 823 153 391 796 224;
  • 40) 0.000 000 029 823 153 391 796 224 × 2 = 0 + 0.000 000 059 646 306 783 592 448;
  • 41) 0.000 000 059 646 306 783 592 448 × 2 = 0 + 0.000 000 119 292 613 567 184 896;
  • 42) 0.000 000 119 292 613 567 184 896 × 2 = 0 + 0.000 000 238 585 227 134 369 792;
  • 43) 0.000 000 238 585 227 134 369 792 × 2 = 0 + 0.000 000 477 170 454 268 739 584;
  • 44) 0.000 000 477 170 454 268 739 584 × 2 = 0 + 0.000 000 954 340 908 537 479 168;
  • 45) 0.000 000 954 340 908 537 479 168 × 2 = 0 + 0.000 001 908 681 817 074 958 336;
  • 46) 0.000 001 908 681 817 074 958 336 × 2 = 0 + 0.000 003 817 363 634 149 916 672;
  • 47) 0.000 003 817 363 634 149 916 672 × 2 = 0 + 0.000 007 634 727 268 299 833 344;
  • 48) 0.000 007 634 727 268 299 833 344 × 2 = 0 + 0.000 015 269 454 536 599 666 688;
  • 49) 0.000 015 269 454 536 599 666 688 × 2 = 0 + 0.000 030 538 909 073 199 333 376;
  • 50) 0.000 030 538 909 073 199 333 376 × 2 = 0 + 0.000 061 077 818 146 398 666 752;
  • 51) 0.000 061 077 818 146 398 666 752 × 2 = 0 + 0.000 122 155 636 292 797 333 504;
  • 52) 0.000 122 155 636 292 797 333 504 × 2 = 0 + 0.000 244 311 272 585 594 667 008;
  • 53) 0.000 244 311 272 585 594 667 008 × 2 = 0 + 0.000 488 622 545 171 189 334 016;
  • 54) 0.000 488 622 545 171 189 334 016 × 2 = 0 + 0.000 977 245 090 342 378 668 032;
  • 55) 0.000 977 245 090 342 378 668 032 × 2 = 0 + 0.001 954 490 180 684 757 336 064;
  • 56) 0.001 954 490 180 684 757 336 064 × 2 = 0 + 0.003 908 980 361 369 514 672 128;
  • 57) 0.003 908 980 361 369 514 672 128 × 2 = 0 + 0.007 817 960 722 739 029 344 256;
  • 58) 0.007 817 960 722 739 029 344 256 × 2 = 0 + 0.015 635 921 445 478 058 688 512;
  • 59) 0.015 635 921 445 478 058 688 512 × 2 = 0 + 0.031 271 842 890 956 117 377 024;
  • 60) 0.031 271 842 890 956 117 377 024 × 2 = 0 + 0.062 543 685 781 912 234 754 048;
  • 61) 0.062 543 685 781 912 234 754 048 × 2 = 0 + 0.125 087 371 563 824 469 508 096;
  • 62) 0.125 087 371 563 824 469 508 096 × 2 = 0 + 0.250 174 743 127 648 939 016 192;
  • 63) 0.250 174 743 127 648 939 016 192 × 2 = 0 + 0.500 349 486 255 297 878 032 384;
  • 64) 0.500 349 486 255 297 878 032 384 × 2 = 1 + 0.000 698 972 510 595 756 064 768;
  • 65) 0.000 698 972 510 595 756 064 768 × 2 = 0 + 0.001 397 945 021 191 512 129 536;
  • 66) 0.001 397 945 021 191 512 129 536 × 2 = 0 + 0.002 795 890 042 383 024 259 072;
  • 67) 0.002 795 890 042 383 024 259 072 × 2 = 0 + 0.005 591 780 084 766 048 518 144;
  • 68) 0.005 591 780 084 766 048 518 144 × 2 = 0 + 0.011 183 560 169 532 097 036 288;
  • 69) 0.011 183 560 169 532 097 036 288 × 2 = 0 + 0.022 367 120 339 064 194 072 576;
  • 70) 0.022 367 120 339 064 194 072 576 × 2 = 0 + 0.044 734 240 678 128 388 145 152;
  • 71) 0.044 734 240 678 128 388 145 152 × 2 = 0 + 0.089 468 481 356 256 776 290 304;
  • 72) 0.089 468 481 356 256 776 290 304 × 2 = 0 + 0.178 936 962 712 513 552 580 608;
  • 73) 0.178 936 962 712 513 552 580 608 × 2 = 0 + 0.357 873 925 425 027 105 161 216;
  • 74) 0.357 873 925 425 027 105 161 216 × 2 = 0 + 0.715 747 850 850 054 210 322 432;
  • 75) 0.715 747 850 850 054 210 322 432 × 2 = 1 + 0.431 495 701 700 108 420 644 864;
  • 76) 0.431 495 701 700 108 420 644 864 × 2 = 0 + 0.862 991 403 400 216 841 289 728;
  • 77) 0.862 991 403 400 216 841 289 728 × 2 = 1 + 0.725 982 806 800 433 682 579 456;
  • 78) 0.725 982 806 800 433 682 579 456 × 2 = 1 + 0.451 965 613 600 867 365 158 912;
  • 79) 0.451 965 613 600 867 365 158 912 × 2 = 0 + 0.903 931 227 201 734 730 317 824;
  • 80) 0.903 931 227 201 734 730 317 824 × 2 = 1 + 0.807 862 454 403 469 460 635 648;
  • 81) 0.807 862 454 403 469 460 635 648 × 2 = 1 + 0.615 724 908 806 938 921 271 296;
  • 82) 0.615 724 908 806 938 921 271 296 × 2 = 1 + 0.231 449 817 613 877 842 542 592;
  • 83) 0.231 449 817 613 877 842 542 592 × 2 = 0 + 0.462 899 635 227 755 685 085 184;
  • 84) 0.462 899 635 227 755 685 085 184 × 2 = 0 + 0.925 799 270 455 511 370 170 368;
  • 85) 0.925 799 270 455 511 370 170 368 × 2 = 1 + 0.851 598 540 911 022 740 340 736;
  • 86) 0.851 598 540 911 022 740 340 736 × 2 = 1 + 0.703 197 081 822 045 480 681 472;
  • 87) 0.703 197 081 822 045 480 681 472 × 2 = 1 + 0.406 394 163 644 090 961 362 944;
  • 88) 0.406 394 163 644 090 961 362 944 × 2 = 0 + 0.812 788 327 288 181 922 725 888;
  • 89) 0.812 788 327 288 181 922 725 888 × 2 = 1 + 0.625 576 654 576 363 845 451 776;
  • 90) 0.625 576 654 576 363 845 451 776 × 2 = 1 + 0.251 153 309 152 727 690 903 552;
  • 91) 0.251 153 309 152 727 690 903 552 × 2 = 0 + 0.502 306 618 305 455 381 807 104;
  • 92) 0.502 306 618 305 455 381 807 104 × 2 = 1 + 0.004 613 236 610 910 763 614 208;
  • 93) 0.004 613 236 610 910 763 614 208 × 2 = 0 + 0.009 226 473 221 821 527 228 416;
  • 94) 0.009 226 473 221 821 527 228 416 × 2 = 0 + 0.018 452 946 443 643 054 456 832;
  • 95) 0.018 452 946 443 643 054 456 832 × 2 = 0 + 0.036 905 892 887 286 108 913 664;
  • 96) 0.036 905 892 887 286 108 913 664 × 2 = 0 + 0.073 811 785 774 572 217 827 328;
  • 97) 0.073 811 785 774 572 217 827 328 × 2 = 0 + 0.147 623 571 549 144 435 654 656;
  • 98) 0.147 623 571 549 144 435 654 656 × 2 = 0 + 0.295 247 143 098 288 871 309 312;
  • 99) 0.295 247 143 098 288 871 309 312 × 2 = 0 + 0.590 494 286 196 577 742 618 624;
  • 100) 0.590 494 286 196 577 742 618 624 × 2 = 1 + 0.180 988 572 393 155 485 237 248;
  • 101) 0.180 988 572 393 155 485 237 248 × 2 = 0 + 0.361 977 144 786 310 970 474 496;
  • 102) 0.361 977 144 786 310 970 474 496 × 2 = 0 + 0.723 954 289 572 621 940 948 992;
  • 103) 0.723 954 289 572 621 940 948 992 × 2 = 1 + 0.447 908 579 145 243 881 897 984;
  • 104) 0.447 908 579 145 243 881 897 984 × 2 = 0 + 0.895 817 158 290 487 763 795 968;
  • 105) 0.895 817 158 290 487 763 795 968 × 2 = 1 + 0.791 634 316 580 975 527 591 936;
  • 106) 0.791 634 316 580 975 527 591 936 × 2 = 1 + 0.583 268 633 161 951 055 183 872;
  • 107) 0.583 268 633 161 951 055 183 872 × 2 = 1 + 0.166 537 266 323 902 110 367 744;
  • 108) 0.166 537 266 323 902 110 367 744 × 2 = 0 + 0.333 074 532 647 804 220 735 488;
  • 109) 0.333 074 532 647 804 220 735 488 × 2 = 0 + 0.666 149 065 295 608 441 470 976;
  • 110) 0.666 149 065 295 608 441 470 976 × 2 = 1 + 0.332 298 130 591 216 882 941 952;
  • 111) 0.332 298 130 591 216 882 941 952 × 2 = 0 + 0.664 596 261 182 433 765 883 904;
  • 112) 0.664 596 261 182 433 765 883 904 × 2 = 1 + 0.329 192 522 364 867 531 767 808;
  • 113) 0.329 192 522 364 867 531 767 808 × 2 = 0 + 0.658 385 044 729 735 063 535 616;
  • 114) 0.658 385 044 729 735 063 535 616 × 2 = 1 + 0.316 770 089 459 470 127 071 232;
  • 115) 0.316 770 089 459 470 127 071 232 × 2 = 0 + 0.633 540 178 918 940 254 142 464;
  • 116) 0.633 540 178 918 940 254 142 464 × 2 = 1 + 0.267 080 357 837 880 508 284 928;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 000 000 000 000 000 054 248(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101(2)

5. Positive number before normalization:

0.000 000 000 000 000 000 054 248(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 64 positions to the right, so that only one non zero digit remains to the left of it:


0.000 000 000 000 000 000 054 248(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101(2) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101(2) × 20 =


1.0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101(2) × 2-64


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -64


Mantissa (not normalized):
1.0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-64 + 2(11-1) - 1 =


(-64 + 1 023)(10) =


959(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 959 ÷ 2 = 479 + 1;
  • 479 ÷ 2 = 239 + 1;
  • 239 ÷ 2 = 119 + 1;
  • 119 ÷ 2 = 59 + 1;
  • 59 ÷ 2 = 29 + 1;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


959(10) =


011 1011 1111(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101 =


0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1011 1111


Mantissa (52 bits) =
0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101


Decimal number 0.000 000 000 000 000 000 054 248 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1011 1111 - 0000 0000 0010 1101 1100 1110 1101 0000 0001 0010 1110 0101 0101


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100