823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 823 549.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 823 549 ÷ 2 = 411 774 + 1;
  • 411 774 ÷ 2 = 205 887 + 0;
  • 205 887 ÷ 2 = 102 943 + 1;
  • 102 943 ÷ 2 = 51 471 + 1;
  • 51 471 ÷ 2 = 25 735 + 1;
  • 25 735 ÷ 2 = 12 867 + 1;
  • 12 867 ÷ 2 = 6 433 + 1;
  • 6 433 ÷ 2 = 3 216 + 1;
  • 3 216 ÷ 2 = 1 608 + 0;
  • 1 608 ÷ 2 = 804 + 0;
  • 804 ÷ 2 = 402 + 0;
  • 402 ÷ 2 = 201 + 0;
  • 201 ÷ 2 = 100 + 1;
  • 100 ÷ 2 = 50 + 0;
  • 50 ÷ 2 = 25 + 0;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

823 549(10) =


1100 1001 0000 1111 1101(2)


3. Convert to binary (base 2) the fractional part: 0.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24 × 2 = 1 + 0.329 165 285 517 407 102 374 132 844 008 149 965 550 058 48;
  • 2) 0.329 165 285 517 407 102 374 132 844 008 149 965 550 058 48 × 2 = 0 + 0.658 330 571 034 814 204 748 265 688 016 299 931 100 116 96;
  • 3) 0.658 330 571 034 814 204 748 265 688 016 299 931 100 116 96 × 2 = 1 + 0.316 661 142 069 628 409 496 531 376 032 599 862 200 233 92;
  • 4) 0.316 661 142 069 628 409 496 531 376 032 599 862 200 233 92 × 2 = 0 + 0.633 322 284 139 256 818 993 062 752 065 199 724 400 467 84;
  • 5) 0.633 322 284 139 256 818 993 062 752 065 199 724 400 467 84 × 2 = 1 + 0.266 644 568 278 513 637 986 125 504 130 399 448 800 935 68;
  • 6) 0.266 644 568 278 513 637 986 125 504 130 399 448 800 935 68 × 2 = 0 + 0.533 289 136 557 027 275 972 251 008 260 798 897 601 871 36;
  • 7) 0.533 289 136 557 027 275 972 251 008 260 798 897 601 871 36 × 2 = 1 + 0.066 578 273 114 054 551 944 502 016 521 597 795 203 742 72;
  • 8) 0.066 578 273 114 054 551 944 502 016 521 597 795 203 742 72 × 2 = 0 + 0.133 156 546 228 109 103 889 004 033 043 195 590 407 485 44;
  • 9) 0.133 156 546 228 109 103 889 004 033 043 195 590 407 485 44 × 2 = 0 + 0.266 313 092 456 218 207 778 008 066 086 391 180 814 970 88;
  • 10) 0.266 313 092 456 218 207 778 008 066 086 391 180 814 970 88 × 2 = 0 + 0.532 626 184 912 436 415 556 016 132 172 782 361 629 941 76;
  • 11) 0.532 626 184 912 436 415 556 016 132 172 782 361 629 941 76 × 2 = 1 + 0.065 252 369 824 872 831 112 032 264 345 564 723 259 883 52;
  • 12) 0.065 252 369 824 872 831 112 032 264 345 564 723 259 883 52 × 2 = 0 + 0.130 504 739 649 745 662 224 064 528 691 129 446 519 767 04;
  • 13) 0.130 504 739 649 745 662 224 064 528 691 129 446 519 767 04 × 2 = 0 + 0.261 009 479 299 491 324 448 129 057 382 258 893 039 534 08;
  • 14) 0.261 009 479 299 491 324 448 129 057 382 258 893 039 534 08 × 2 = 0 + 0.522 018 958 598 982 648 896 258 114 764 517 786 079 068 16;
  • 15) 0.522 018 958 598 982 648 896 258 114 764 517 786 079 068 16 × 2 = 1 + 0.044 037 917 197 965 297 792 516 229 529 035 572 158 136 32;
  • 16) 0.044 037 917 197 965 297 792 516 229 529 035 572 158 136 32 × 2 = 0 + 0.088 075 834 395 930 595 585 032 459 058 071 144 316 272 64;
  • 17) 0.088 075 834 395 930 595 585 032 459 058 071 144 316 272 64 × 2 = 0 + 0.176 151 668 791 861 191 170 064 918 116 142 288 632 545 28;
  • 18) 0.176 151 668 791 861 191 170 064 918 116 142 288 632 545 28 × 2 = 0 + 0.352 303 337 583 722 382 340 129 836 232 284 577 265 090 56;
  • 19) 0.352 303 337 583 722 382 340 129 836 232 284 577 265 090 56 × 2 = 0 + 0.704 606 675 167 444 764 680 259 672 464 569 154 530 181 12;
  • 20) 0.704 606 675 167 444 764 680 259 672 464 569 154 530 181 12 × 2 = 1 + 0.409 213 350 334 889 529 360 519 344 929 138 309 060 362 24;
  • 21) 0.409 213 350 334 889 529 360 519 344 929 138 309 060 362 24 × 2 = 0 + 0.818 426 700 669 779 058 721 038 689 858 276 618 120 724 48;
  • 22) 0.818 426 700 669 779 058 721 038 689 858 276 618 120 724 48 × 2 = 1 + 0.636 853 401 339 558 117 442 077 379 716 553 236 241 448 96;
  • 23) 0.636 853 401 339 558 117 442 077 379 716 553 236 241 448 96 × 2 = 1 + 0.273 706 802 679 116 234 884 154 759 433 106 472 482 897 92;
  • 24) 0.273 706 802 679 116 234 884 154 759 433 106 472 482 897 92 × 2 = 0 + 0.547 413 605 358 232 469 768 309 518 866 212 944 965 795 84;
  • 25) 0.547 413 605 358 232 469 768 309 518 866 212 944 965 795 84 × 2 = 1 + 0.094 827 210 716 464 939 536 619 037 732 425 889 931 591 68;
  • 26) 0.094 827 210 716 464 939 536 619 037 732 425 889 931 591 68 × 2 = 0 + 0.189 654 421 432 929 879 073 238 075 464 851 779 863 183 36;
  • 27) 0.189 654 421 432 929 879 073 238 075 464 851 779 863 183 36 × 2 = 0 + 0.379 308 842 865 859 758 146 476 150 929 703 559 726 366 72;
  • 28) 0.379 308 842 865 859 758 146 476 150 929 703 559 726 366 72 × 2 = 0 + 0.758 617 685 731 719 516 292 952 301 859 407 119 452 733 44;
  • 29) 0.758 617 685 731 719 516 292 952 301 859 407 119 452 733 44 × 2 = 1 + 0.517 235 371 463 439 032 585 904 603 718 814 238 905 466 88;
  • 30) 0.517 235 371 463 439 032 585 904 603 718 814 238 905 466 88 × 2 = 1 + 0.034 470 742 926 878 065 171 809 207 437 628 477 810 933 76;
  • 31) 0.034 470 742 926 878 065 171 809 207 437 628 477 810 933 76 × 2 = 0 + 0.068 941 485 853 756 130 343 618 414 875 256 955 621 867 52;
  • 32) 0.068 941 485 853 756 130 343 618 414 875 256 955 621 867 52 × 2 = 0 + 0.137 882 971 707 512 260 687 236 829 750 513 911 243 735 04;
  • 33) 0.137 882 971 707 512 260 687 236 829 750 513 911 243 735 04 × 2 = 0 + 0.275 765 943 415 024 521 374 473 659 501 027 822 487 470 08;
  • 34) 0.275 765 943 415 024 521 374 473 659 501 027 822 487 470 08 × 2 = 0 + 0.551 531 886 830 049 042 748 947 319 002 055 644 974 940 16;
  • 35) 0.551 531 886 830 049 042 748 947 319 002 055 644 974 940 16 × 2 = 1 + 0.103 063 773 660 098 085 497 894 638 004 111 289 949 880 32;
  • 36) 0.103 063 773 660 098 085 497 894 638 004 111 289 949 880 32 × 2 = 0 + 0.206 127 547 320 196 170 995 789 276 008 222 579 899 760 64;
  • 37) 0.206 127 547 320 196 170 995 789 276 008 222 579 899 760 64 × 2 = 0 + 0.412 255 094 640 392 341 991 578 552 016 445 159 799 521 28;
  • 38) 0.412 255 094 640 392 341 991 578 552 016 445 159 799 521 28 × 2 = 0 + 0.824 510 189 280 784 683 983 157 104 032 890 319 599 042 56;
  • 39) 0.824 510 189 280 784 683 983 157 104 032 890 319 599 042 56 × 2 = 1 + 0.649 020 378 561 569 367 966 314 208 065 780 639 198 085 12;
  • 40) 0.649 020 378 561 569 367 966 314 208 065 780 639 198 085 12 × 2 = 1 + 0.298 040 757 123 138 735 932 628 416 131 561 278 396 170 24;
  • 41) 0.298 040 757 123 138 735 932 628 416 131 561 278 396 170 24 × 2 = 0 + 0.596 081 514 246 277 471 865 256 832 263 122 556 792 340 48;
  • 42) 0.596 081 514 246 277 471 865 256 832 263 122 556 792 340 48 × 2 = 1 + 0.192 163 028 492 554 943 730 513 664 526 245 113 584 680 96;
  • 43) 0.192 163 028 492 554 943 730 513 664 526 245 113 584 680 96 × 2 = 0 + 0.384 326 056 985 109 887 461 027 329 052 490 227 169 361 92;
  • 44) 0.384 326 056 985 109 887 461 027 329 052 490 227 169 361 92 × 2 = 0 + 0.768 652 113 970 219 774 922 054 658 104 980 454 338 723 84;
  • 45) 0.768 652 113 970 219 774 922 054 658 104 980 454 338 723 84 × 2 = 1 + 0.537 304 227 940 439 549 844 109 316 209 960 908 677 447 68;
  • 46) 0.537 304 227 940 439 549 844 109 316 209 960 908 677 447 68 × 2 = 1 + 0.074 608 455 880 879 099 688 218 632 419 921 817 354 895 36;
  • 47) 0.074 608 455 880 879 099 688 218 632 419 921 817 354 895 36 × 2 = 0 + 0.149 216 911 761 758 199 376 437 264 839 843 634 709 790 72;
  • 48) 0.149 216 911 761 758 199 376 437 264 839 843 634 709 790 72 × 2 = 0 + 0.298 433 823 523 516 398 752 874 529 679 687 269 419 581 44;
  • 49) 0.298 433 823 523 516 398 752 874 529 679 687 269 419 581 44 × 2 = 0 + 0.596 867 647 047 032 797 505 749 059 359 374 538 839 162 88;
  • 50) 0.596 867 647 047 032 797 505 749 059 359 374 538 839 162 88 × 2 = 1 + 0.193 735 294 094 065 595 011 498 118 718 749 077 678 325 76;
  • 51) 0.193 735 294 094 065 595 011 498 118 718 749 077 678 325 76 × 2 = 0 + 0.387 470 588 188 131 190 022 996 237 437 498 155 356 651 52;
  • 52) 0.387 470 588 188 131 190 022 996 237 437 498 155 356 651 52 × 2 = 0 + 0.774 941 176 376 262 380 045 992 474 874 996 310 713 303 04;
  • 53) 0.774 941 176 376 262 380 045 992 474 874 996 310 713 303 04 × 2 = 1 + 0.549 882 352 752 524 760 091 984 949 749 992 621 426 606 08;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24(10) =


0.1010 1010 0010 0010 0001 0110 1000 1100 0010 0011 0100 1100 0100 1(2)

5. Positive number before normalization:

823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24(10) =


1100 1001 0000 1111 1101.1010 1010 0010 0010 0001 0110 1000 1100 0010 0011 0100 1100 0100 1(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 19 positions to the left, so that only one non zero digit remains to the left of it:


823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24(10) =


1100 1001 0000 1111 1101.1010 1010 0010 0010 0001 0110 1000 1100 0010 0011 0100 1100 0100 1(2) =


1100 1001 0000 1111 1101.1010 1010 0010 0010 0001 0110 1000 1100 0010 0011 0100 1100 0100 1(2) × 20 =


1.1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000 0100 0110 1001 1000 1001(2) × 219


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 19


Mantissa (not normalized):
1.1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000 0100 0110 1001 1000 1001


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


19 + 2(11-1) - 1 =


(19 + 1 023)(10) =


1 042(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 042 ÷ 2 = 521 + 0;
  • 521 ÷ 2 = 260 + 1;
  • 260 ÷ 2 = 130 + 0;
  • 130 ÷ 2 = 65 + 0;
  • 65 ÷ 2 = 32 + 1;
  • 32 ÷ 2 = 16 + 0;
  • 16 ÷ 2 = 8 + 0;
  • 8 ÷ 2 = 4 + 0;
  • 4 ÷ 2 = 2 + 0;
  • 2 ÷ 2 = 1 + 0;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1042(10) =


100 0001 0010(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000 0100 0110 1001 1000 1001 =


1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
100 0001 0010


Mantissa (52 bits) =
1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000


Decimal number 823 549.664 582 642 758 703 551 187 066 422 004 074 982 775 029 24 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 100 0001 0010 - 1001 0010 0001 1111 1011 0101 0100 0100 0100 0010 1101 0001 1000


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100