162 061 060 059 158 157 156 055 154 052 876 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 162 061 060 059 158 157 156 055 154 052 876(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
162 061 060 059 158 157 156 055 154 052 876(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 162 061 060 059 158 157 156 055 154 052 876 ÷ 2 = 81 030 530 029 579 078 578 027 577 026 438 + 0;
  • 81 030 530 029 579 078 578 027 577 026 438 ÷ 2 = 40 515 265 014 789 539 289 013 788 513 219 + 0;
  • 40 515 265 014 789 539 289 013 788 513 219 ÷ 2 = 20 257 632 507 394 769 644 506 894 256 609 + 1;
  • 20 257 632 507 394 769 644 506 894 256 609 ÷ 2 = 10 128 816 253 697 384 822 253 447 128 304 + 1;
  • 10 128 816 253 697 384 822 253 447 128 304 ÷ 2 = 5 064 408 126 848 692 411 126 723 564 152 + 0;
  • 5 064 408 126 848 692 411 126 723 564 152 ÷ 2 = 2 532 204 063 424 346 205 563 361 782 076 + 0;
  • 2 532 204 063 424 346 205 563 361 782 076 ÷ 2 = 1 266 102 031 712 173 102 781 680 891 038 + 0;
  • 1 266 102 031 712 173 102 781 680 891 038 ÷ 2 = 633 051 015 856 086 551 390 840 445 519 + 0;
  • 633 051 015 856 086 551 390 840 445 519 ÷ 2 = 316 525 507 928 043 275 695 420 222 759 + 1;
  • 316 525 507 928 043 275 695 420 222 759 ÷ 2 = 158 262 753 964 021 637 847 710 111 379 + 1;
  • 158 262 753 964 021 637 847 710 111 379 ÷ 2 = 79 131 376 982 010 818 923 855 055 689 + 1;
  • 79 131 376 982 010 818 923 855 055 689 ÷ 2 = 39 565 688 491 005 409 461 927 527 844 + 1;
  • 39 565 688 491 005 409 461 927 527 844 ÷ 2 = 19 782 844 245 502 704 730 963 763 922 + 0;
  • 19 782 844 245 502 704 730 963 763 922 ÷ 2 = 9 891 422 122 751 352 365 481 881 961 + 0;
  • 9 891 422 122 751 352 365 481 881 961 ÷ 2 = 4 945 711 061 375 676 182 740 940 980 + 1;
  • 4 945 711 061 375 676 182 740 940 980 ÷ 2 = 2 472 855 530 687 838 091 370 470 490 + 0;
  • 2 472 855 530 687 838 091 370 470 490 ÷ 2 = 1 236 427 765 343 919 045 685 235 245 + 0;
  • 1 236 427 765 343 919 045 685 235 245 ÷ 2 = 618 213 882 671 959 522 842 617 622 + 1;
  • 618 213 882 671 959 522 842 617 622 ÷ 2 = 309 106 941 335 979 761 421 308 811 + 0;
  • 309 106 941 335 979 761 421 308 811 ÷ 2 = 154 553 470 667 989 880 710 654 405 + 1;
  • 154 553 470 667 989 880 710 654 405 ÷ 2 = 77 276 735 333 994 940 355 327 202 + 1;
  • 77 276 735 333 994 940 355 327 202 ÷ 2 = 38 638 367 666 997 470 177 663 601 + 0;
  • 38 638 367 666 997 470 177 663 601 ÷ 2 = 19 319 183 833 498 735 088 831 800 + 1;
  • 19 319 183 833 498 735 088 831 800 ÷ 2 = 9 659 591 916 749 367 544 415 900 + 0;
  • 9 659 591 916 749 367 544 415 900 ÷ 2 = 4 829 795 958 374 683 772 207 950 + 0;
  • 4 829 795 958 374 683 772 207 950 ÷ 2 = 2 414 897 979 187 341 886 103 975 + 0;
  • 2 414 897 979 187 341 886 103 975 ÷ 2 = 1 207 448 989 593 670 943 051 987 + 1;
  • 1 207 448 989 593 670 943 051 987 ÷ 2 = 603 724 494 796 835 471 525 993 + 1;
  • 603 724 494 796 835 471 525 993 ÷ 2 = 301 862 247 398 417 735 762 996 + 1;
  • 301 862 247 398 417 735 762 996 ÷ 2 = 150 931 123 699 208 867 881 498 + 0;
  • 150 931 123 699 208 867 881 498 ÷ 2 = 75 465 561 849 604 433 940 749 + 0;
  • 75 465 561 849 604 433 940 749 ÷ 2 = 37 732 780 924 802 216 970 374 + 1;
  • 37 732 780 924 802 216 970 374 ÷ 2 = 18 866 390 462 401 108 485 187 + 0;
  • 18 866 390 462 401 108 485 187 ÷ 2 = 9 433 195 231 200 554 242 593 + 1;
  • 9 433 195 231 200 554 242 593 ÷ 2 = 4 716 597 615 600 277 121 296 + 1;
  • 4 716 597 615 600 277 121 296 ÷ 2 = 2 358 298 807 800 138 560 648 + 0;
  • 2 358 298 807 800 138 560 648 ÷ 2 = 1 179 149 403 900 069 280 324 + 0;
  • 1 179 149 403 900 069 280 324 ÷ 2 = 589 574 701 950 034 640 162 + 0;
  • 589 574 701 950 034 640 162 ÷ 2 = 294 787 350 975 017 320 081 + 0;
  • 294 787 350 975 017 320 081 ÷ 2 = 147 393 675 487 508 660 040 + 1;
  • 147 393 675 487 508 660 040 ÷ 2 = 73 696 837 743 754 330 020 + 0;
  • 73 696 837 743 754 330 020 ÷ 2 = 36 848 418 871 877 165 010 + 0;
  • 36 848 418 871 877 165 010 ÷ 2 = 18 424 209 435 938 582 505 + 0;
  • 18 424 209 435 938 582 505 ÷ 2 = 9 212 104 717 969 291 252 + 1;
  • 9 212 104 717 969 291 252 ÷ 2 = 4 606 052 358 984 645 626 + 0;
  • 4 606 052 358 984 645 626 ÷ 2 = 2 303 026 179 492 322 813 + 0;
  • 2 303 026 179 492 322 813 ÷ 2 = 1 151 513 089 746 161 406 + 1;
  • 1 151 513 089 746 161 406 ÷ 2 = 575 756 544 873 080 703 + 0;
  • 575 756 544 873 080 703 ÷ 2 = 287 878 272 436 540 351 + 1;
  • 287 878 272 436 540 351 ÷ 2 = 143 939 136 218 270 175 + 1;
  • 143 939 136 218 270 175 ÷ 2 = 71 969 568 109 135 087 + 1;
  • 71 969 568 109 135 087 ÷ 2 = 35 984 784 054 567 543 + 1;
  • 35 984 784 054 567 543 ÷ 2 = 17 992 392 027 283 771 + 1;
  • 17 992 392 027 283 771 ÷ 2 = 8 996 196 013 641 885 + 1;
  • 8 996 196 013 641 885 ÷ 2 = 4 498 098 006 820 942 + 1;
  • 4 498 098 006 820 942 ÷ 2 = 2 249 049 003 410 471 + 0;
  • 2 249 049 003 410 471 ÷ 2 = 1 124 524 501 705 235 + 1;
  • 1 124 524 501 705 235 ÷ 2 = 562 262 250 852 617 + 1;
  • 562 262 250 852 617 ÷ 2 = 281 131 125 426 308 + 1;
  • 281 131 125 426 308 ÷ 2 = 140 565 562 713 154 + 0;
  • 140 565 562 713 154 ÷ 2 = 70 282 781 356 577 + 0;
  • 70 282 781 356 577 ÷ 2 = 35 141 390 678 288 + 1;
  • 35 141 390 678 288 ÷ 2 = 17 570 695 339 144 + 0;
  • 17 570 695 339 144 ÷ 2 = 8 785 347 669 572 + 0;
  • 8 785 347 669 572 ÷ 2 = 4 392 673 834 786 + 0;
  • 4 392 673 834 786 ÷ 2 = 2 196 336 917 393 + 0;
  • 2 196 336 917 393 ÷ 2 = 1 098 168 458 696 + 1;
  • 1 098 168 458 696 ÷ 2 = 549 084 229 348 + 0;
  • 549 084 229 348 ÷ 2 = 274 542 114 674 + 0;
  • 274 542 114 674 ÷ 2 = 137 271 057 337 + 0;
  • 137 271 057 337 ÷ 2 = 68 635 528 668 + 1;
  • 68 635 528 668 ÷ 2 = 34 317 764 334 + 0;
  • 34 317 764 334 ÷ 2 = 17 158 882 167 + 0;
  • 17 158 882 167 ÷ 2 = 8 579 441 083 + 1;
  • 8 579 441 083 ÷ 2 = 4 289 720 541 + 1;
  • 4 289 720 541 ÷ 2 = 2 144 860 270 + 1;
  • 2 144 860 270 ÷ 2 = 1 072 430 135 + 0;
  • 1 072 430 135 ÷ 2 = 536 215 067 + 1;
  • 536 215 067 ÷ 2 = 268 107 533 + 1;
  • 268 107 533 ÷ 2 = 134 053 766 + 1;
  • 134 053 766 ÷ 2 = 67 026 883 + 0;
  • 67 026 883 ÷ 2 = 33 513 441 + 1;
  • 33 513 441 ÷ 2 = 16 756 720 + 1;
  • 16 756 720 ÷ 2 = 8 378 360 + 0;
  • 8 378 360 ÷ 2 = 4 189 180 + 0;
  • 4 189 180 ÷ 2 = 2 094 590 + 0;
  • 2 094 590 ÷ 2 = 1 047 295 + 0;
  • 1 047 295 ÷ 2 = 523 647 + 1;
  • 523 647 ÷ 2 = 261 823 + 1;
  • 261 823 ÷ 2 = 130 911 + 1;
  • 130 911 ÷ 2 = 65 455 + 1;
  • 65 455 ÷ 2 = 32 727 + 1;
  • 32 727 ÷ 2 = 16 363 + 1;
  • 16 363 ÷ 2 = 8 181 + 1;
  • 8 181 ÷ 2 = 4 090 + 1;
  • 4 090 ÷ 2 = 2 045 + 0;
  • 2 045 ÷ 2 = 1 022 + 1;
  • 1 022 ÷ 2 = 511 + 0;
  • 511 ÷ 2 = 255 + 1;
  • 255 ÷ 2 = 127 + 1;
  • 127 ÷ 2 = 63 + 1;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

162 061 060 059 158 157 156 055 154 052 876(10) =


111 1111 1101 0111 1111 1000 0110 1110 1110 0100 0100 0010 0111 0111 1111 0100 1000 1000 0110 1001 1100 0101 1010 0100 1111 0000 1100(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 106 positions to the left, so that only one non zero digit remains to the left of it:


162 061 060 059 158 157 156 055 154 052 876(10) =


111 1111 1101 0111 1111 1000 0110 1110 1110 0100 0100 0010 0111 0111 1111 0100 1000 1000 0110 1001 1100 0101 1010 0100 1111 0000 1100(2) =


111 1111 1101 0111 1111 1000 0110 1110 1110 0100 0100 0010 0111 0111 1111 0100 1000 1000 0110 1001 1100 0101 1010 0100 1111 0000 1100(2) × 20 =


1.1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101 1111 1101 0010 0010 0001 1010 0111 0001 0110 1001 0011 1100 0011 00(2) × 2106


4. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 106


Mantissa (not normalized):
1.1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101 1111 1101 0010 0010 0001 1010 0111 0001 0110 1001 0011 1100 0011 00


5. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


106 + 2(11-1) - 1 =


(106 + 1 023)(10) =


1 129(10)


6. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 1 129 ÷ 2 = 564 + 1;
  • 564 ÷ 2 = 282 + 0;
  • 282 ÷ 2 = 141 + 0;
  • 141 ÷ 2 = 70 + 1;
  • 70 ÷ 2 = 35 + 0;
  • 35 ÷ 2 = 17 + 1;
  • 17 ÷ 2 = 8 + 1;
  • 8 ÷ 2 = 4 + 0;
  • 4 ÷ 2 = 2 + 0;
  • 2 ÷ 2 = 1 + 0;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


1129(10) =


100 0110 1001(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101 11 1111 0100 1000 1000 0110 1001 1100 0101 1010 0100 1111 0000 1100 =


1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101


9. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
100 0110 1001


Mantissa (52 bits) =
1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101


Decimal number 162 061 060 059 158 157 156 055 154 052 876 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 100 0110 1001 - 1111 1111 0101 1111 1110 0001 1011 1011 1001 0001 0000 1001 1101


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100