1 000 110 001 110 110 999 999 999 999 659 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 1 000 110 001 110 110 999 999 999 999 659(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
1 000 110 001 110 110 999 999 999 999 659(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 1 000 110 001 110 110 999 999 999 999 659 ÷ 2 = 500 055 000 555 055 499 999 999 999 829 + 1;
  • 500 055 000 555 055 499 999 999 999 829 ÷ 2 = 250 027 500 277 527 749 999 999 999 914 + 1;
  • 250 027 500 277 527 749 999 999 999 914 ÷ 2 = 125 013 750 138 763 874 999 999 999 957 + 0;
  • 125 013 750 138 763 874 999 999 999 957 ÷ 2 = 62 506 875 069 381 937 499 999 999 978 + 1;
  • 62 506 875 069 381 937 499 999 999 978 ÷ 2 = 31 253 437 534 690 968 749 999 999 989 + 0;
  • 31 253 437 534 690 968 749 999 999 989 ÷ 2 = 15 626 718 767 345 484 374 999 999 994 + 1;
  • 15 626 718 767 345 484 374 999 999 994 ÷ 2 = 7 813 359 383 672 742 187 499 999 997 + 0;
  • 7 813 359 383 672 742 187 499 999 997 ÷ 2 = 3 906 679 691 836 371 093 749 999 998 + 1;
  • 3 906 679 691 836 371 093 749 999 998 ÷ 2 = 1 953 339 845 918 185 546 874 999 999 + 0;
  • 1 953 339 845 918 185 546 874 999 999 ÷ 2 = 976 669 922 959 092 773 437 499 999 + 1;
  • 976 669 922 959 092 773 437 499 999 ÷ 2 = 488 334 961 479 546 386 718 749 999 + 1;
  • 488 334 961 479 546 386 718 749 999 ÷ 2 = 244 167 480 739 773 193 359 374 999 + 1;
  • 244 167 480 739 773 193 359 374 999 ÷ 2 = 122 083 740 369 886 596 679 687 499 + 1;
  • 122 083 740 369 886 596 679 687 499 ÷ 2 = 61 041 870 184 943 298 339 843 749 + 1;
  • 61 041 870 184 943 298 339 843 749 ÷ 2 = 30 520 935 092 471 649 169 921 874 + 1;
  • 30 520 935 092 471 649 169 921 874 ÷ 2 = 15 260 467 546 235 824 584 960 937 + 0;
  • 15 260 467 546 235 824 584 960 937 ÷ 2 = 7 630 233 773 117 912 292 480 468 + 1;
  • 7 630 233 773 117 912 292 480 468 ÷ 2 = 3 815 116 886 558 956 146 240 234 + 0;
  • 3 815 116 886 558 956 146 240 234 ÷ 2 = 1 907 558 443 279 478 073 120 117 + 0;
  • 1 907 558 443 279 478 073 120 117 ÷ 2 = 953 779 221 639 739 036 560 058 + 1;
  • 953 779 221 639 739 036 560 058 ÷ 2 = 476 889 610 819 869 518 280 029 + 0;
  • 476 889 610 819 869 518 280 029 ÷ 2 = 238 444 805 409 934 759 140 014 + 1;
  • 238 444 805 409 934 759 140 014 ÷ 2 = 119 222 402 704 967 379 570 007 + 0;
  • 119 222 402 704 967 379 570 007 ÷ 2 = 59 611 201 352 483 689 785 003 + 1;
  • 59 611 201 352 483 689 785 003 ÷ 2 = 29 805 600 676 241 844 892 501 + 1;
  • 29 805 600 676 241 844 892 501 ÷ 2 = 14 902 800 338 120 922 446 250 + 1;
  • 14 902 800 338 120 922 446 250 ÷ 2 = 7 451 400 169 060 461 223 125 + 0;
  • 7 451 400 169 060 461 223 125 ÷ 2 = 3 725 700 084 530 230 611 562 + 1;
  • 3 725 700 084 530 230 611 562 ÷ 2 = 1 862 850 042 265 115 305 781 + 0;
  • 1 862 850 042 265 115 305 781 ÷ 2 = 931 425 021 132 557 652 890 + 1;
  • 931 425 021 132 557 652 890 ÷ 2 = 465 712 510 566 278 826 445 + 0;
  • 465 712 510 566 278 826 445 ÷ 2 = 232 856 255 283 139 413 222 + 1;
  • 232 856 255 283 139 413 222 ÷ 2 = 116 428 127 641 569 706 611 + 0;
  • 116 428 127 641 569 706 611 ÷ 2 = 58 214 063 820 784 853 305 + 1;
  • 58 214 063 820 784 853 305 ÷ 2 = 29 107 031 910 392 426 652 + 1;
  • 29 107 031 910 392 426 652 ÷ 2 = 14 553 515 955 196 213 326 + 0;
  • 14 553 515 955 196 213 326 ÷ 2 = 7 276 757 977 598 106 663 + 0;
  • 7 276 757 977 598 106 663 ÷ 2 = 3 638 378 988 799 053 331 + 1;
  • 3 638 378 988 799 053 331 ÷ 2 = 1 819 189 494 399 526 665 + 1;
  • 1 819 189 494 399 526 665 ÷ 2 = 909 594 747 199 763 332 + 1;
  • 909 594 747 199 763 332 ÷ 2 = 454 797 373 599 881 666 + 0;
  • 454 797 373 599 881 666 ÷ 2 = 227 398 686 799 940 833 + 0;
  • 227 398 686 799 940 833 ÷ 2 = 113 699 343 399 970 416 + 1;
  • 113 699 343 399 970 416 ÷ 2 = 56 849 671 699 985 208 + 0;
  • 56 849 671 699 985 208 ÷ 2 = 28 424 835 849 992 604 + 0;
  • 28 424 835 849 992 604 ÷ 2 = 14 212 417 924 996 302 + 0;
  • 14 212 417 924 996 302 ÷ 2 = 7 106 208 962 498 151 + 0;
  • 7 106 208 962 498 151 ÷ 2 = 3 553 104 481 249 075 + 1;
  • 3 553 104 481 249 075 ÷ 2 = 1 776 552 240 624 537 + 1;
  • 1 776 552 240 624 537 ÷ 2 = 888 276 120 312 268 + 1;
  • 888 276 120 312 268 ÷ 2 = 444 138 060 156 134 + 0;
  • 444 138 060 156 134 ÷ 2 = 222 069 030 078 067 + 0;
  • 222 069 030 078 067 ÷ 2 = 111 034 515 039 033 + 1;
  • 111 034 515 039 033 ÷ 2 = 55 517 257 519 516 + 1;
  • 55 517 257 519 516 ÷ 2 = 27 758 628 759 758 + 0;
  • 27 758 628 759 758 ÷ 2 = 13 879 314 379 879 + 0;
  • 13 879 314 379 879 ÷ 2 = 6 939 657 189 939 + 1;
  • 6 939 657 189 939 ÷ 2 = 3 469 828 594 969 + 1;
  • 3 469 828 594 969 ÷ 2 = 1 734 914 297 484 + 1;
  • 1 734 914 297 484 ÷ 2 = 867 457 148 742 + 0;
  • 867 457 148 742 ÷ 2 = 433 728 574 371 + 0;
  • 433 728 574 371 ÷ 2 = 216 864 287 185 + 1;
  • 216 864 287 185 ÷ 2 = 108 432 143 592 + 1;
  • 108 432 143 592 ÷ 2 = 54 216 071 796 + 0;
  • 54 216 071 796 ÷ 2 = 27 108 035 898 + 0;
  • 27 108 035 898 ÷ 2 = 13 554 017 949 + 0;
  • 13 554 017 949 ÷ 2 = 6 777 008 974 + 1;
  • 6 777 008 974 ÷ 2 = 3 388 504 487 + 0;
  • 3 388 504 487 ÷ 2 = 1 694 252 243 + 1;
  • 1 694 252 243 ÷ 2 = 847 126 121 + 1;
  • 847 126 121 ÷ 2 = 423 563 060 + 1;
  • 423 563 060 ÷ 2 = 211 781 530 + 0;
  • 211 781 530 ÷ 2 = 105 890 765 + 0;
  • 105 890 765 ÷ 2 = 52 945 382 + 1;
  • 52 945 382 ÷ 2 = 26 472 691 + 0;
  • 26 472 691 ÷ 2 = 13 236 345 + 1;
  • 13 236 345 ÷ 2 = 6 618 172 + 1;
  • 6 618 172 ÷ 2 = 3 309 086 + 0;
  • 3 309 086 ÷ 2 = 1 654 543 + 0;
  • 1 654 543 ÷ 2 = 827 271 + 1;
  • 827 271 ÷ 2 = 413 635 + 1;
  • 413 635 ÷ 2 = 206 817 + 1;
  • 206 817 ÷ 2 = 103 408 + 1;
  • 103 408 ÷ 2 = 51 704 + 0;
  • 51 704 ÷ 2 = 25 852 + 0;
  • 25 852 ÷ 2 = 12 926 + 0;
  • 12 926 ÷ 2 = 6 463 + 0;
  • 6 463 ÷ 2 = 3 231 + 1;
  • 3 231 ÷ 2 = 1 615 + 1;
  • 1 615 ÷ 2 = 807 + 1;
  • 807 ÷ 2 = 403 + 1;
  • 403 ÷ 2 = 201 + 1;
  • 201 ÷ 2 = 100 + 1;
  • 100 ÷ 2 = 50 + 0;
  • 50 ÷ 2 = 25 + 0;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

1 000 110 001 110 110 999 999 999 999 659(10) =


1100 1001 1111 1000 0111 1001 1010 0111 0100 0110 0111 0011 0011 1000 0100 1110 0110 1010 1011 1010 1001 0111 1110 1010 1011(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 99 positions to the left, so that only one non zero digit remains to the left of it:


1 000 110 001 110 110 999 999 999 999 659(10) =


1100 1001 1111 1000 0111 1001 1010 0111 0100 0110 0111 0011 0011 1000 0100 1110 0110 1010 1011 1010 1001 0111 1110 1010 1011(2) =


1100 1001 1111 1000 0111 1001 1010 0111 0100 0110 0111 0011 0011 1000 0100 1110 0110 1010 1011 1010 1001 0111 1110 1010 1011(2) × 20 =


1.1001 0011 1111 0000 1111 0011 0100 1110 1000 1100 1110 0110 0111 0000 1001 1100 1101 0101 0111 0101 0010 1111 1101 0101 011(2) × 299


4. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 99


Mantissa (not normalized):
1.1001 0011 1111 0000 1111 0011 0100 1110 1000 1100 1110 0110 0111 0000 1001 1100 1101 0101 0111 0101 0010 1111 1101 0101 011


5. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


99 + 2(8-1) - 1 =


(99 + 127)(10) =


226(10)


6. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 226 ÷ 2 = 113 + 0;
  • 113 ÷ 2 = 56 + 1;
  • 56 ÷ 2 = 28 + 0;
  • 28 ÷ 2 = 14 + 0;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


226(10) =


1110 0010(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 100 1001 1111 1000 0111 1001 1010 0111 0100 0110 0111 0011 0011 1000 0100 1110 0110 1010 1011 1010 1001 0111 1110 1010 1011 =


100 1001 1111 1000 0111 1001


9. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1110 0010


Mantissa (23 bits) =
100 1001 1111 1000 0111 1001


Decimal number 1 000 110 001 110 110 999 999 999 999 659 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1110 0010 - 100 1001 1111 1000 0111 1001


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111