1 000 001 000 110 999 999 999 999 999 294 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 1 000 001 000 110 999 999 999 999 999 294(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
1 000 001 000 110 999 999 999 999 999 294(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 1 000 001 000 110 999 999 999 999 999 294 ÷ 2 = 500 000 500 055 499 999 999 999 999 647 + 0;
  • 500 000 500 055 499 999 999 999 999 647 ÷ 2 = 250 000 250 027 749 999 999 999 999 823 + 1;
  • 250 000 250 027 749 999 999 999 999 823 ÷ 2 = 125 000 125 013 874 999 999 999 999 911 + 1;
  • 125 000 125 013 874 999 999 999 999 911 ÷ 2 = 62 500 062 506 937 499 999 999 999 955 + 1;
  • 62 500 062 506 937 499 999 999 999 955 ÷ 2 = 31 250 031 253 468 749 999 999 999 977 + 1;
  • 31 250 031 253 468 749 999 999 999 977 ÷ 2 = 15 625 015 626 734 374 999 999 999 988 + 1;
  • 15 625 015 626 734 374 999 999 999 988 ÷ 2 = 7 812 507 813 367 187 499 999 999 994 + 0;
  • 7 812 507 813 367 187 499 999 999 994 ÷ 2 = 3 906 253 906 683 593 749 999 999 997 + 0;
  • 3 906 253 906 683 593 749 999 999 997 ÷ 2 = 1 953 126 953 341 796 874 999 999 998 + 1;
  • 1 953 126 953 341 796 874 999 999 998 ÷ 2 = 976 563 476 670 898 437 499 999 999 + 0;
  • 976 563 476 670 898 437 499 999 999 ÷ 2 = 488 281 738 335 449 218 749 999 999 + 1;
  • 488 281 738 335 449 218 749 999 999 ÷ 2 = 244 140 869 167 724 609 374 999 999 + 1;
  • 244 140 869 167 724 609 374 999 999 ÷ 2 = 122 070 434 583 862 304 687 499 999 + 1;
  • 122 070 434 583 862 304 687 499 999 ÷ 2 = 61 035 217 291 931 152 343 749 999 + 1;
  • 61 035 217 291 931 152 343 749 999 ÷ 2 = 30 517 608 645 965 576 171 874 999 + 1;
  • 30 517 608 645 965 576 171 874 999 ÷ 2 = 15 258 804 322 982 788 085 937 499 + 1;
  • 15 258 804 322 982 788 085 937 499 ÷ 2 = 7 629 402 161 491 394 042 968 749 + 1;
  • 7 629 402 161 491 394 042 968 749 ÷ 2 = 3 814 701 080 745 697 021 484 374 + 1;
  • 3 814 701 080 745 697 021 484 374 ÷ 2 = 1 907 350 540 372 848 510 742 187 + 0;
  • 1 907 350 540 372 848 510 742 187 ÷ 2 = 953 675 270 186 424 255 371 093 + 1;
  • 953 675 270 186 424 255 371 093 ÷ 2 = 476 837 635 093 212 127 685 546 + 1;
  • 476 837 635 093 212 127 685 546 ÷ 2 = 238 418 817 546 606 063 842 773 + 0;
  • 238 418 817 546 606 063 842 773 ÷ 2 = 119 209 408 773 303 031 921 386 + 1;
  • 119 209 408 773 303 031 921 386 ÷ 2 = 59 604 704 386 651 515 960 693 + 0;
  • 59 604 704 386 651 515 960 693 ÷ 2 = 29 802 352 193 325 757 980 346 + 1;
  • 29 802 352 193 325 757 980 346 ÷ 2 = 14 901 176 096 662 878 990 173 + 0;
  • 14 901 176 096 662 878 990 173 ÷ 2 = 7 450 588 048 331 439 495 086 + 1;
  • 7 450 588 048 331 439 495 086 ÷ 2 = 3 725 294 024 165 719 747 543 + 0;
  • 3 725 294 024 165 719 747 543 ÷ 2 = 1 862 647 012 082 859 873 771 + 1;
  • 1 862 647 012 082 859 873 771 ÷ 2 = 931 323 506 041 429 936 885 + 1;
  • 931 323 506 041 429 936 885 ÷ 2 = 465 661 753 020 714 968 442 + 1;
  • 465 661 753 020 714 968 442 ÷ 2 = 232 830 876 510 357 484 221 + 0;
  • 232 830 876 510 357 484 221 ÷ 2 = 116 415 438 255 178 742 110 + 1;
  • 116 415 438 255 178 742 110 ÷ 2 = 58 207 719 127 589 371 055 + 0;
  • 58 207 719 127 589 371 055 ÷ 2 = 29 103 859 563 794 685 527 + 1;
  • 29 103 859 563 794 685 527 ÷ 2 = 14 551 929 781 897 342 763 + 1;
  • 14 551 929 781 897 342 763 ÷ 2 = 7 275 964 890 948 671 381 + 1;
  • 7 275 964 890 948 671 381 ÷ 2 = 3 637 982 445 474 335 690 + 1;
  • 3 637 982 445 474 335 690 ÷ 2 = 1 818 991 222 737 167 845 + 0;
  • 1 818 991 222 737 167 845 ÷ 2 = 909 495 611 368 583 922 + 1;
  • 909 495 611 368 583 922 ÷ 2 = 454 747 805 684 291 961 + 0;
  • 454 747 805 684 291 961 ÷ 2 = 227 373 902 842 145 980 + 1;
  • 227 373 902 842 145 980 ÷ 2 = 113 686 951 421 072 990 + 0;
  • 113 686 951 421 072 990 ÷ 2 = 56 843 475 710 536 495 + 0;
  • 56 843 475 710 536 495 ÷ 2 = 28 421 737 855 268 247 + 1;
  • 28 421 737 855 268 247 ÷ 2 = 14 210 868 927 634 123 + 1;
  • 14 210 868 927 634 123 ÷ 2 = 7 105 434 463 817 061 + 1;
  • 7 105 434 463 817 061 ÷ 2 = 3 552 717 231 908 530 + 1;
  • 3 552 717 231 908 530 ÷ 2 = 1 776 358 615 954 265 + 0;
  • 1 776 358 615 954 265 ÷ 2 = 888 179 307 977 132 + 1;
  • 888 179 307 977 132 ÷ 2 = 444 089 653 988 566 + 0;
  • 444 089 653 988 566 ÷ 2 = 222 044 826 994 283 + 0;
  • 222 044 826 994 283 ÷ 2 = 111 022 413 497 141 + 1;
  • 111 022 413 497 141 ÷ 2 = 55 511 206 748 570 + 1;
  • 55 511 206 748 570 ÷ 2 = 27 755 603 374 285 + 0;
  • 27 755 603 374 285 ÷ 2 = 13 877 801 687 142 + 1;
  • 13 877 801 687 142 ÷ 2 = 6 938 900 843 571 + 0;
  • 6 938 900 843 571 ÷ 2 = 3 469 450 421 785 + 1;
  • 3 469 450 421 785 ÷ 2 = 1 734 725 210 892 + 1;
  • 1 734 725 210 892 ÷ 2 = 867 362 605 446 + 0;
  • 867 362 605 446 ÷ 2 = 433 681 302 723 + 0;
  • 433 681 302 723 ÷ 2 = 216 840 651 361 + 1;
  • 216 840 651 361 ÷ 2 = 108 420 325 680 + 1;
  • 108 420 325 680 ÷ 2 = 54 210 162 840 + 0;
  • 54 210 162 840 ÷ 2 = 27 105 081 420 + 0;
  • 27 105 081 420 ÷ 2 = 13 552 540 710 + 0;
  • 13 552 540 710 ÷ 2 = 6 776 270 355 + 0;
  • 6 776 270 355 ÷ 2 = 3 388 135 177 + 1;
  • 3 388 135 177 ÷ 2 = 1 694 067 588 + 1;
  • 1 694 067 588 ÷ 2 = 847 033 794 + 0;
  • 847 033 794 ÷ 2 = 423 516 897 + 0;
  • 423 516 897 ÷ 2 = 211 758 448 + 1;
  • 211 758 448 ÷ 2 = 105 879 224 + 0;
  • 105 879 224 ÷ 2 = 52 939 612 + 0;
  • 52 939 612 ÷ 2 = 26 469 806 + 0;
  • 26 469 806 ÷ 2 = 13 234 903 + 0;
  • 13 234 903 ÷ 2 = 6 617 451 + 1;
  • 6 617 451 ÷ 2 = 3 308 725 + 1;
  • 3 308 725 ÷ 2 = 1 654 362 + 1;
  • 1 654 362 ÷ 2 = 827 181 + 0;
  • 827 181 ÷ 2 = 413 590 + 1;
  • 413 590 ÷ 2 = 206 795 + 0;
  • 206 795 ÷ 2 = 103 397 + 1;
  • 103 397 ÷ 2 = 51 698 + 1;
  • 51 698 ÷ 2 = 25 849 + 0;
  • 25 849 ÷ 2 = 12 924 + 1;
  • 12 924 ÷ 2 = 6 462 + 0;
  • 6 462 ÷ 2 = 3 231 + 0;
  • 3 231 ÷ 2 = 1 615 + 1;
  • 1 615 ÷ 2 = 807 + 1;
  • 807 ÷ 2 = 403 + 1;
  • 403 ÷ 2 = 201 + 1;
  • 201 ÷ 2 = 100 + 1;
  • 100 ÷ 2 = 50 + 0;
  • 50 ÷ 2 = 25 + 0;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

1 000 001 000 110 999 999 999 999 999 294(10) =


1100 1001 1111 0010 1101 0111 0000 1001 1000 0110 0110 1011 0010 1111 0010 1011 1101 0111 0101 0101 1011 1111 1101 0011 1110(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 99 positions to the left, so that only one non zero digit remains to the left of it:


1 000 001 000 110 999 999 999 999 999 294(10) =


1100 1001 1111 0010 1101 0111 0000 1001 1000 0110 0110 1011 0010 1111 0010 1011 1101 0111 0101 0101 1011 1111 1101 0011 1110(2) =


1100 1001 1111 0010 1101 0111 0000 1001 1000 0110 0110 1011 0010 1111 0010 1011 1101 0111 0101 0101 1011 1111 1101 0011 1110(2) × 20 =


1.1001 0011 1110 0101 1010 1110 0001 0011 0000 1100 1101 0110 0101 1110 0101 0111 1010 1110 1010 1011 0111 1111 1010 0111 110(2) × 299


4. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 99


Mantissa (not normalized):
1.1001 0011 1110 0101 1010 1110 0001 0011 0000 1100 1101 0110 0101 1110 0101 0111 1010 1110 1010 1011 0111 1111 1010 0111 110


5. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


99 + 2(8-1) - 1 =


(99 + 127)(10) =


226(10)


6. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 226 ÷ 2 = 113 + 0;
  • 113 ÷ 2 = 56 + 1;
  • 56 ÷ 2 = 28 + 0;
  • 28 ÷ 2 = 14 + 0;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


226(10) =


1110 0010(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 100 1001 1111 0010 1101 0111 0000 1001 1000 0110 0110 1011 0010 1111 0010 1011 1101 0111 0101 0101 1011 1111 1101 0011 1110 =


100 1001 1111 0010 1101 0111


9. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1110 0010


Mantissa (23 bits) =
100 1001 1111 0010 1101 0111


Decimal number 1 000 001 000 110 999 999 999 999 999 294 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1110 0010 - 100 1001 1111 0010 1101 0111


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111