10 111 109 999 999 999 999 999 999 999 798 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 10 111 109 999 999 999 999 999 999 999 798(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
10 111 109 999 999 999 999 999 999 999 798(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 10 111 109 999 999 999 999 999 999 999 798 ÷ 2 = 5 055 554 999 999 999 999 999 999 999 899 + 0;
  • 5 055 554 999 999 999 999 999 999 999 899 ÷ 2 = 2 527 777 499 999 999 999 999 999 999 949 + 1;
  • 2 527 777 499 999 999 999 999 999 999 949 ÷ 2 = 1 263 888 749 999 999 999 999 999 999 974 + 1;
  • 1 263 888 749 999 999 999 999 999 999 974 ÷ 2 = 631 944 374 999 999 999 999 999 999 987 + 0;
  • 631 944 374 999 999 999 999 999 999 987 ÷ 2 = 315 972 187 499 999 999 999 999 999 993 + 1;
  • 315 972 187 499 999 999 999 999 999 993 ÷ 2 = 157 986 093 749 999 999 999 999 999 996 + 1;
  • 157 986 093 749 999 999 999 999 999 996 ÷ 2 = 78 993 046 874 999 999 999 999 999 998 + 0;
  • 78 993 046 874 999 999 999 999 999 998 ÷ 2 = 39 496 523 437 499 999 999 999 999 999 + 0;
  • 39 496 523 437 499 999 999 999 999 999 ÷ 2 = 19 748 261 718 749 999 999 999 999 999 + 1;
  • 19 748 261 718 749 999 999 999 999 999 ÷ 2 = 9 874 130 859 374 999 999 999 999 999 + 1;
  • 9 874 130 859 374 999 999 999 999 999 ÷ 2 = 4 937 065 429 687 499 999 999 999 999 + 1;
  • 4 937 065 429 687 499 999 999 999 999 ÷ 2 = 2 468 532 714 843 749 999 999 999 999 + 1;
  • 2 468 532 714 843 749 999 999 999 999 ÷ 2 = 1 234 266 357 421 874 999 999 999 999 + 1;
  • 1 234 266 357 421 874 999 999 999 999 ÷ 2 = 617 133 178 710 937 499 999 999 999 + 1;
  • 617 133 178 710 937 499 999 999 999 ÷ 2 = 308 566 589 355 468 749 999 999 999 + 1;
  • 308 566 589 355 468 749 999 999 999 ÷ 2 = 154 283 294 677 734 374 999 999 999 + 1;
  • 154 283 294 677 734 374 999 999 999 ÷ 2 = 77 141 647 338 867 187 499 999 999 + 1;
  • 77 141 647 338 867 187 499 999 999 ÷ 2 = 38 570 823 669 433 593 749 999 999 + 1;
  • 38 570 823 669 433 593 749 999 999 ÷ 2 = 19 285 411 834 716 796 874 999 999 + 1;
  • 19 285 411 834 716 796 874 999 999 ÷ 2 = 9 642 705 917 358 398 437 499 999 + 1;
  • 9 642 705 917 358 398 437 499 999 ÷ 2 = 4 821 352 958 679 199 218 749 999 + 1;
  • 4 821 352 958 679 199 218 749 999 ÷ 2 = 2 410 676 479 339 599 609 374 999 + 1;
  • 2 410 676 479 339 599 609 374 999 ÷ 2 = 1 205 338 239 669 799 804 687 499 + 1;
  • 1 205 338 239 669 799 804 687 499 ÷ 2 = 602 669 119 834 899 902 343 749 + 1;
  • 602 669 119 834 899 902 343 749 ÷ 2 = 301 334 559 917 449 951 171 874 + 1;
  • 301 334 559 917 449 951 171 874 ÷ 2 = 150 667 279 958 724 975 585 937 + 0;
  • 150 667 279 958 724 975 585 937 ÷ 2 = 75 333 639 979 362 487 792 968 + 1;
  • 75 333 639 979 362 487 792 968 ÷ 2 = 37 666 819 989 681 243 896 484 + 0;
  • 37 666 819 989 681 243 896 484 ÷ 2 = 18 833 409 994 840 621 948 242 + 0;
  • 18 833 409 994 840 621 948 242 ÷ 2 = 9 416 704 997 420 310 974 121 + 0;
  • 9 416 704 997 420 310 974 121 ÷ 2 = 4 708 352 498 710 155 487 060 + 1;
  • 4 708 352 498 710 155 487 060 ÷ 2 = 2 354 176 249 355 077 743 530 + 0;
  • 2 354 176 249 355 077 743 530 ÷ 2 = 1 177 088 124 677 538 871 765 + 0;
  • 1 177 088 124 677 538 871 765 ÷ 2 = 588 544 062 338 769 435 882 + 1;
  • 588 544 062 338 769 435 882 ÷ 2 = 294 272 031 169 384 717 941 + 0;
  • 294 272 031 169 384 717 941 ÷ 2 = 147 136 015 584 692 358 970 + 1;
  • 147 136 015 584 692 358 970 ÷ 2 = 73 568 007 792 346 179 485 + 0;
  • 73 568 007 792 346 179 485 ÷ 2 = 36 784 003 896 173 089 742 + 1;
  • 36 784 003 896 173 089 742 ÷ 2 = 18 392 001 948 086 544 871 + 0;
  • 18 392 001 948 086 544 871 ÷ 2 = 9 196 000 974 043 272 435 + 1;
  • 9 196 000 974 043 272 435 ÷ 2 = 4 598 000 487 021 636 217 + 1;
  • 4 598 000 487 021 636 217 ÷ 2 = 2 299 000 243 510 818 108 + 1;
  • 2 299 000 243 510 818 108 ÷ 2 = 1 149 500 121 755 409 054 + 0;
  • 1 149 500 121 755 409 054 ÷ 2 = 574 750 060 877 704 527 + 0;
  • 574 750 060 877 704 527 ÷ 2 = 287 375 030 438 852 263 + 1;
  • 287 375 030 438 852 263 ÷ 2 = 143 687 515 219 426 131 + 1;
  • 143 687 515 219 426 131 ÷ 2 = 71 843 757 609 713 065 + 1;
  • 71 843 757 609 713 065 ÷ 2 = 35 921 878 804 856 532 + 1;
  • 35 921 878 804 856 532 ÷ 2 = 17 960 939 402 428 266 + 0;
  • 17 960 939 402 428 266 ÷ 2 = 8 980 469 701 214 133 + 0;
  • 8 980 469 701 214 133 ÷ 2 = 4 490 234 850 607 066 + 1;
  • 4 490 234 850 607 066 ÷ 2 = 2 245 117 425 303 533 + 0;
  • 2 245 117 425 303 533 ÷ 2 = 1 122 558 712 651 766 + 1;
  • 1 122 558 712 651 766 ÷ 2 = 561 279 356 325 883 + 0;
  • 561 279 356 325 883 ÷ 2 = 280 639 678 162 941 + 1;
  • 280 639 678 162 941 ÷ 2 = 140 319 839 081 470 + 1;
  • 140 319 839 081 470 ÷ 2 = 70 159 919 540 735 + 0;
  • 70 159 919 540 735 ÷ 2 = 35 079 959 770 367 + 1;
  • 35 079 959 770 367 ÷ 2 = 17 539 979 885 183 + 1;
  • 17 539 979 885 183 ÷ 2 = 8 769 989 942 591 + 1;
  • 8 769 989 942 591 ÷ 2 = 4 384 994 971 295 + 1;
  • 4 384 994 971 295 ÷ 2 = 2 192 497 485 647 + 1;
  • 2 192 497 485 647 ÷ 2 = 1 096 248 742 823 + 1;
  • 1 096 248 742 823 ÷ 2 = 548 124 371 411 + 1;
  • 548 124 371 411 ÷ 2 = 274 062 185 705 + 1;
  • 274 062 185 705 ÷ 2 = 137 031 092 852 + 1;
  • 137 031 092 852 ÷ 2 = 68 515 546 426 + 0;
  • 68 515 546 426 ÷ 2 = 34 257 773 213 + 0;
  • 34 257 773 213 ÷ 2 = 17 128 886 606 + 1;
  • 17 128 886 606 ÷ 2 = 8 564 443 303 + 0;
  • 8 564 443 303 ÷ 2 = 4 282 221 651 + 1;
  • 4 282 221 651 ÷ 2 = 2 141 110 825 + 1;
  • 2 141 110 825 ÷ 2 = 1 070 555 412 + 1;
  • 1 070 555 412 ÷ 2 = 535 277 706 + 0;
  • 535 277 706 ÷ 2 = 267 638 853 + 0;
  • 267 638 853 ÷ 2 = 133 819 426 + 1;
  • 133 819 426 ÷ 2 = 66 909 713 + 0;
  • 66 909 713 ÷ 2 = 33 454 856 + 1;
  • 33 454 856 ÷ 2 = 16 727 428 + 0;
  • 16 727 428 ÷ 2 = 8 363 714 + 0;
  • 8 363 714 ÷ 2 = 4 181 857 + 0;
  • 4 181 857 ÷ 2 = 2 090 928 + 1;
  • 2 090 928 ÷ 2 = 1 045 464 + 0;
  • 1 045 464 ÷ 2 = 522 732 + 0;
  • 522 732 ÷ 2 = 261 366 + 0;
  • 261 366 ÷ 2 = 130 683 + 0;
  • 130 683 ÷ 2 = 65 341 + 1;
  • 65 341 ÷ 2 = 32 670 + 1;
  • 32 670 ÷ 2 = 16 335 + 0;
  • 16 335 ÷ 2 = 8 167 + 1;
  • 8 167 ÷ 2 = 4 083 + 1;
  • 4 083 ÷ 2 = 2 041 + 1;
  • 2 041 ÷ 2 = 1 020 + 1;
  • 1 020 ÷ 2 = 510 + 0;
  • 510 ÷ 2 = 255 + 0;
  • 255 ÷ 2 = 127 + 1;
  • 127 ÷ 2 = 63 + 1;
  • 63 ÷ 2 = 31 + 1;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

10 111 109 999 999 999 999 999 999 999 798(10) =


111 1111 1001 1110 1100 0010 0010 1001 1101 0011 1111 1110 1101 0100 1111 0011 1010 1010 0100 0101 1111 1111 1111 1111 0011 0110(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 102 positions to the left, so that only one non zero digit remains to the left of it:


10 111 109 999 999 999 999 999 999 999 798(10) =


111 1111 1001 1110 1100 0010 0010 1001 1101 0011 1111 1110 1101 0100 1111 0011 1010 1010 0100 0101 1111 1111 1111 1111 0011 0110(2) =


111 1111 1001 1110 1100 0010 0010 1001 1101 0011 1111 1110 1101 0100 1111 0011 1010 1010 0100 0101 1111 1111 1111 1111 0011 0110(2) × 20 =


1.1111 1110 0111 1011 0000 1000 1010 0111 0100 1111 1111 1011 0101 0011 1100 1110 1010 1001 0001 0111 1111 1111 1111 1100 1101 10(2) × 2102


4. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 102


Mantissa (not normalized):
1.1111 1110 0111 1011 0000 1000 1010 0111 0100 1111 1111 1011 0101 0011 1100 1110 1010 1001 0001 0111 1111 1111 1111 1100 1101 10


5. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


102 + 2(8-1) - 1 =


(102 + 127)(10) =


229(10)


6. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 229 ÷ 2 = 114 + 1;
  • 114 ÷ 2 = 57 + 0;
  • 57 ÷ 2 = 28 + 1;
  • 28 ÷ 2 = 14 + 0;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


229(10) =


1110 0101(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 111 1111 0011 1101 1000 0100 010 1001 1101 0011 1111 1110 1101 0100 1111 0011 1010 1010 0100 0101 1111 1111 1111 1111 0011 0110 =


111 1111 0011 1101 1000 0100


9. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1110 0101


Mantissa (23 bits) =
111 1111 0011 1101 1000 0100


Decimal number 10 111 109 999 999 999 999 999 999 999 798 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1110 0101 - 111 1111 0011 1101 1000 0100


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111