1 000 001 101 109 999 999 999 999 999 414 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 1 000 001 101 109 999 999 999 999 999 414(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
1 000 001 101 109 999 999 999 999 999 414(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 1 000 001 101 109 999 999 999 999 999 414 ÷ 2 = 500 000 550 554 999 999 999 999 999 707 + 0;
  • 500 000 550 554 999 999 999 999 999 707 ÷ 2 = 250 000 275 277 499 999 999 999 999 853 + 1;
  • 250 000 275 277 499 999 999 999 999 853 ÷ 2 = 125 000 137 638 749 999 999 999 999 926 + 1;
  • 125 000 137 638 749 999 999 999 999 926 ÷ 2 = 62 500 068 819 374 999 999 999 999 963 + 0;
  • 62 500 068 819 374 999 999 999 999 963 ÷ 2 = 31 250 034 409 687 499 999 999 999 981 + 1;
  • 31 250 034 409 687 499 999 999 999 981 ÷ 2 = 15 625 017 204 843 749 999 999 999 990 + 1;
  • 15 625 017 204 843 749 999 999 999 990 ÷ 2 = 7 812 508 602 421 874 999 999 999 995 + 0;
  • 7 812 508 602 421 874 999 999 999 995 ÷ 2 = 3 906 254 301 210 937 499 999 999 997 + 1;
  • 3 906 254 301 210 937 499 999 999 997 ÷ 2 = 1 953 127 150 605 468 749 999 999 998 + 1;
  • 1 953 127 150 605 468 749 999 999 998 ÷ 2 = 976 563 575 302 734 374 999 999 999 + 0;
  • 976 563 575 302 734 374 999 999 999 ÷ 2 = 488 281 787 651 367 187 499 999 999 + 1;
  • 488 281 787 651 367 187 499 999 999 ÷ 2 = 244 140 893 825 683 593 749 999 999 + 1;
  • 244 140 893 825 683 593 749 999 999 ÷ 2 = 122 070 446 912 841 796 874 999 999 + 1;
  • 122 070 446 912 841 796 874 999 999 ÷ 2 = 61 035 223 456 420 898 437 499 999 + 1;
  • 61 035 223 456 420 898 437 499 999 ÷ 2 = 30 517 611 728 210 449 218 749 999 + 1;
  • 30 517 611 728 210 449 218 749 999 ÷ 2 = 15 258 805 864 105 224 609 374 999 + 1;
  • 15 258 805 864 105 224 609 374 999 ÷ 2 = 7 629 402 932 052 612 304 687 499 + 1;
  • 7 629 402 932 052 612 304 687 499 ÷ 2 = 3 814 701 466 026 306 152 343 749 + 1;
  • 3 814 701 466 026 306 152 343 749 ÷ 2 = 1 907 350 733 013 153 076 171 874 + 1;
  • 1 907 350 733 013 153 076 171 874 ÷ 2 = 953 675 366 506 576 538 085 937 + 0;
  • 953 675 366 506 576 538 085 937 ÷ 2 = 476 837 683 253 288 269 042 968 + 1;
  • 476 837 683 253 288 269 042 968 ÷ 2 = 238 418 841 626 644 134 521 484 + 0;
  • 238 418 841 626 644 134 521 484 ÷ 2 = 119 209 420 813 322 067 260 742 + 0;
  • 119 209 420 813 322 067 260 742 ÷ 2 = 59 604 710 406 661 033 630 371 + 0;
  • 59 604 710 406 661 033 630 371 ÷ 2 = 29 802 355 203 330 516 815 185 + 1;
  • 29 802 355 203 330 516 815 185 ÷ 2 = 14 901 177 601 665 258 407 592 + 1;
  • 14 901 177 601 665 258 407 592 ÷ 2 = 7 450 588 800 832 629 203 796 + 0;
  • 7 450 588 800 832 629 203 796 ÷ 2 = 3 725 294 400 416 314 601 898 + 0;
  • 3 725 294 400 416 314 601 898 ÷ 2 = 1 862 647 200 208 157 300 949 + 0;
  • 1 862 647 200 208 157 300 949 ÷ 2 = 931 323 600 104 078 650 474 + 1;
  • 931 323 600 104 078 650 474 ÷ 2 = 465 661 800 052 039 325 237 + 0;
  • 465 661 800 052 039 325 237 ÷ 2 = 232 830 900 026 019 662 618 + 1;
  • 232 830 900 026 019 662 618 ÷ 2 = 116 415 450 013 009 831 309 + 0;
  • 116 415 450 013 009 831 309 ÷ 2 = 58 207 725 006 504 915 654 + 1;
  • 58 207 725 006 504 915 654 ÷ 2 = 29 103 862 503 252 457 827 + 0;
  • 29 103 862 503 252 457 827 ÷ 2 = 14 551 931 251 626 228 913 + 1;
  • 14 551 931 251 626 228 913 ÷ 2 = 7 275 965 625 813 114 456 + 1;
  • 7 275 965 625 813 114 456 ÷ 2 = 3 637 982 812 906 557 228 + 0;
  • 3 637 982 812 906 557 228 ÷ 2 = 1 818 991 406 453 278 614 + 0;
  • 1 818 991 406 453 278 614 ÷ 2 = 909 495 703 226 639 307 + 0;
  • 909 495 703 226 639 307 ÷ 2 = 454 747 851 613 319 653 + 1;
  • 454 747 851 613 319 653 ÷ 2 = 227 373 925 806 659 826 + 1;
  • 227 373 925 806 659 826 ÷ 2 = 113 686 962 903 329 913 + 0;
  • 113 686 962 903 329 913 ÷ 2 = 56 843 481 451 664 956 + 1;
  • 56 843 481 451 664 956 ÷ 2 = 28 421 740 725 832 478 + 0;
  • 28 421 740 725 832 478 ÷ 2 = 14 210 870 362 916 239 + 0;
  • 14 210 870 362 916 239 ÷ 2 = 7 105 435 181 458 119 + 1;
  • 7 105 435 181 458 119 ÷ 2 = 3 552 717 590 729 059 + 1;
  • 3 552 717 590 729 059 ÷ 2 = 1 776 358 795 364 529 + 1;
  • 1 776 358 795 364 529 ÷ 2 = 888 179 397 682 264 + 1;
  • 888 179 397 682 264 ÷ 2 = 444 089 698 841 132 + 0;
  • 444 089 698 841 132 ÷ 2 = 222 044 849 420 566 + 0;
  • 222 044 849 420 566 ÷ 2 = 111 022 424 710 283 + 0;
  • 111 022 424 710 283 ÷ 2 = 55 511 212 355 141 + 1;
  • 55 511 212 355 141 ÷ 2 = 27 755 606 177 570 + 1;
  • 27 755 606 177 570 ÷ 2 = 13 877 803 088 785 + 0;
  • 13 877 803 088 785 ÷ 2 = 6 938 901 544 392 + 1;
  • 6 938 901 544 392 ÷ 2 = 3 469 450 772 196 + 0;
  • 3 469 450 772 196 ÷ 2 = 1 734 725 386 098 + 0;
  • 1 734 725 386 098 ÷ 2 = 867 362 693 049 + 0;
  • 867 362 693 049 ÷ 2 = 433 681 346 524 + 1;
  • 433 681 346 524 ÷ 2 = 216 840 673 262 + 0;
  • 216 840 673 262 ÷ 2 = 108 420 336 631 + 0;
  • 108 420 336 631 ÷ 2 = 54 210 168 315 + 1;
  • 54 210 168 315 ÷ 2 = 27 105 084 157 + 1;
  • 27 105 084 157 ÷ 2 = 13 552 542 078 + 1;
  • 13 552 542 078 ÷ 2 = 6 776 271 039 + 0;
  • 6 776 271 039 ÷ 2 = 3 388 135 519 + 1;
  • 3 388 135 519 ÷ 2 = 1 694 067 759 + 1;
  • 1 694 067 759 ÷ 2 = 847 033 879 + 1;
  • 847 033 879 ÷ 2 = 423 516 939 + 1;
  • 423 516 939 ÷ 2 = 211 758 469 + 1;
  • 211 758 469 ÷ 2 = 105 879 234 + 1;
  • 105 879 234 ÷ 2 = 52 939 617 + 0;
  • 52 939 617 ÷ 2 = 26 469 808 + 1;
  • 26 469 808 ÷ 2 = 13 234 904 + 0;
  • 13 234 904 ÷ 2 = 6 617 452 + 0;
  • 6 617 452 ÷ 2 = 3 308 726 + 0;
  • 3 308 726 ÷ 2 = 1 654 363 + 0;
  • 1 654 363 ÷ 2 = 827 181 + 1;
  • 827 181 ÷ 2 = 413 590 + 1;
  • 413 590 ÷ 2 = 206 795 + 0;
  • 206 795 ÷ 2 = 103 397 + 1;
  • 103 397 ÷ 2 = 51 698 + 1;
  • 51 698 ÷ 2 = 25 849 + 0;
  • 25 849 ÷ 2 = 12 924 + 1;
  • 12 924 ÷ 2 = 6 462 + 0;
  • 6 462 ÷ 2 = 3 231 + 0;
  • 3 231 ÷ 2 = 1 615 + 1;
  • 1 615 ÷ 2 = 807 + 1;
  • 807 ÷ 2 = 403 + 1;
  • 403 ÷ 2 = 201 + 1;
  • 201 ÷ 2 = 100 + 1;
  • 100 ÷ 2 = 50 + 0;
  • 50 ÷ 2 = 25 + 0;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

1 000 001 101 109 999 999 999 999 999 414(10) =


1100 1001 1111 0010 1101 1000 0101 1111 1011 1001 0001 0110 0011 1100 1011 0001 1010 1010 0011 0001 0111 1111 1101 1011 0110(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 99 positions to the left, so that only one non zero digit remains to the left of it:


1 000 001 101 109 999 999 999 999 999 414(10) =


1100 1001 1111 0010 1101 1000 0101 1111 1011 1001 0001 0110 0011 1100 1011 0001 1010 1010 0011 0001 0111 1111 1101 1011 0110(2) =


1100 1001 1111 0010 1101 1000 0101 1111 1011 1001 0001 0110 0011 1100 1011 0001 1010 1010 0011 0001 0111 1111 1101 1011 0110(2) × 20 =


1.1001 0011 1110 0101 1011 0000 1011 1111 0111 0010 0010 1100 0111 1001 0110 0011 0101 0100 0110 0010 1111 1111 1011 0110 110(2) × 299


4. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 99


Mantissa (not normalized):
1.1001 0011 1110 0101 1011 0000 1011 1111 0111 0010 0010 1100 0111 1001 0110 0011 0101 0100 0110 0010 1111 1111 1011 0110 110


5. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


99 + 2(8-1) - 1 =


(99 + 127)(10) =


226(10)


6. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 226 ÷ 2 = 113 + 0;
  • 113 ÷ 2 = 56 + 1;
  • 56 ÷ 2 = 28 + 0;
  • 28 ÷ 2 = 14 + 0;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


226(10) =


1110 0010(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 100 1001 1111 0010 1101 1000 0101 1111 1011 1001 0001 0110 0011 1100 1011 0001 1010 1010 0011 0001 0111 1111 1101 1011 0110 =


100 1001 1111 0010 1101 1000


9. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1110 0010


Mantissa (23 bits) =
100 1001 1111 0010 1101 1000


Decimal number 1 000 001 101 109 999 999 999 999 999 414 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1110 0010 - 100 1001 1111 0010 1101 1000


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111