1 010 010 110 010 100 000 100 109 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 1 010 010 110 010 100 000 100 109(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
1 010 010 110 010 100 000 100 109(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 1 010 010 110 010 100 000 100 109 ÷ 2 = 505 005 055 005 050 000 050 054 + 1;
  • 505 005 055 005 050 000 050 054 ÷ 2 = 252 502 527 502 525 000 025 027 + 0;
  • 252 502 527 502 525 000 025 027 ÷ 2 = 126 251 263 751 262 500 012 513 + 1;
  • 126 251 263 751 262 500 012 513 ÷ 2 = 63 125 631 875 631 250 006 256 + 1;
  • 63 125 631 875 631 250 006 256 ÷ 2 = 31 562 815 937 815 625 003 128 + 0;
  • 31 562 815 937 815 625 003 128 ÷ 2 = 15 781 407 968 907 812 501 564 + 0;
  • 15 781 407 968 907 812 501 564 ÷ 2 = 7 890 703 984 453 906 250 782 + 0;
  • 7 890 703 984 453 906 250 782 ÷ 2 = 3 945 351 992 226 953 125 391 + 0;
  • 3 945 351 992 226 953 125 391 ÷ 2 = 1 972 675 996 113 476 562 695 + 1;
  • 1 972 675 996 113 476 562 695 ÷ 2 = 986 337 998 056 738 281 347 + 1;
  • 986 337 998 056 738 281 347 ÷ 2 = 493 168 999 028 369 140 673 + 1;
  • 493 168 999 028 369 140 673 ÷ 2 = 246 584 499 514 184 570 336 + 1;
  • 246 584 499 514 184 570 336 ÷ 2 = 123 292 249 757 092 285 168 + 0;
  • 123 292 249 757 092 285 168 ÷ 2 = 61 646 124 878 546 142 584 + 0;
  • 61 646 124 878 546 142 584 ÷ 2 = 30 823 062 439 273 071 292 + 0;
  • 30 823 062 439 273 071 292 ÷ 2 = 15 411 531 219 636 535 646 + 0;
  • 15 411 531 219 636 535 646 ÷ 2 = 7 705 765 609 818 267 823 + 0;
  • 7 705 765 609 818 267 823 ÷ 2 = 3 852 882 804 909 133 911 + 1;
  • 3 852 882 804 909 133 911 ÷ 2 = 1 926 441 402 454 566 955 + 1;
  • 1 926 441 402 454 566 955 ÷ 2 = 963 220 701 227 283 477 + 1;
  • 963 220 701 227 283 477 ÷ 2 = 481 610 350 613 641 738 + 1;
  • 481 610 350 613 641 738 ÷ 2 = 240 805 175 306 820 869 + 0;
  • 240 805 175 306 820 869 ÷ 2 = 120 402 587 653 410 434 + 1;
  • 120 402 587 653 410 434 ÷ 2 = 60 201 293 826 705 217 + 0;
  • 60 201 293 826 705 217 ÷ 2 = 30 100 646 913 352 608 + 1;
  • 30 100 646 913 352 608 ÷ 2 = 15 050 323 456 676 304 + 0;
  • 15 050 323 456 676 304 ÷ 2 = 7 525 161 728 338 152 + 0;
  • 7 525 161 728 338 152 ÷ 2 = 3 762 580 864 169 076 + 0;
  • 3 762 580 864 169 076 ÷ 2 = 1 881 290 432 084 538 + 0;
  • 1 881 290 432 084 538 ÷ 2 = 940 645 216 042 269 + 0;
  • 940 645 216 042 269 ÷ 2 = 470 322 608 021 134 + 1;
  • 470 322 608 021 134 ÷ 2 = 235 161 304 010 567 + 0;
  • 235 161 304 010 567 ÷ 2 = 117 580 652 005 283 + 1;
  • 117 580 652 005 283 ÷ 2 = 58 790 326 002 641 + 1;
  • 58 790 326 002 641 ÷ 2 = 29 395 163 001 320 + 1;
  • 29 395 163 001 320 ÷ 2 = 14 697 581 500 660 + 0;
  • 14 697 581 500 660 ÷ 2 = 7 348 790 750 330 + 0;
  • 7 348 790 750 330 ÷ 2 = 3 674 395 375 165 + 0;
  • 3 674 395 375 165 ÷ 2 = 1 837 197 687 582 + 1;
  • 1 837 197 687 582 ÷ 2 = 918 598 843 791 + 0;
  • 918 598 843 791 ÷ 2 = 459 299 421 895 + 1;
  • 459 299 421 895 ÷ 2 = 229 649 710 947 + 1;
  • 229 649 710 947 ÷ 2 = 114 824 855 473 + 1;
  • 114 824 855 473 ÷ 2 = 57 412 427 736 + 1;
  • 57 412 427 736 ÷ 2 = 28 706 213 868 + 0;
  • 28 706 213 868 ÷ 2 = 14 353 106 934 + 0;
  • 14 353 106 934 ÷ 2 = 7 176 553 467 + 0;
  • 7 176 553 467 ÷ 2 = 3 588 276 733 + 1;
  • 3 588 276 733 ÷ 2 = 1 794 138 366 + 1;
  • 1 794 138 366 ÷ 2 = 897 069 183 + 0;
  • 897 069 183 ÷ 2 = 448 534 591 + 1;
  • 448 534 591 ÷ 2 = 224 267 295 + 1;
  • 224 267 295 ÷ 2 = 112 133 647 + 1;
  • 112 133 647 ÷ 2 = 56 066 823 + 1;
  • 56 066 823 ÷ 2 = 28 033 411 + 1;
  • 28 033 411 ÷ 2 = 14 016 705 + 1;
  • 14 016 705 ÷ 2 = 7 008 352 + 1;
  • 7 008 352 ÷ 2 = 3 504 176 + 0;
  • 3 504 176 ÷ 2 = 1 752 088 + 0;
  • 1 752 088 ÷ 2 = 876 044 + 0;
  • 876 044 ÷ 2 = 438 022 + 0;
  • 438 022 ÷ 2 = 219 011 + 0;
  • 219 011 ÷ 2 = 109 505 + 1;
  • 109 505 ÷ 2 = 54 752 + 1;
  • 54 752 ÷ 2 = 27 376 + 0;
  • 27 376 ÷ 2 = 13 688 + 0;
  • 13 688 ÷ 2 = 6 844 + 0;
  • 6 844 ÷ 2 = 3 422 + 0;
  • 3 422 ÷ 2 = 1 711 + 0;
  • 1 711 ÷ 2 = 855 + 1;
  • 855 ÷ 2 = 427 + 1;
  • 427 ÷ 2 = 213 + 1;
  • 213 ÷ 2 = 106 + 1;
  • 106 ÷ 2 = 53 + 0;
  • 53 ÷ 2 = 26 + 1;
  • 26 ÷ 2 = 13 + 0;
  • 13 ÷ 2 = 6 + 1;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the positive number.

Take all the remainders starting from the bottom of the list constructed above.

1 010 010 110 010 100 000 100 109(10) =


1101 0101 1110 0000 1100 0001 1111 1101 1000 1111 0100 0111 0100 0001 0101 1110 0000 1111 0000 1101(2)


3. Normalize the binary representation of the number.

Shift the decimal mark 79 positions to the left, so that only one non zero digit remains to the left of it:


1 010 010 110 010 100 000 100 109(10) =


1101 0101 1110 0000 1100 0001 1111 1101 1000 1111 0100 0111 0100 0001 0101 1110 0000 1111 0000 1101(2) =


1101 0101 1110 0000 1100 0001 1111 1101 1000 1111 0100 0111 0100 0001 0101 1110 0000 1111 0000 1101(2) × 20 =


1.1010 1011 1100 0001 1000 0011 1111 1011 0001 1110 1000 1110 1000 0010 1011 1100 0001 1110 0001 101(2) × 279


4. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 79


Mantissa (not normalized):
1.1010 1011 1100 0001 1000 0011 1111 1011 0001 1110 1000 1110 1000 0010 1011 1100 0001 1110 0001 101


5. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


79 + 2(8-1) - 1 =


(79 + 127)(10) =


206(10)


6. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 206 ÷ 2 = 103 + 0;
  • 103 ÷ 2 = 51 + 1;
  • 51 ÷ 2 = 25 + 1;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

7. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


206(10) =


1100 1110(2)


8. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 101 0101 1110 0000 1100 0001 1111 1101 1000 1111 0100 0111 0100 0001 0101 1110 0000 1111 0000 1101 =


101 0101 1110 0000 1100 0001


9. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1100 1110


Mantissa (23 bits) =
101 0101 1110 0000 1100 0001


Decimal number 1 010 010 110 010 100 000 100 109 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1100 1110 - 101 0101 1110 0000 1100 0001


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111