34 028 349 133 317 412 298 432 094 444 834 721 134.64 Converted to 32 Bit Single Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 34 028 349 133 317 412 298 432 094 444 834 721 134.64(10) to 32 bit single precision IEEE 754 binary floating point representation standard (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

What are the steps to convert decimal number
34 028 349 133 317 412 298 432 094 444 834 721 134.64(10) to 32 bit single precision IEEE 754 binary floating point representation (1 bit for sign, 8 bits for exponent, 23 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 34 028 349 133 317 412 298 432 094 444 834 721 134.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 34 028 349 133 317 412 298 432 094 444 834 721 134 ÷ 2 = 17 014 174 566 658 706 149 216 047 222 417 360 567 + 0;
  • 17 014 174 566 658 706 149 216 047 222 417 360 567 ÷ 2 = 8 507 087 283 329 353 074 608 023 611 208 680 283 + 1;
  • 8 507 087 283 329 353 074 608 023 611 208 680 283 ÷ 2 = 4 253 543 641 664 676 537 304 011 805 604 340 141 + 1;
  • 4 253 543 641 664 676 537 304 011 805 604 340 141 ÷ 2 = 2 126 771 820 832 338 268 652 005 902 802 170 070 + 1;
  • 2 126 771 820 832 338 268 652 005 902 802 170 070 ÷ 2 = 1 063 385 910 416 169 134 326 002 951 401 085 035 + 0;
  • 1 063 385 910 416 169 134 326 002 951 401 085 035 ÷ 2 = 531 692 955 208 084 567 163 001 475 700 542 517 + 1;
  • 531 692 955 208 084 567 163 001 475 700 542 517 ÷ 2 = 265 846 477 604 042 283 581 500 737 850 271 258 + 1;
  • 265 846 477 604 042 283 581 500 737 850 271 258 ÷ 2 = 132 923 238 802 021 141 790 750 368 925 135 629 + 0;
  • 132 923 238 802 021 141 790 750 368 925 135 629 ÷ 2 = 66 461 619 401 010 570 895 375 184 462 567 814 + 1;
  • 66 461 619 401 010 570 895 375 184 462 567 814 ÷ 2 = 33 230 809 700 505 285 447 687 592 231 283 907 + 0;
  • 33 230 809 700 505 285 447 687 592 231 283 907 ÷ 2 = 16 615 404 850 252 642 723 843 796 115 641 953 + 1;
  • 16 615 404 850 252 642 723 843 796 115 641 953 ÷ 2 = 8 307 702 425 126 321 361 921 898 057 820 976 + 1;
  • 8 307 702 425 126 321 361 921 898 057 820 976 ÷ 2 = 4 153 851 212 563 160 680 960 949 028 910 488 + 0;
  • 4 153 851 212 563 160 680 960 949 028 910 488 ÷ 2 = 2 076 925 606 281 580 340 480 474 514 455 244 + 0;
  • 2 076 925 606 281 580 340 480 474 514 455 244 ÷ 2 = 1 038 462 803 140 790 170 240 237 257 227 622 + 0;
  • 1 038 462 803 140 790 170 240 237 257 227 622 ÷ 2 = 519 231 401 570 395 085 120 118 628 613 811 + 0;
  • 519 231 401 570 395 085 120 118 628 613 811 ÷ 2 = 259 615 700 785 197 542 560 059 314 306 905 + 1;
  • 259 615 700 785 197 542 560 059 314 306 905 ÷ 2 = 129 807 850 392 598 771 280 029 657 153 452 + 1;
  • 129 807 850 392 598 771 280 029 657 153 452 ÷ 2 = 64 903 925 196 299 385 640 014 828 576 726 + 0;
  • 64 903 925 196 299 385 640 014 828 576 726 ÷ 2 = 32 451 962 598 149 692 820 007 414 288 363 + 0;
  • 32 451 962 598 149 692 820 007 414 288 363 ÷ 2 = 16 225 981 299 074 846 410 003 707 144 181 + 1;
  • 16 225 981 299 074 846 410 003 707 144 181 ÷ 2 = 8 112 990 649 537 423 205 001 853 572 090 + 1;
  • 8 112 990 649 537 423 205 001 853 572 090 ÷ 2 = 4 056 495 324 768 711 602 500 926 786 045 + 0;
  • 4 056 495 324 768 711 602 500 926 786 045 ÷ 2 = 2 028 247 662 384 355 801 250 463 393 022 + 1;
  • 2 028 247 662 384 355 801 250 463 393 022 ÷ 2 = 1 014 123 831 192 177 900 625 231 696 511 + 0;
  • 1 014 123 831 192 177 900 625 231 696 511 ÷ 2 = 507 061 915 596 088 950 312 615 848 255 + 1;
  • 507 061 915 596 088 950 312 615 848 255 ÷ 2 = 253 530 957 798 044 475 156 307 924 127 + 1;
  • 253 530 957 798 044 475 156 307 924 127 ÷ 2 = 126 765 478 899 022 237 578 153 962 063 + 1;
  • 126 765 478 899 022 237 578 153 962 063 ÷ 2 = 63 382 739 449 511 118 789 076 981 031 + 1;
  • 63 382 739 449 511 118 789 076 981 031 ÷ 2 = 31 691 369 724 755 559 394 538 490 515 + 1;
  • 31 691 369 724 755 559 394 538 490 515 ÷ 2 = 15 845 684 862 377 779 697 269 245 257 + 1;
  • 15 845 684 862 377 779 697 269 245 257 ÷ 2 = 7 922 842 431 188 889 848 634 622 628 + 1;
  • 7 922 842 431 188 889 848 634 622 628 ÷ 2 = 3 961 421 215 594 444 924 317 311 314 + 0;
  • 3 961 421 215 594 444 924 317 311 314 ÷ 2 = 1 980 710 607 797 222 462 158 655 657 + 0;
  • 1 980 710 607 797 222 462 158 655 657 ÷ 2 = 990 355 303 898 611 231 079 327 828 + 1;
  • 990 355 303 898 611 231 079 327 828 ÷ 2 = 495 177 651 949 305 615 539 663 914 + 0;
  • 495 177 651 949 305 615 539 663 914 ÷ 2 = 247 588 825 974 652 807 769 831 957 + 0;
  • 247 588 825 974 652 807 769 831 957 ÷ 2 = 123 794 412 987 326 403 884 915 978 + 1;
  • 123 794 412 987 326 403 884 915 978 ÷ 2 = 61 897 206 493 663 201 942 457 989 + 0;
  • 61 897 206 493 663 201 942 457 989 ÷ 2 = 30 948 603 246 831 600 971 228 994 + 1;
  • 30 948 603 246 831 600 971 228 994 ÷ 2 = 15 474 301 623 415 800 485 614 497 + 0;
  • 15 474 301 623 415 800 485 614 497 ÷ 2 = 7 737 150 811 707 900 242 807 248 + 1;
  • 7 737 150 811 707 900 242 807 248 ÷ 2 = 3 868 575 405 853 950 121 403 624 + 0;
  • 3 868 575 405 853 950 121 403 624 ÷ 2 = 1 934 287 702 926 975 060 701 812 + 0;
  • 1 934 287 702 926 975 060 701 812 ÷ 2 = 967 143 851 463 487 530 350 906 + 0;
  • 967 143 851 463 487 530 350 906 ÷ 2 = 483 571 925 731 743 765 175 453 + 0;
  • 483 571 925 731 743 765 175 453 ÷ 2 = 241 785 962 865 871 882 587 726 + 1;
  • 241 785 962 865 871 882 587 726 ÷ 2 = 120 892 981 432 935 941 293 863 + 0;
  • 120 892 981 432 935 941 293 863 ÷ 2 = 60 446 490 716 467 970 646 931 + 1;
  • 60 446 490 716 467 970 646 931 ÷ 2 = 30 223 245 358 233 985 323 465 + 1;
  • 30 223 245 358 233 985 323 465 ÷ 2 = 15 111 622 679 116 992 661 732 + 1;
  • 15 111 622 679 116 992 661 732 ÷ 2 = 7 555 811 339 558 496 330 866 + 0;
  • 7 555 811 339 558 496 330 866 ÷ 2 = 3 777 905 669 779 248 165 433 + 0;
  • 3 777 905 669 779 248 165 433 ÷ 2 = 1 888 952 834 889 624 082 716 + 1;
  • 1 888 952 834 889 624 082 716 ÷ 2 = 944 476 417 444 812 041 358 + 0;
  • 944 476 417 444 812 041 358 ÷ 2 = 472 238 208 722 406 020 679 + 0;
  • 472 238 208 722 406 020 679 ÷ 2 = 236 119 104 361 203 010 339 + 1;
  • 236 119 104 361 203 010 339 ÷ 2 = 118 059 552 180 601 505 169 + 1;
  • 118 059 552 180 601 505 169 ÷ 2 = 59 029 776 090 300 752 584 + 1;
  • 59 029 776 090 300 752 584 ÷ 2 = 29 514 888 045 150 376 292 + 0;
  • 29 514 888 045 150 376 292 ÷ 2 = 14 757 444 022 575 188 146 + 0;
  • 14 757 444 022 575 188 146 ÷ 2 = 7 378 722 011 287 594 073 + 0;
  • 7 378 722 011 287 594 073 ÷ 2 = 3 689 361 005 643 797 036 + 1;
  • 3 689 361 005 643 797 036 ÷ 2 = 1 844 680 502 821 898 518 + 0;
  • 1 844 680 502 821 898 518 ÷ 2 = 922 340 251 410 949 259 + 0;
  • 922 340 251 410 949 259 ÷ 2 = 461 170 125 705 474 629 + 1;
  • 461 170 125 705 474 629 ÷ 2 = 230 585 062 852 737 314 + 1;
  • 230 585 062 852 737 314 ÷ 2 = 115 292 531 426 368 657 + 0;
  • 115 292 531 426 368 657 ÷ 2 = 57 646 265 713 184 328 + 1;
  • 57 646 265 713 184 328 ÷ 2 = 28 823 132 856 592 164 + 0;
  • 28 823 132 856 592 164 ÷ 2 = 14 411 566 428 296 082 + 0;
  • 14 411 566 428 296 082 ÷ 2 = 7 205 783 214 148 041 + 0;
  • 7 205 783 214 148 041 ÷ 2 = 3 602 891 607 074 020 + 1;
  • 3 602 891 607 074 020 ÷ 2 = 1 801 445 803 537 010 + 0;
  • 1 801 445 803 537 010 ÷ 2 = 900 722 901 768 505 + 0;
  • 900 722 901 768 505 ÷ 2 = 450 361 450 884 252 + 1;
  • 450 361 450 884 252 ÷ 2 = 225 180 725 442 126 + 0;
  • 225 180 725 442 126 ÷ 2 = 112 590 362 721 063 + 0;
  • 112 590 362 721 063 ÷ 2 = 56 295 181 360 531 + 1;
  • 56 295 181 360 531 ÷ 2 = 28 147 590 680 265 + 1;
  • 28 147 590 680 265 ÷ 2 = 14 073 795 340 132 + 1;
  • 14 073 795 340 132 ÷ 2 = 7 036 897 670 066 + 0;
  • 7 036 897 670 066 ÷ 2 = 3 518 448 835 033 + 0;
  • 3 518 448 835 033 ÷ 2 = 1 759 224 417 516 + 1;
  • 1 759 224 417 516 ÷ 2 = 879 612 208 758 + 0;
  • 879 612 208 758 ÷ 2 = 439 806 104 379 + 0;
  • 439 806 104 379 ÷ 2 = 219 903 052 189 + 1;
  • 219 903 052 189 ÷ 2 = 109 951 526 094 + 1;
  • 109 951 526 094 ÷ 2 = 54 975 763 047 + 0;
  • 54 975 763 047 ÷ 2 = 27 487 881 523 + 1;
  • 27 487 881 523 ÷ 2 = 13 743 940 761 + 1;
  • 13 743 940 761 ÷ 2 = 6 871 970 380 + 1;
  • 6 871 970 380 ÷ 2 = 3 435 985 190 + 0;
  • 3 435 985 190 ÷ 2 = 1 717 992 595 + 0;
  • 1 717 992 595 ÷ 2 = 858 996 297 + 1;
  • 858 996 297 ÷ 2 = 429 498 148 + 1;
  • 429 498 148 ÷ 2 = 214 749 074 + 0;
  • 214 749 074 ÷ 2 = 107 374 537 + 0;
  • 107 374 537 ÷ 2 = 53 687 268 + 1;
  • 53 687 268 ÷ 2 = 26 843 634 + 0;
  • 26 843 634 ÷ 2 = 13 421 817 + 0;
  • 13 421 817 ÷ 2 = 6 710 908 + 1;
  • 6 710 908 ÷ 2 = 3 355 454 + 0;
  • 3 355 454 ÷ 2 = 1 677 727 + 0;
  • 1 677 727 ÷ 2 = 838 863 + 1;
  • 838 863 ÷ 2 = 419 431 + 1;
  • 419 431 ÷ 2 = 209 715 + 1;
  • 209 715 ÷ 2 = 104 857 + 1;
  • 104 857 ÷ 2 = 52 428 + 1;
  • 52 428 ÷ 2 = 26 214 + 0;
  • 26 214 ÷ 2 = 13 107 + 0;
  • 13 107 ÷ 2 = 6 553 + 1;
  • 6 553 ÷ 2 = 3 276 + 1;
  • 3 276 ÷ 2 = 1 638 + 0;
  • 1 638 ÷ 2 = 819 + 0;
  • 819 ÷ 2 = 409 + 1;
  • 409 ÷ 2 = 204 + 1;
  • 204 ÷ 2 = 102 + 0;
  • 102 ÷ 2 = 51 + 0;
  • 51 ÷ 2 = 25 + 1;
  • 25 ÷ 2 = 12 + 1;
  • 12 ÷ 2 = 6 + 0;
  • 6 ÷ 2 = 3 + 0;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

34 028 349 133 317 412 298 432 094 444 834 721 134(10) =


1 1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110(2)


3. Convert to binary (base 2) the fractional part: 0.64.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.64 × 2 = 1 + 0.28;
  • 2) 0.28 × 2 = 0 + 0.56;
  • 3) 0.56 × 2 = 1 + 0.12;
  • 4) 0.12 × 2 = 0 + 0.24;
  • 5) 0.24 × 2 = 0 + 0.48;
  • 6) 0.48 × 2 = 0 + 0.96;
  • 7) 0.96 × 2 = 1 + 0.92;
  • 8) 0.92 × 2 = 1 + 0.84;
  • 9) 0.84 × 2 = 1 + 0.68;
  • 10) 0.68 × 2 = 1 + 0.36;
  • 11) 0.36 × 2 = 0 + 0.72;
  • 12) 0.72 × 2 = 1 + 0.44;
  • 13) 0.44 × 2 = 0 + 0.88;
  • 14) 0.88 × 2 = 1 + 0.76;
  • 15) 0.76 × 2 = 1 + 0.52;
  • 16) 0.52 × 2 = 1 + 0.04;
  • 17) 0.04 × 2 = 0 + 0.08;
  • 18) 0.08 × 2 = 0 + 0.16;
  • 19) 0.16 × 2 = 0 + 0.32;
  • 20) 0.32 × 2 = 0 + 0.64;
  • 21) 0.64 × 2 = 1 + 0.28;
  • 22) 0.28 × 2 = 0 + 0.56;
  • 23) 0.56 × 2 = 1 + 0.12;
  • 24) 0.12 × 2 = 0 + 0.24;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.64(10) =


0.1010 0011 1101 0111 0000 1010(2)

5. Positive number before normalization:

34 028 349 133 317 412 298 432 094 444 834 721 134.64(10) =


1 1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110.1010 0011 1101 0111 0000 1010(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 124 positions to the left, so that only one non zero digit remains to the left of it:


34 028 349 133 317 412 298 432 094 444 834 721 134.64(10) =


1 1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110.1010 0011 1101 0111 0000 1010(2) =


1 1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110.1010 0011 1101 0111 0000 1010(2) × 20 =


1.1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110 1010 0011 1101 0111 0000 1010(2) × 2124


7. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): 124


Mantissa (not normalized):
1.1001 1001 1001 1001 1111 0010 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110 1010 0011 1101 0111 0000 1010


8. Adjust the exponent.

Use the 8 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(8-1) - 1 =


124 + 2(8-1) - 1 =


(124 + 127)(10) =


251(10)


9. Convert the adjusted exponent from the decimal (base 10) to 8 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 251 ÷ 2 = 125 + 1;
  • 125 ÷ 2 = 62 + 1;
  • 62 ÷ 2 = 31 + 0;
  • 31 ÷ 2 = 15 + 1;
  • 15 ÷ 2 = 7 + 1;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


251(10) =


1111 1011(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 23 bits, by removing the excess bits, from the right (if any of the excess bits is set on 1, we are losing precision...).


Mantissa (normalized) =


1. 100 1100 1100 1100 1111 1001 0 0100 1100 1110 1100 1001 1100 1001 0001 0110 0100 0111 0010 0111 0100 0010 1010 0100 1111 1110 1011 0011 0000 1101 0110 1110 1010 0011 1101 0111 0000 1010 =


100 1100 1100 1100 1111 1001


12. The three elements that make up the number's 32 bit single precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (8 bits) =
1111 1011


Mantissa (23 bits) =
100 1100 1100 1100 1111 1001


Decimal number 34 028 349 133 317 412 298 432 094 444 834 721 134.64 converted to 32 bit single precision IEEE 754 binary floating point representation:

0 - 1111 1011 - 100 1100 1100 1100 1111 1001


How to convert decimal numbers from base ten to 32 bit single precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 32 bit single precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the base ten positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, by shifting the decimal point (or if you prefer, the decimal mark) "n" positions either to the left or to the right, so that only one non zero digit remains to the left of the decimal point.
  • 7. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign if the case) and adjust its length to 23 bits, either by removing the excess bits from the right (losing precision...) or by adding extra '0' bits to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -25.347 from decimal system (base ten) to 32 bit single precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-25.347| = 25.347

  • 2. First convert the integer part, 25. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 25 ÷ 2 = 12 + 1;
    • 12 ÷ 2 = 6 + 0;
    • 6 ÷ 2 = 3 + 0;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    25(10) = 1 1001(2)

  • 4. Then convert the fractional part, 0.347. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.347 × 2 = 0 + 0.694;
    • 2) 0.694 × 2 = 1 + 0.388;
    • 3) 0.388 × 2 = 0 + 0.776;
    • 4) 0.776 × 2 = 1 + 0.552;
    • 5) 0.552 × 2 = 1 + 0.104;
    • 6) 0.104 × 2 = 0 + 0.208;
    • 7) 0.208 × 2 = 0 + 0.416;
    • 8) 0.416 × 2 = 0 + 0.832;
    • 9) 0.832 × 2 = 1 + 0.664;
    • 10) 0.664 × 2 = 1 + 0.328;
    • 11) 0.328 × 2 = 0 + 0.656;
    • 12) 0.656 × 2 = 1 + 0.312;
    • 13) 0.312 × 2 = 0 + 0.624;
    • 14) 0.624 × 2 = 1 + 0.248;
    • 15) 0.248 × 2 = 0 + 0.496;
    • 16) 0.496 × 2 = 0 + 0.992;
    • 17) 0.992 × 2 = 1 + 0.984;
    • 18) 0.984 × 2 = 1 + 0.968;
    • 19) 0.968 × 2 = 1 + 0.936;
    • 20) 0.936 × 2 = 1 + 0.872;
    • 21) 0.872 × 2 = 1 + 0.744;
    • 22) 0.744 × 2 = 1 + 0.488;
    • 23) 0.488 × 2 = 0 + 0.976;
    • 24) 0.976 × 2 = 1 + 0.952;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 23) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.347(10) = 0.0101 1000 1101 0100 1111 1101(2)

  • 6. Summarizing - the positive number before normalization:

    25.347(10) = 1 1001.0101 1000 1101 0100 1111 1101(2)

  • 7. Normalize the binary representation of the number, shifting the decimal point 4 positions to the left so that only one non-zero digit stays to the left of the decimal point:

    25.347(10) =
    1 1001.0101 1000 1101 0100 1111 1101(2) =
    1 1001.0101 1000 1101 0100 1111 1101(2) × 20 =
    1.1001 0101 1000 1101 0100 1111 1101(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 32 bit single precision IEEE 754 binary floating point:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

  • 9. Adjust the exponent in 8 bit excess/bias notation and then convert it from decimal (base 10) to 8 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as already demonstrated above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(8-1) - 1 = (4 + 127)(10) = 131(10) =
    1000 0011(2)

  • 10. Normalize the mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal point) and adjust its length to 23 bits, by removing the excess bits from the right (losing precision...):

    Mantissa (not-normalized): 1.1001 0101 1000 1101 0100 1111 1101

    Mantissa (normalized): 100 1010 1100 0110 1010 0111

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 1000 0011

    Mantissa (23 bits) = 100 1010 1100 0110 1010 0111

  • Number -25.347, converted from the decimal system (base 10) to 32 bit single precision IEEE 754 binary floating point =
    1 - 1000 0011 - 100 1010 1100 0110 1010 0111