0.000 000 000 000 000 000 000 000 341 8 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.000 000 000 000 000 000 000 000 341 8(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.000 000 000 000 000 000 000 000 341 8(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.000 000 000 000 000 000 000 000 341 8.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 000 000 000 000 000 000 000 341 8 × 2 = 0 + 0.000 000 000 000 000 000 000 000 683 6;
  • 2) 0.000 000 000 000 000 000 000 000 683 6 × 2 = 0 + 0.000 000 000 000 000 000 000 001 367 2;
  • 3) 0.000 000 000 000 000 000 000 001 367 2 × 2 = 0 + 0.000 000 000 000 000 000 000 002 734 4;
  • 4) 0.000 000 000 000 000 000 000 002 734 4 × 2 = 0 + 0.000 000 000 000 000 000 000 005 468 8;
  • 5) 0.000 000 000 000 000 000 000 005 468 8 × 2 = 0 + 0.000 000 000 000 000 000 000 010 937 6;
  • 6) 0.000 000 000 000 000 000 000 010 937 6 × 2 = 0 + 0.000 000 000 000 000 000 000 021 875 2;
  • 7) 0.000 000 000 000 000 000 000 021 875 2 × 2 = 0 + 0.000 000 000 000 000 000 000 043 750 4;
  • 8) 0.000 000 000 000 000 000 000 043 750 4 × 2 = 0 + 0.000 000 000 000 000 000 000 087 500 8;
  • 9) 0.000 000 000 000 000 000 000 087 500 8 × 2 = 0 + 0.000 000 000 000 000 000 000 175 001 6;
  • 10) 0.000 000 000 000 000 000 000 175 001 6 × 2 = 0 + 0.000 000 000 000 000 000 000 350 003 2;
  • 11) 0.000 000 000 000 000 000 000 350 003 2 × 2 = 0 + 0.000 000 000 000 000 000 000 700 006 4;
  • 12) 0.000 000 000 000 000 000 000 700 006 4 × 2 = 0 + 0.000 000 000 000 000 000 001 400 012 8;
  • 13) 0.000 000 000 000 000 000 001 400 012 8 × 2 = 0 + 0.000 000 000 000 000 000 002 800 025 6;
  • 14) 0.000 000 000 000 000 000 002 800 025 6 × 2 = 0 + 0.000 000 000 000 000 000 005 600 051 2;
  • 15) 0.000 000 000 000 000 000 005 600 051 2 × 2 = 0 + 0.000 000 000 000 000 000 011 200 102 4;
  • 16) 0.000 000 000 000 000 000 011 200 102 4 × 2 = 0 + 0.000 000 000 000 000 000 022 400 204 8;
  • 17) 0.000 000 000 000 000 000 022 400 204 8 × 2 = 0 + 0.000 000 000 000 000 000 044 800 409 6;
  • 18) 0.000 000 000 000 000 000 044 800 409 6 × 2 = 0 + 0.000 000 000 000 000 000 089 600 819 2;
  • 19) 0.000 000 000 000 000 000 089 600 819 2 × 2 = 0 + 0.000 000 000 000 000 000 179 201 638 4;
  • 20) 0.000 000 000 000 000 000 179 201 638 4 × 2 = 0 + 0.000 000 000 000 000 000 358 403 276 8;
  • 21) 0.000 000 000 000 000 000 358 403 276 8 × 2 = 0 + 0.000 000 000 000 000 000 716 806 553 6;
  • 22) 0.000 000 000 000 000 000 716 806 553 6 × 2 = 0 + 0.000 000 000 000 000 001 433 613 107 2;
  • 23) 0.000 000 000 000 000 001 433 613 107 2 × 2 = 0 + 0.000 000 000 000 000 002 867 226 214 4;
  • 24) 0.000 000 000 000 000 002 867 226 214 4 × 2 = 0 + 0.000 000 000 000 000 005 734 452 428 8;
  • 25) 0.000 000 000 000 000 005 734 452 428 8 × 2 = 0 + 0.000 000 000 000 000 011 468 904 857 6;
  • 26) 0.000 000 000 000 000 011 468 904 857 6 × 2 = 0 + 0.000 000 000 000 000 022 937 809 715 2;
  • 27) 0.000 000 000 000 000 022 937 809 715 2 × 2 = 0 + 0.000 000 000 000 000 045 875 619 430 4;
  • 28) 0.000 000 000 000 000 045 875 619 430 4 × 2 = 0 + 0.000 000 000 000 000 091 751 238 860 8;
  • 29) 0.000 000 000 000 000 091 751 238 860 8 × 2 = 0 + 0.000 000 000 000 000 183 502 477 721 6;
  • 30) 0.000 000 000 000 000 183 502 477 721 6 × 2 = 0 + 0.000 000 000 000 000 367 004 955 443 2;
  • 31) 0.000 000 000 000 000 367 004 955 443 2 × 2 = 0 + 0.000 000 000 000 000 734 009 910 886 4;
  • 32) 0.000 000 000 000 000 734 009 910 886 4 × 2 = 0 + 0.000 000 000 000 001 468 019 821 772 8;
  • 33) 0.000 000 000 000 001 468 019 821 772 8 × 2 = 0 + 0.000 000 000 000 002 936 039 643 545 6;
  • 34) 0.000 000 000 000 002 936 039 643 545 6 × 2 = 0 + 0.000 000 000 000 005 872 079 287 091 2;
  • 35) 0.000 000 000 000 005 872 079 287 091 2 × 2 = 0 + 0.000 000 000 000 011 744 158 574 182 4;
  • 36) 0.000 000 000 000 011 744 158 574 182 4 × 2 = 0 + 0.000 000 000 000 023 488 317 148 364 8;
  • 37) 0.000 000 000 000 023 488 317 148 364 8 × 2 = 0 + 0.000 000 000 000 046 976 634 296 729 6;
  • 38) 0.000 000 000 000 046 976 634 296 729 6 × 2 = 0 + 0.000 000 000 000 093 953 268 593 459 2;
  • 39) 0.000 000 000 000 093 953 268 593 459 2 × 2 = 0 + 0.000 000 000 000 187 906 537 186 918 4;
  • 40) 0.000 000 000 000 187 906 537 186 918 4 × 2 = 0 + 0.000 000 000 000 375 813 074 373 836 8;
  • 41) 0.000 000 000 000 375 813 074 373 836 8 × 2 = 0 + 0.000 000 000 000 751 626 148 747 673 6;
  • 42) 0.000 000 000 000 751 626 148 747 673 6 × 2 = 0 + 0.000 000 000 001 503 252 297 495 347 2;
  • 43) 0.000 000 000 001 503 252 297 495 347 2 × 2 = 0 + 0.000 000 000 003 006 504 594 990 694 4;
  • 44) 0.000 000 000 003 006 504 594 990 694 4 × 2 = 0 + 0.000 000 000 006 013 009 189 981 388 8;
  • 45) 0.000 000 000 006 013 009 189 981 388 8 × 2 = 0 + 0.000 000 000 012 026 018 379 962 777 6;
  • 46) 0.000 000 000 012 026 018 379 962 777 6 × 2 = 0 + 0.000 000 000 024 052 036 759 925 555 2;
  • 47) 0.000 000 000 024 052 036 759 925 555 2 × 2 = 0 + 0.000 000 000 048 104 073 519 851 110 4;
  • 48) 0.000 000 000 048 104 073 519 851 110 4 × 2 = 0 + 0.000 000 000 096 208 147 039 702 220 8;
  • 49) 0.000 000 000 096 208 147 039 702 220 8 × 2 = 0 + 0.000 000 000 192 416 294 079 404 441 6;
  • 50) 0.000 000 000 192 416 294 079 404 441 6 × 2 = 0 + 0.000 000 000 384 832 588 158 808 883 2;
  • 51) 0.000 000 000 384 832 588 158 808 883 2 × 2 = 0 + 0.000 000 000 769 665 176 317 617 766 4;
  • 52) 0.000 000 000 769 665 176 317 617 766 4 × 2 = 0 + 0.000 000 001 539 330 352 635 235 532 8;
  • 53) 0.000 000 001 539 330 352 635 235 532 8 × 2 = 0 + 0.000 000 003 078 660 705 270 471 065 6;
  • 54) 0.000 000 003 078 660 705 270 471 065 6 × 2 = 0 + 0.000 000 006 157 321 410 540 942 131 2;
  • 55) 0.000 000 006 157 321 410 540 942 131 2 × 2 = 0 + 0.000 000 012 314 642 821 081 884 262 4;
  • 56) 0.000 000 012 314 642 821 081 884 262 4 × 2 = 0 + 0.000 000 024 629 285 642 163 768 524 8;
  • 57) 0.000 000 024 629 285 642 163 768 524 8 × 2 = 0 + 0.000 000 049 258 571 284 327 537 049 6;
  • 58) 0.000 000 049 258 571 284 327 537 049 6 × 2 = 0 + 0.000 000 098 517 142 568 655 074 099 2;
  • 59) 0.000 000 098 517 142 568 655 074 099 2 × 2 = 0 + 0.000 000 197 034 285 137 310 148 198 4;
  • 60) 0.000 000 197 034 285 137 310 148 198 4 × 2 = 0 + 0.000 000 394 068 570 274 620 296 396 8;
  • 61) 0.000 000 394 068 570 274 620 296 396 8 × 2 = 0 + 0.000 000 788 137 140 549 240 592 793 6;
  • 62) 0.000 000 788 137 140 549 240 592 793 6 × 2 = 0 + 0.000 001 576 274 281 098 481 185 587 2;
  • 63) 0.000 001 576 274 281 098 481 185 587 2 × 2 = 0 + 0.000 003 152 548 562 196 962 371 174 4;
  • 64) 0.000 003 152 548 562 196 962 371 174 4 × 2 = 0 + 0.000 006 305 097 124 393 924 742 348 8;
  • 65) 0.000 006 305 097 124 393 924 742 348 8 × 2 = 0 + 0.000 012 610 194 248 787 849 484 697 6;
  • 66) 0.000 012 610 194 248 787 849 484 697 6 × 2 = 0 + 0.000 025 220 388 497 575 698 969 395 2;
  • 67) 0.000 025 220 388 497 575 698 969 395 2 × 2 = 0 + 0.000 050 440 776 995 151 397 938 790 4;
  • 68) 0.000 050 440 776 995 151 397 938 790 4 × 2 = 0 + 0.000 100 881 553 990 302 795 877 580 8;
  • 69) 0.000 100 881 553 990 302 795 877 580 8 × 2 = 0 + 0.000 201 763 107 980 605 591 755 161 6;
  • 70) 0.000 201 763 107 980 605 591 755 161 6 × 2 = 0 + 0.000 403 526 215 961 211 183 510 323 2;
  • 71) 0.000 403 526 215 961 211 183 510 323 2 × 2 = 0 + 0.000 807 052 431 922 422 367 020 646 4;
  • 72) 0.000 807 052 431 922 422 367 020 646 4 × 2 = 0 + 0.001 614 104 863 844 844 734 041 292 8;
  • 73) 0.001 614 104 863 844 844 734 041 292 8 × 2 = 0 + 0.003 228 209 727 689 689 468 082 585 6;
  • 74) 0.003 228 209 727 689 689 468 082 585 6 × 2 = 0 + 0.006 456 419 455 379 378 936 165 171 2;
  • 75) 0.006 456 419 455 379 378 936 165 171 2 × 2 = 0 + 0.012 912 838 910 758 757 872 330 342 4;
  • 76) 0.012 912 838 910 758 757 872 330 342 4 × 2 = 0 + 0.025 825 677 821 517 515 744 660 684 8;
  • 77) 0.025 825 677 821 517 515 744 660 684 8 × 2 = 0 + 0.051 651 355 643 035 031 489 321 369 6;
  • 78) 0.051 651 355 643 035 031 489 321 369 6 × 2 = 0 + 0.103 302 711 286 070 062 978 642 739 2;
  • 79) 0.103 302 711 286 070 062 978 642 739 2 × 2 = 0 + 0.206 605 422 572 140 125 957 285 478 4;
  • 80) 0.206 605 422 572 140 125 957 285 478 4 × 2 = 0 + 0.413 210 845 144 280 251 914 570 956 8;
  • 81) 0.413 210 845 144 280 251 914 570 956 8 × 2 = 0 + 0.826 421 690 288 560 503 829 141 913 6;
  • 82) 0.826 421 690 288 560 503 829 141 913 6 × 2 = 1 + 0.652 843 380 577 121 007 658 283 827 2;
  • 83) 0.652 843 380 577 121 007 658 283 827 2 × 2 = 1 + 0.305 686 761 154 242 015 316 567 654 4;
  • 84) 0.305 686 761 154 242 015 316 567 654 4 × 2 = 0 + 0.611 373 522 308 484 030 633 135 308 8;
  • 85) 0.611 373 522 308 484 030 633 135 308 8 × 2 = 1 + 0.222 747 044 616 968 061 266 270 617 6;
  • 86) 0.222 747 044 616 968 061 266 270 617 6 × 2 = 0 + 0.445 494 089 233 936 122 532 541 235 2;
  • 87) 0.445 494 089 233 936 122 532 541 235 2 × 2 = 0 + 0.890 988 178 467 872 245 065 082 470 4;
  • 88) 0.890 988 178 467 872 245 065 082 470 4 × 2 = 1 + 0.781 976 356 935 744 490 130 164 940 8;
  • 89) 0.781 976 356 935 744 490 130 164 940 8 × 2 = 1 + 0.563 952 713 871 488 980 260 329 881 6;
  • 90) 0.563 952 713 871 488 980 260 329 881 6 × 2 = 1 + 0.127 905 427 742 977 960 520 659 763 2;
  • 91) 0.127 905 427 742 977 960 520 659 763 2 × 2 = 0 + 0.255 810 855 485 955 921 041 319 526 4;
  • 92) 0.255 810 855 485 955 921 041 319 526 4 × 2 = 0 + 0.511 621 710 971 911 842 082 639 052 8;
  • 93) 0.511 621 710 971 911 842 082 639 052 8 × 2 = 1 + 0.023 243 421 943 823 684 165 278 105 6;
  • 94) 0.023 243 421 943 823 684 165 278 105 6 × 2 = 0 + 0.046 486 843 887 647 368 330 556 211 2;
  • 95) 0.046 486 843 887 647 368 330 556 211 2 × 2 = 0 + 0.092 973 687 775 294 736 661 112 422 4;
  • 96) 0.092 973 687 775 294 736 661 112 422 4 × 2 = 0 + 0.185 947 375 550 589 473 322 224 844 8;
  • 97) 0.185 947 375 550 589 473 322 224 844 8 × 2 = 0 + 0.371 894 751 101 178 946 644 449 689 6;
  • 98) 0.371 894 751 101 178 946 644 449 689 6 × 2 = 0 + 0.743 789 502 202 357 893 288 899 379 2;
  • 99) 0.743 789 502 202 357 893 288 899 379 2 × 2 = 1 + 0.487 579 004 404 715 786 577 798 758 4;
  • 100) 0.487 579 004 404 715 786 577 798 758 4 × 2 = 0 + 0.975 158 008 809 431 573 155 597 516 8;
  • 101) 0.975 158 008 809 431 573 155 597 516 8 × 2 = 1 + 0.950 316 017 618 863 146 311 195 033 6;
  • 102) 0.950 316 017 618 863 146 311 195 033 6 × 2 = 1 + 0.900 632 035 237 726 292 622 390 067 2;
  • 103) 0.900 632 035 237 726 292 622 390 067 2 × 2 = 1 + 0.801 264 070 475 452 585 244 780 134 4;
  • 104) 0.801 264 070 475 452 585 244 780 134 4 × 2 = 1 + 0.602 528 140 950 905 170 489 560 268 8;
  • 105) 0.602 528 140 950 905 170 489 560 268 8 × 2 = 1 + 0.205 056 281 901 810 340 979 120 537 6;
  • 106) 0.205 056 281 901 810 340 979 120 537 6 × 2 = 0 + 0.410 112 563 803 620 681 958 241 075 2;
  • 107) 0.410 112 563 803 620 681 958 241 075 2 × 2 = 0 + 0.820 225 127 607 241 363 916 482 150 4;
  • 108) 0.820 225 127 607 241 363 916 482 150 4 × 2 = 1 + 0.640 450 255 214 482 727 832 964 300 8;
  • 109) 0.640 450 255 214 482 727 832 964 300 8 × 2 = 1 + 0.280 900 510 428 965 455 665 928 601 6;
  • 110) 0.280 900 510 428 965 455 665 928 601 6 × 2 = 0 + 0.561 801 020 857 930 911 331 857 203 2;
  • 111) 0.561 801 020 857 930 911 331 857 203 2 × 2 = 1 + 0.123 602 041 715 861 822 663 714 406 4;
  • 112) 0.123 602 041 715 861 822 663 714 406 4 × 2 = 0 + 0.247 204 083 431 723 645 327 428 812 8;
  • 113) 0.247 204 083 431 723 645 327 428 812 8 × 2 = 0 + 0.494 408 166 863 447 290 654 857 625 6;
  • 114) 0.494 408 166 863 447 290 654 857 625 6 × 2 = 0 + 0.988 816 333 726 894 581 309 715 251 2;
  • 115) 0.988 816 333 726 894 581 309 715 251 2 × 2 = 1 + 0.977 632 667 453 789 162 619 430 502 4;
  • 116) 0.977 632 667 453 789 162 619 430 502 4 × 2 = 1 + 0.955 265 334 907 578 325 238 861 004 8;
  • 117) 0.955 265 334 907 578 325 238 861 004 8 × 2 = 1 + 0.910 530 669 815 156 650 477 722 009 6;
  • 118) 0.910 530 669 815 156 650 477 722 009 6 × 2 = 1 + 0.821 061 339 630 313 300 955 444 019 2;
  • 119) 0.821 061 339 630 313 300 955 444 019 2 × 2 = 1 + 0.642 122 679 260 626 601 910 888 038 4;
  • 120) 0.642 122 679 260 626 601 910 888 038 4 × 2 = 1 + 0.284 245 358 521 253 203 821 776 076 8;
  • 121) 0.284 245 358 521 253 203 821 776 076 8 × 2 = 0 + 0.568 490 717 042 506 407 643 552 153 6;
  • 122) 0.568 490 717 042 506 407 643 552 153 6 × 2 = 1 + 0.136 981 434 085 012 815 287 104 307 2;
  • 123) 0.136 981 434 085 012 815 287 104 307 2 × 2 = 0 + 0.273 962 868 170 025 630 574 208 614 4;
  • 124) 0.273 962 868 170 025 630 574 208 614 4 × 2 = 0 + 0.547 925 736 340 051 261 148 417 228 8;
  • 125) 0.547 925 736 340 051 261 148 417 228 8 × 2 = 1 + 0.095 851 472 680 102 522 296 834 457 6;
  • 126) 0.095 851 472 680 102 522 296 834 457 6 × 2 = 0 + 0.191 702 945 360 205 044 593 668 915 2;
  • 127) 0.191 702 945 360 205 044 593 668 915 2 × 2 = 0 + 0.383 405 890 720 410 089 187 337 830 4;
  • 128) 0.383 405 890 720 410 089 187 337 830 4 × 2 = 0 + 0.766 811 781 440 820 178 374 675 660 8;
  • 129) 0.766 811 781 440 820 178 374 675 660 8 × 2 = 1 + 0.533 623 562 881 640 356 749 351 321 6;
  • 130) 0.533 623 562 881 640 356 749 351 321 6 × 2 = 1 + 0.067 247 125 763 280 713 498 702 643 2;
  • 131) 0.067 247 125 763 280 713 498 702 643 2 × 2 = 0 + 0.134 494 251 526 561 426 997 405 286 4;
  • 132) 0.134 494 251 526 561 426 997 405 286 4 × 2 = 0 + 0.268 988 503 053 122 853 994 810 572 8;
  • 133) 0.268 988 503 053 122 853 994 810 572 8 × 2 = 0 + 0.537 977 006 106 245 707 989 621 145 6;
  • 134) 0.537 977 006 106 245 707 989 621 145 6 × 2 = 1 + 0.075 954 012 212 491 415 979 242 291 2;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 000 000 000 000 000 000 000 341 8(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 1001 1100 1000 0010 1111 1001 1010 0011 1111 0100 1000 1100 01(2)

5. Positive number before normalization:

0.000 000 000 000 000 000 000 000 341 8(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 1001 1100 1000 0010 1111 1001 1010 0011 1111 0100 1000 1100 01(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 82 positions to the right, so that only one non zero digit remains to the left of it:


0.000 000 000 000 000 000 000 000 341 8(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 1001 1100 1000 0010 1111 1001 1010 0011 1111 0100 1000 1100 01(2) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 1001 1100 1000 0010 1111 1001 1010 0011 1111 0100 1000 1100 01(2) × 20 =


1.1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001(2) × 2-82


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -82


Mantissa (not normalized):
1.1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-82 + 2(11-1) - 1 =


(-82 + 1 023)(10) =


941(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 941 ÷ 2 = 470 + 1;
  • 470 ÷ 2 = 235 + 0;
  • 235 ÷ 2 = 117 + 1;
  • 117 ÷ 2 = 58 + 1;
  • 58 ÷ 2 = 29 + 0;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


941(10) =


011 1010 1101(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001 =


1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1010 1101


Mantissa (52 bits) =
1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001


Decimal number 0.000 000 000 000 000 000 000 000 341 8 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1010 1101 - 1010 0111 0010 0000 1011 1110 0110 1000 1111 1101 0010 0011 0001


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100