0.000 000 000 000 000 000 000 138 064 889 7 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.000 000 000 000 000 000 000 138 064 889 7(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.000 000 000 000 000 000 000 138 064 889 7(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.000 000 000 000 000 000 000 138 064 889 7.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 000 000 000 000 000 000 138 064 889 7 × 2 = 0 + 0.000 000 000 000 000 000 000 276 129 779 4;
  • 2) 0.000 000 000 000 000 000 000 276 129 779 4 × 2 = 0 + 0.000 000 000 000 000 000 000 552 259 558 8;
  • 3) 0.000 000 000 000 000 000 000 552 259 558 8 × 2 = 0 + 0.000 000 000 000 000 000 001 104 519 117 6;
  • 4) 0.000 000 000 000 000 000 001 104 519 117 6 × 2 = 0 + 0.000 000 000 000 000 000 002 209 038 235 2;
  • 5) 0.000 000 000 000 000 000 002 209 038 235 2 × 2 = 0 + 0.000 000 000 000 000 000 004 418 076 470 4;
  • 6) 0.000 000 000 000 000 000 004 418 076 470 4 × 2 = 0 + 0.000 000 000 000 000 000 008 836 152 940 8;
  • 7) 0.000 000 000 000 000 000 008 836 152 940 8 × 2 = 0 + 0.000 000 000 000 000 000 017 672 305 881 6;
  • 8) 0.000 000 000 000 000 000 017 672 305 881 6 × 2 = 0 + 0.000 000 000 000 000 000 035 344 611 763 2;
  • 9) 0.000 000 000 000 000 000 035 344 611 763 2 × 2 = 0 + 0.000 000 000 000 000 000 070 689 223 526 4;
  • 10) 0.000 000 000 000 000 000 070 689 223 526 4 × 2 = 0 + 0.000 000 000 000 000 000 141 378 447 052 8;
  • 11) 0.000 000 000 000 000 000 141 378 447 052 8 × 2 = 0 + 0.000 000 000 000 000 000 282 756 894 105 6;
  • 12) 0.000 000 000 000 000 000 282 756 894 105 6 × 2 = 0 + 0.000 000 000 000 000 000 565 513 788 211 2;
  • 13) 0.000 000 000 000 000 000 565 513 788 211 2 × 2 = 0 + 0.000 000 000 000 000 001 131 027 576 422 4;
  • 14) 0.000 000 000 000 000 001 131 027 576 422 4 × 2 = 0 + 0.000 000 000 000 000 002 262 055 152 844 8;
  • 15) 0.000 000 000 000 000 002 262 055 152 844 8 × 2 = 0 + 0.000 000 000 000 000 004 524 110 305 689 6;
  • 16) 0.000 000 000 000 000 004 524 110 305 689 6 × 2 = 0 + 0.000 000 000 000 000 009 048 220 611 379 2;
  • 17) 0.000 000 000 000 000 009 048 220 611 379 2 × 2 = 0 + 0.000 000 000 000 000 018 096 441 222 758 4;
  • 18) 0.000 000 000 000 000 018 096 441 222 758 4 × 2 = 0 + 0.000 000 000 000 000 036 192 882 445 516 8;
  • 19) 0.000 000 000 000 000 036 192 882 445 516 8 × 2 = 0 + 0.000 000 000 000 000 072 385 764 891 033 6;
  • 20) 0.000 000 000 000 000 072 385 764 891 033 6 × 2 = 0 + 0.000 000 000 000 000 144 771 529 782 067 2;
  • 21) 0.000 000 000 000 000 144 771 529 782 067 2 × 2 = 0 + 0.000 000 000 000 000 289 543 059 564 134 4;
  • 22) 0.000 000 000 000 000 289 543 059 564 134 4 × 2 = 0 + 0.000 000 000 000 000 579 086 119 128 268 8;
  • 23) 0.000 000 000 000 000 579 086 119 128 268 8 × 2 = 0 + 0.000 000 000 000 001 158 172 238 256 537 6;
  • 24) 0.000 000 000 000 001 158 172 238 256 537 6 × 2 = 0 + 0.000 000 000 000 002 316 344 476 513 075 2;
  • 25) 0.000 000 000 000 002 316 344 476 513 075 2 × 2 = 0 + 0.000 000 000 000 004 632 688 953 026 150 4;
  • 26) 0.000 000 000 000 004 632 688 953 026 150 4 × 2 = 0 + 0.000 000 000 000 009 265 377 906 052 300 8;
  • 27) 0.000 000 000 000 009 265 377 906 052 300 8 × 2 = 0 + 0.000 000 000 000 018 530 755 812 104 601 6;
  • 28) 0.000 000 000 000 018 530 755 812 104 601 6 × 2 = 0 + 0.000 000 000 000 037 061 511 624 209 203 2;
  • 29) 0.000 000 000 000 037 061 511 624 209 203 2 × 2 = 0 + 0.000 000 000 000 074 123 023 248 418 406 4;
  • 30) 0.000 000 000 000 074 123 023 248 418 406 4 × 2 = 0 + 0.000 000 000 000 148 246 046 496 836 812 8;
  • 31) 0.000 000 000 000 148 246 046 496 836 812 8 × 2 = 0 + 0.000 000 000 000 296 492 092 993 673 625 6;
  • 32) 0.000 000 000 000 296 492 092 993 673 625 6 × 2 = 0 + 0.000 000 000 000 592 984 185 987 347 251 2;
  • 33) 0.000 000 000 000 592 984 185 987 347 251 2 × 2 = 0 + 0.000 000 000 001 185 968 371 974 694 502 4;
  • 34) 0.000 000 000 001 185 968 371 974 694 502 4 × 2 = 0 + 0.000 000 000 002 371 936 743 949 389 004 8;
  • 35) 0.000 000 000 002 371 936 743 949 389 004 8 × 2 = 0 + 0.000 000 000 004 743 873 487 898 778 009 6;
  • 36) 0.000 000 000 004 743 873 487 898 778 009 6 × 2 = 0 + 0.000 000 000 009 487 746 975 797 556 019 2;
  • 37) 0.000 000 000 009 487 746 975 797 556 019 2 × 2 = 0 + 0.000 000 000 018 975 493 951 595 112 038 4;
  • 38) 0.000 000 000 018 975 493 951 595 112 038 4 × 2 = 0 + 0.000 000 000 037 950 987 903 190 224 076 8;
  • 39) 0.000 000 000 037 950 987 903 190 224 076 8 × 2 = 0 + 0.000 000 000 075 901 975 806 380 448 153 6;
  • 40) 0.000 000 000 075 901 975 806 380 448 153 6 × 2 = 0 + 0.000 000 000 151 803 951 612 760 896 307 2;
  • 41) 0.000 000 000 151 803 951 612 760 896 307 2 × 2 = 0 + 0.000 000 000 303 607 903 225 521 792 614 4;
  • 42) 0.000 000 000 303 607 903 225 521 792 614 4 × 2 = 0 + 0.000 000 000 607 215 806 451 043 585 228 8;
  • 43) 0.000 000 000 607 215 806 451 043 585 228 8 × 2 = 0 + 0.000 000 001 214 431 612 902 087 170 457 6;
  • 44) 0.000 000 001 214 431 612 902 087 170 457 6 × 2 = 0 + 0.000 000 002 428 863 225 804 174 340 915 2;
  • 45) 0.000 000 002 428 863 225 804 174 340 915 2 × 2 = 0 + 0.000 000 004 857 726 451 608 348 681 830 4;
  • 46) 0.000 000 004 857 726 451 608 348 681 830 4 × 2 = 0 + 0.000 000 009 715 452 903 216 697 363 660 8;
  • 47) 0.000 000 009 715 452 903 216 697 363 660 8 × 2 = 0 + 0.000 000 019 430 905 806 433 394 727 321 6;
  • 48) 0.000 000 019 430 905 806 433 394 727 321 6 × 2 = 0 + 0.000 000 038 861 811 612 866 789 454 643 2;
  • 49) 0.000 000 038 861 811 612 866 789 454 643 2 × 2 = 0 + 0.000 000 077 723 623 225 733 578 909 286 4;
  • 50) 0.000 000 077 723 623 225 733 578 909 286 4 × 2 = 0 + 0.000 000 155 447 246 451 467 157 818 572 8;
  • 51) 0.000 000 155 447 246 451 467 157 818 572 8 × 2 = 0 + 0.000 000 310 894 492 902 934 315 637 145 6;
  • 52) 0.000 000 310 894 492 902 934 315 637 145 6 × 2 = 0 + 0.000 000 621 788 985 805 868 631 274 291 2;
  • 53) 0.000 000 621 788 985 805 868 631 274 291 2 × 2 = 0 + 0.000 001 243 577 971 611 737 262 548 582 4;
  • 54) 0.000 001 243 577 971 611 737 262 548 582 4 × 2 = 0 + 0.000 002 487 155 943 223 474 525 097 164 8;
  • 55) 0.000 002 487 155 943 223 474 525 097 164 8 × 2 = 0 + 0.000 004 974 311 886 446 949 050 194 329 6;
  • 56) 0.000 004 974 311 886 446 949 050 194 329 6 × 2 = 0 + 0.000 009 948 623 772 893 898 100 388 659 2;
  • 57) 0.000 009 948 623 772 893 898 100 388 659 2 × 2 = 0 + 0.000 019 897 247 545 787 796 200 777 318 4;
  • 58) 0.000 019 897 247 545 787 796 200 777 318 4 × 2 = 0 + 0.000 039 794 495 091 575 592 401 554 636 8;
  • 59) 0.000 039 794 495 091 575 592 401 554 636 8 × 2 = 0 + 0.000 079 588 990 183 151 184 803 109 273 6;
  • 60) 0.000 079 588 990 183 151 184 803 109 273 6 × 2 = 0 + 0.000 159 177 980 366 302 369 606 218 547 2;
  • 61) 0.000 159 177 980 366 302 369 606 218 547 2 × 2 = 0 + 0.000 318 355 960 732 604 739 212 437 094 4;
  • 62) 0.000 318 355 960 732 604 739 212 437 094 4 × 2 = 0 + 0.000 636 711 921 465 209 478 424 874 188 8;
  • 63) 0.000 636 711 921 465 209 478 424 874 188 8 × 2 = 0 + 0.001 273 423 842 930 418 956 849 748 377 6;
  • 64) 0.001 273 423 842 930 418 956 849 748 377 6 × 2 = 0 + 0.002 546 847 685 860 837 913 699 496 755 2;
  • 65) 0.002 546 847 685 860 837 913 699 496 755 2 × 2 = 0 + 0.005 093 695 371 721 675 827 398 993 510 4;
  • 66) 0.005 093 695 371 721 675 827 398 993 510 4 × 2 = 0 + 0.010 187 390 743 443 351 654 797 987 020 8;
  • 67) 0.010 187 390 743 443 351 654 797 987 020 8 × 2 = 0 + 0.020 374 781 486 886 703 309 595 974 041 6;
  • 68) 0.020 374 781 486 886 703 309 595 974 041 6 × 2 = 0 + 0.040 749 562 973 773 406 619 191 948 083 2;
  • 69) 0.040 749 562 973 773 406 619 191 948 083 2 × 2 = 0 + 0.081 499 125 947 546 813 238 383 896 166 4;
  • 70) 0.081 499 125 947 546 813 238 383 896 166 4 × 2 = 0 + 0.162 998 251 895 093 626 476 767 792 332 8;
  • 71) 0.162 998 251 895 093 626 476 767 792 332 8 × 2 = 0 + 0.325 996 503 790 187 252 953 535 584 665 6;
  • 72) 0.325 996 503 790 187 252 953 535 584 665 6 × 2 = 0 + 0.651 993 007 580 374 505 907 071 169 331 2;
  • 73) 0.651 993 007 580 374 505 907 071 169 331 2 × 2 = 1 + 0.303 986 015 160 749 011 814 142 338 662 4;
  • 74) 0.303 986 015 160 749 011 814 142 338 662 4 × 2 = 0 + 0.607 972 030 321 498 023 628 284 677 324 8;
  • 75) 0.607 972 030 321 498 023 628 284 677 324 8 × 2 = 1 + 0.215 944 060 642 996 047 256 569 354 649 6;
  • 76) 0.215 944 060 642 996 047 256 569 354 649 6 × 2 = 0 + 0.431 888 121 285 992 094 513 138 709 299 2;
  • 77) 0.431 888 121 285 992 094 513 138 709 299 2 × 2 = 0 + 0.863 776 242 571 984 189 026 277 418 598 4;
  • 78) 0.863 776 242 571 984 189 026 277 418 598 4 × 2 = 1 + 0.727 552 485 143 968 378 052 554 837 196 8;
  • 79) 0.727 552 485 143 968 378 052 554 837 196 8 × 2 = 1 + 0.455 104 970 287 936 756 105 109 674 393 6;
  • 80) 0.455 104 970 287 936 756 105 109 674 393 6 × 2 = 0 + 0.910 209 940 575 873 512 210 219 348 787 2;
  • 81) 0.910 209 940 575 873 512 210 219 348 787 2 × 2 = 1 + 0.820 419 881 151 747 024 420 438 697 574 4;
  • 82) 0.820 419 881 151 747 024 420 438 697 574 4 × 2 = 1 + 0.640 839 762 303 494 048 840 877 395 148 8;
  • 83) 0.640 839 762 303 494 048 840 877 395 148 8 × 2 = 1 + 0.281 679 524 606 988 097 681 754 790 297 6;
  • 84) 0.281 679 524 606 988 097 681 754 790 297 6 × 2 = 0 + 0.563 359 049 213 976 195 363 509 580 595 2;
  • 85) 0.563 359 049 213 976 195 363 509 580 595 2 × 2 = 1 + 0.126 718 098 427 952 390 727 019 161 190 4;
  • 86) 0.126 718 098 427 952 390 727 019 161 190 4 × 2 = 0 + 0.253 436 196 855 904 781 454 038 322 380 8;
  • 87) 0.253 436 196 855 904 781 454 038 322 380 8 × 2 = 0 + 0.506 872 393 711 809 562 908 076 644 761 6;
  • 88) 0.506 872 393 711 809 562 908 076 644 761 6 × 2 = 1 + 0.013 744 787 423 619 125 816 153 289 523 2;
  • 89) 0.013 744 787 423 619 125 816 153 289 523 2 × 2 = 0 + 0.027 489 574 847 238 251 632 306 579 046 4;
  • 90) 0.027 489 574 847 238 251 632 306 579 046 4 × 2 = 0 + 0.054 979 149 694 476 503 264 613 158 092 8;
  • 91) 0.054 979 149 694 476 503 264 613 158 092 8 × 2 = 0 + 0.109 958 299 388 953 006 529 226 316 185 6;
  • 92) 0.109 958 299 388 953 006 529 226 316 185 6 × 2 = 0 + 0.219 916 598 777 906 013 058 452 632 371 2;
  • 93) 0.219 916 598 777 906 013 058 452 632 371 2 × 2 = 0 + 0.439 833 197 555 812 026 116 905 264 742 4;
  • 94) 0.439 833 197 555 812 026 116 905 264 742 4 × 2 = 0 + 0.879 666 395 111 624 052 233 810 529 484 8;
  • 95) 0.879 666 395 111 624 052 233 810 529 484 8 × 2 = 1 + 0.759 332 790 223 248 104 467 621 058 969 6;
  • 96) 0.759 332 790 223 248 104 467 621 058 969 6 × 2 = 1 + 0.518 665 580 446 496 208 935 242 117 939 2;
  • 97) 0.518 665 580 446 496 208 935 242 117 939 2 × 2 = 1 + 0.037 331 160 892 992 417 870 484 235 878 4;
  • 98) 0.037 331 160 892 992 417 870 484 235 878 4 × 2 = 0 + 0.074 662 321 785 984 835 740 968 471 756 8;
  • 99) 0.074 662 321 785 984 835 740 968 471 756 8 × 2 = 0 + 0.149 324 643 571 969 671 481 936 943 513 6;
  • 100) 0.149 324 643 571 969 671 481 936 943 513 6 × 2 = 0 + 0.298 649 287 143 939 342 963 873 887 027 2;
  • 101) 0.298 649 287 143 939 342 963 873 887 027 2 × 2 = 0 + 0.597 298 574 287 878 685 927 747 774 054 4;
  • 102) 0.597 298 574 287 878 685 927 747 774 054 4 × 2 = 1 + 0.194 597 148 575 757 371 855 495 548 108 8;
  • 103) 0.194 597 148 575 757 371 855 495 548 108 8 × 2 = 0 + 0.389 194 297 151 514 743 710 991 096 217 6;
  • 104) 0.389 194 297 151 514 743 710 991 096 217 6 × 2 = 0 + 0.778 388 594 303 029 487 421 982 192 435 2;
  • 105) 0.778 388 594 303 029 487 421 982 192 435 2 × 2 = 1 + 0.556 777 188 606 058 974 843 964 384 870 4;
  • 106) 0.556 777 188 606 058 974 843 964 384 870 4 × 2 = 1 + 0.113 554 377 212 117 949 687 928 769 740 8;
  • 107) 0.113 554 377 212 117 949 687 928 769 740 8 × 2 = 0 + 0.227 108 754 424 235 899 375 857 539 481 6;
  • 108) 0.227 108 754 424 235 899 375 857 539 481 6 × 2 = 0 + 0.454 217 508 848 471 798 751 715 078 963 2;
  • 109) 0.454 217 508 848 471 798 751 715 078 963 2 × 2 = 0 + 0.908 435 017 696 943 597 503 430 157 926 4;
  • 110) 0.908 435 017 696 943 597 503 430 157 926 4 × 2 = 1 + 0.816 870 035 393 887 195 006 860 315 852 8;
  • 111) 0.816 870 035 393 887 195 006 860 315 852 8 × 2 = 1 + 0.633 740 070 787 774 390 013 720 631 705 6;
  • 112) 0.633 740 070 787 774 390 013 720 631 705 6 × 2 = 1 + 0.267 480 141 575 548 780 027 441 263 411 2;
  • 113) 0.267 480 141 575 548 780 027 441 263 411 2 × 2 = 0 + 0.534 960 283 151 097 560 054 882 526 822 4;
  • 114) 0.534 960 283 151 097 560 054 882 526 822 4 × 2 = 1 + 0.069 920 566 302 195 120 109 765 053 644 8;
  • 115) 0.069 920 566 302 195 120 109 765 053 644 8 × 2 = 0 + 0.139 841 132 604 390 240 219 530 107 289 6;
  • 116) 0.139 841 132 604 390 240 219 530 107 289 6 × 2 = 0 + 0.279 682 265 208 780 480 439 060 214 579 2;
  • 117) 0.279 682 265 208 780 480 439 060 214 579 2 × 2 = 0 + 0.559 364 530 417 560 960 878 120 429 158 4;
  • 118) 0.559 364 530 417 560 960 878 120 429 158 4 × 2 = 1 + 0.118 729 060 835 121 921 756 240 858 316 8;
  • 119) 0.118 729 060 835 121 921 756 240 858 316 8 × 2 = 0 + 0.237 458 121 670 243 843 512 481 716 633 6;
  • 120) 0.237 458 121 670 243 843 512 481 716 633 6 × 2 = 0 + 0.474 916 243 340 487 687 024 963 433 267 2;
  • 121) 0.474 916 243 340 487 687 024 963 433 267 2 × 2 = 0 + 0.949 832 486 680 975 374 049 926 866 534 4;
  • 122) 0.949 832 486 680 975 374 049 926 866 534 4 × 2 = 1 + 0.899 664 973 361 950 748 099 853 733 068 8;
  • 123) 0.899 664 973 361 950 748 099 853 733 068 8 × 2 = 1 + 0.799 329 946 723 901 496 199 707 466 137 6;
  • 124) 0.799 329 946 723 901 496 199 707 466 137 6 × 2 = 1 + 0.598 659 893 447 802 992 399 414 932 275 2;
  • 125) 0.598 659 893 447 802 992 399 414 932 275 2 × 2 = 1 + 0.197 319 786 895 605 984 798 829 864 550 4;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 000 000 000 000 000 000 138 064 889 7(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1010 0110 1110 1001 0000 0011 1000 0100 1100 0111 0100 0100 0111 1(2)

5. Positive number before normalization:

0.000 000 000 000 000 000 000 138 064 889 7(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1010 0110 1110 1001 0000 0011 1000 0100 1100 0111 0100 0100 0111 1(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 73 positions to the right, so that only one non zero digit remains to the left of it:


0.000 000 000 000 000 000 000 138 064 889 7(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1010 0110 1110 1001 0000 0011 1000 0100 1100 0111 0100 0100 0111 1(2) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1010 0110 1110 1001 0000 0011 1000 0100 1100 0111 0100 0100 0111 1(2) × 20 =


1.0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111(2) × 2-73


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -73


Mantissa (not normalized):
1.0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-73 + 2(11-1) - 1 =


(-73 + 1 023)(10) =


950(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 950 ÷ 2 = 475 + 0;
  • 475 ÷ 2 = 237 + 1;
  • 237 ÷ 2 = 118 + 1;
  • 118 ÷ 2 = 59 + 0;
  • 59 ÷ 2 = 29 + 1;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


950(10) =


011 1011 0110(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111 =


0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1011 0110


Mantissa (52 bits) =
0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111


Decimal number 0.000 000 000 000 000 000 000 138 064 889 7 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1011 0110 - 0100 1101 1101 0010 0000 0111 0000 1001 1000 1110 1000 1000 1111


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100