0.000 000 000 000 000 000 054 210 108 624 274 99 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.000 000 000 000 000 000 054 210 108 624 274 99(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.000 000 000 000 000 000 054 210 108 624 274 99(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.000 000 000 000 000 000 054 210 108 624 274 99.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 000 000 000 000 000 054 210 108 624 274 99 × 2 = 0 + 0.000 000 000 000 000 000 108 420 217 248 549 98;
  • 2) 0.000 000 000 000 000 000 108 420 217 248 549 98 × 2 = 0 + 0.000 000 000 000 000 000 216 840 434 497 099 96;
  • 3) 0.000 000 000 000 000 000 216 840 434 497 099 96 × 2 = 0 + 0.000 000 000 000 000 000 433 680 868 994 199 92;
  • 4) 0.000 000 000 000 000 000 433 680 868 994 199 92 × 2 = 0 + 0.000 000 000 000 000 000 867 361 737 988 399 84;
  • 5) 0.000 000 000 000 000 000 867 361 737 988 399 84 × 2 = 0 + 0.000 000 000 000 000 001 734 723 475 976 799 68;
  • 6) 0.000 000 000 000 000 001 734 723 475 976 799 68 × 2 = 0 + 0.000 000 000 000 000 003 469 446 951 953 599 36;
  • 7) 0.000 000 000 000 000 003 469 446 951 953 599 36 × 2 = 0 + 0.000 000 000 000 000 006 938 893 903 907 198 72;
  • 8) 0.000 000 000 000 000 006 938 893 903 907 198 72 × 2 = 0 + 0.000 000 000 000 000 013 877 787 807 814 397 44;
  • 9) 0.000 000 000 000 000 013 877 787 807 814 397 44 × 2 = 0 + 0.000 000 000 000 000 027 755 575 615 628 794 88;
  • 10) 0.000 000 000 000 000 027 755 575 615 628 794 88 × 2 = 0 + 0.000 000 000 000 000 055 511 151 231 257 589 76;
  • 11) 0.000 000 000 000 000 055 511 151 231 257 589 76 × 2 = 0 + 0.000 000 000 000 000 111 022 302 462 515 179 52;
  • 12) 0.000 000 000 000 000 111 022 302 462 515 179 52 × 2 = 0 + 0.000 000 000 000 000 222 044 604 925 030 359 04;
  • 13) 0.000 000 000 000 000 222 044 604 925 030 359 04 × 2 = 0 + 0.000 000 000 000 000 444 089 209 850 060 718 08;
  • 14) 0.000 000 000 000 000 444 089 209 850 060 718 08 × 2 = 0 + 0.000 000 000 000 000 888 178 419 700 121 436 16;
  • 15) 0.000 000 000 000 000 888 178 419 700 121 436 16 × 2 = 0 + 0.000 000 000 000 001 776 356 839 400 242 872 32;
  • 16) 0.000 000 000 000 001 776 356 839 400 242 872 32 × 2 = 0 + 0.000 000 000 000 003 552 713 678 800 485 744 64;
  • 17) 0.000 000 000 000 003 552 713 678 800 485 744 64 × 2 = 0 + 0.000 000 000 000 007 105 427 357 600 971 489 28;
  • 18) 0.000 000 000 000 007 105 427 357 600 971 489 28 × 2 = 0 + 0.000 000 000 000 014 210 854 715 201 942 978 56;
  • 19) 0.000 000 000 000 014 210 854 715 201 942 978 56 × 2 = 0 + 0.000 000 000 000 028 421 709 430 403 885 957 12;
  • 20) 0.000 000 000 000 028 421 709 430 403 885 957 12 × 2 = 0 + 0.000 000 000 000 056 843 418 860 807 771 914 24;
  • 21) 0.000 000 000 000 056 843 418 860 807 771 914 24 × 2 = 0 + 0.000 000 000 000 113 686 837 721 615 543 828 48;
  • 22) 0.000 000 000 000 113 686 837 721 615 543 828 48 × 2 = 0 + 0.000 000 000 000 227 373 675 443 231 087 656 96;
  • 23) 0.000 000 000 000 227 373 675 443 231 087 656 96 × 2 = 0 + 0.000 000 000 000 454 747 350 886 462 175 313 92;
  • 24) 0.000 000 000 000 454 747 350 886 462 175 313 92 × 2 = 0 + 0.000 000 000 000 909 494 701 772 924 350 627 84;
  • 25) 0.000 000 000 000 909 494 701 772 924 350 627 84 × 2 = 0 + 0.000 000 000 001 818 989 403 545 848 701 255 68;
  • 26) 0.000 000 000 001 818 989 403 545 848 701 255 68 × 2 = 0 + 0.000 000 000 003 637 978 807 091 697 402 511 36;
  • 27) 0.000 000 000 003 637 978 807 091 697 402 511 36 × 2 = 0 + 0.000 000 000 007 275 957 614 183 394 805 022 72;
  • 28) 0.000 000 000 007 275 957 614 183 394 805 022 72 × 2 = 0 + 0.000 000 000 014 551 915 228 366 789 610 045 44;
  • 29) 0.000 000 000 014 551 915 228 366 789 610 045 44 × 2 = 0 + 0.000 000 000 029 103 830 456 733 579 220 090 88;
  • 30) 0.000 000 000 029 103 830 456 733 579 220 090 88 × 2 = 0 + 0.000 000 000 058 207 660 913 467 158 440 181 76;
  • 31) 0.000 000 000 058 207 660 913 467 158 440 181 76 × 2 = 0 + 0.000 000 000 116 415 321 826 934 316 880 363 52;
  • 32) 0.000 000 000 116 415 321 826 934 316 880 363 52 × 2 = 0 + 0.000 000 000 232 830 643 653 868 633 760 727 04;
  • 33) 0.000 000 000 232 830 643 653 868 633 760 727 04 × 2 = 0 + 0.000 000 000 465 661 287 307 737 267 521 454 08;
  • 34) 0.000 000 000 465 661 287 307 737 267 521 454 08 × 2 = 0 + 0.000 000 000 931 322 574 615 474 535 042 908 16;
  • 35) 0.000 000 000 931 322 574 615 474 535 042 908 16 × 2 = 0 + 0.000 000 001 862 645 149 230 949 070 085 816 32;
  • 36) 0.000 000 001 862 645 149 230 949 070 085 816 32 × 2 = 0 + 0.000 000 003 725 290 298 461 898 140 171 632 64;
  • 37) 0.000 000 003 725 290 298 461 898 140 171 632 64 × 2 = 0 + 0.000 000 007 450 580 596 923 796 280 343 265 28;
  • 38) 0.000 000 007 450 580 596 923 796 280 343 265 28 × 2 = 0 + 0.000 000 014 901 161 193 847 592 560 686 530 56;
  • 39) 0.000 000 014 901 161 193 847 592 560 686 530 56 × 2 = 0 + 0.000 000 029 802 322 387 695 185 121 373 061 12;
  • 40) 0.000 000 029 802 322 387 695 185 121 373 061 12 × 2 = 0 + 0.000 000 059 604 644 775 390 370 242 746 122 24;
  • 41) 0.000 000 059 604 644 775 390 370 242 746 122 24 × 2 = 0 + 0.000 000 119 209 289 550 780 740 485 492 244 48;
  • 42) 0.000 000 119 209 289 550 780 740 485 492 244 48 × 2 = 0 + 0.000 000 238 418 579 101 561 480 970 984 488 96;
  • 43) 0.000 000 238 418 579 101 561 480 970 984 488 96 × 2 = 0 + 0.000 000 476 837 158 203 122 961 941 968 977 92;
  • 44) 0.000 000 476 837 158 203 122 961 941 968 977 92 × 2 = 0 + 0.000 000 953 674 316 406 245 923 883 937 955 84;
  • 45) 0.000 000 953 674 316 406 245 923 883 937 955 84 × 2 = 0 + 0.000 001 907 348 632 812 491 847 767 875 911 68;
  • 46) 0.000 001 907 348 632 812 491 847 767 875 911 68 × 2 = 0 + 0.000 003 814 697 265 624 983 695 535 751 823 36;
  • 47) 0.000 003 814 697 265 624 983 695 535 751 823 36 × 2 = 0 + 0.000 007 629 394 531 249 967 391 071 503 646 72;
  • 48) 0.000 007 629 394 531 249 967 391 071 503 646 72 × 2 = 0 + 0.000 015 258 789 062 499 934 782 143 007 293 44;
  • 49) 0.000 015 258 789 062 499 934 782 143 007 293 44 × 2 = 0 + 0.000 030 517 578 124 999 869 564 286 014 586 88;
  • 50) 0.000 030 517 578 124 999 869 564 286 014 586 88 × 2 = 0 + 0.000 061 035 156 249 999 739 128 572 029 173 76;
  • 51) 0.000 061 035 156 249 999 739 128 572 029 173 76 × 2 = 0 + 0.000 122 070 312 499 999 478 257 144 058 347 52;
  • 52) 0.000 122 070 312 499 999 478 257 144 058 347 52 × 2 = 0 + 0.000 244 140 624 999 998 956 514 288 116 695 04;
  • 53) 0.000 244 140 624 999 998 956 514 288 116 695 04 × 2 = 0 + 0.000 488 281 249 999 997 913 028 576 233 390 08;
  • 54) 0.000 488 281 249 999 997 913 028 576 233 390 08 × 2 = 0 + 0.000 976 562 499 999 995 826 057 152 466 780 16;
  • 55) 0.000 976 562 499 999 995 826 057 152 466 780 16 × 2 = 0 + 0.001 953 124 999 999 991 652 114 304 933 560 32;
  • 56) 0.001 953 124 999 999 991 652 114 304 933 560 32 × 2 = 0 + 0.003 906 249 999 999 983 304 228 609 867 120 64;
  • 57) 0.003 906 249 999 999 983 304 228 609 867 120 64 × 2 = 0 + 0.007 812 499 999 999 966 608 457 219 734 241 28;
  • 58) 0.007 812 499 999 999 966 608 457 219 734 241 28 × 2 = 0 + 0.015 624 999 999 999 933 216 914 439 468 482 56;
  • 59) 0.015 624 999 999 999 933 216 914 439 468 482 56 × 2 = 0 + 0.031 249 999 999 999 866 433 828 878 936 965 12;
  • 60) 0.031 249 999 999 999 866 433 828 878 936 965 12 × 2 = 0 + 0.062 499 999 999 999 732 867 657 757 873 930 24;
  • 61) 0.062 499 999 999 999 732 867 657 757 873 930 24 × 2 = 0 + 0.124 999 999 999 999 465 735 315 515 747 860 48;
  • 62) 0.124 999 999 999 999 465 735 315 515 747 860 48 × 2 = 0 + 0.249 999 999 999 998 931 470 631 031 495 720 96;
  • 63) 0.249 999 999 999 998 931 470 631 031 495 720 96 × 2 = 0 + 0.499 999 999 999 997 862 941 262 062 991 441 92;
  • 64) 0.499 999 999 999 997 862 941 262 062 991 441 92 × 2 = 0 + 0.999 999 999 999 995 725 882 524 125 982 883 84;
  • 65) 0.999 999 999 999 995 725 882 524 125 982 883 84 × 2 = 1 + 0.999 999 999 999 991 451 765 048 251 965 767 68;
  • 66) 0.999 999 999 999 991 451 765 048 251 965 767 68 × 2 = 1 + 0.999 999 999 999 982 903 530 096 503 931 535 36;
  • 67) 0.999 999 999 999 982 903 530 096 503 931 535 36 × 2 = 1 + 0.999 999 999 999 965 807 060 193 007 863 070 72;
  • 68) 0.999 999 999 999 965 807 060 193 007 863 070 72 × 2 = 1 + 0.999 999 999 999 931 614 120 386 015 726 141 44;
  • 69) 0.999 999 999 999 931 614 120 386 015 726 141 44 × 2 = 1 + 0.999 999 999 999 863 228 240 772 031 452 282 88;
  • 70) 0.999 999 999 999 863 228 240 772 031 452 282 88 × 2 = 1 + 0.999 999 999 999 726 456 481 544 062 904 565 76;
  • 71) 0.999 999 999 999 726 456 481 544 062 904 565 76 × 2 = 1 + 0.999 999 999 999 452 912 963 088 125 809 131 52;
  • 72) 0.999 999 999 999 452 912 963 088 125 809 131 52 × 2 = 1 + 0.999 999 999 998 905 825 926 176 251 618 263 04;
  • 73) 0.999 999 999 998 905 825 926 176 251 618 263 04 × 2 = 1 + 0.999 999 999 997 811 651 852 352 503 236 526 08;
  • 74) 0.999 999 999 997 811 651 852 352 503 236 526 08 × 2 = 1 + 0.999 999 999 995 623 303 704 705 006 473 052 16;
  • 75) 0.999 999 999 995 623 303 704 705 006 473 052 16 × 2 = 1 + 0.999 999 999 991 246 607 409 410 012 946 104 32;
  • 76) 0.999 999 999 991 246 607 409 410 012 946 104 32 × 2 = 1 + 0.999 999 999 982 493 214 818 820 025 892 208 64;
  • 77) 0.999 999 999 982 493 214 818 820 025 892 208 64 × 2 = 1 + 0.999 999 999 964 986 429 637 640 051 784 417 28;
  • 78) 0.999 999 999 964 986 429 637 640 051 784 417 28 × 2 = 1 + 0.999 999 999 929 972 859 275 280 103 568 834 56;
  • 79) 0.999 999 999 929 972 859 275 280 103 568 834 56 × 2 = 1 + 0.999 999 999 859 945 718 550 560 207 137 669 12;
  • 80) 0.999 999 999 859 945 718 550 560 207 137 669 12 × 2 = 1 + 0.999 999 999 719 891 437 101 120 414 275 338 24;
  • 81) 0.999 999 999 719 891 437 101 120 414 275 338 24 × 2 = 1 + 0.999 999 999 439 782 874 202 240 828 550 676 48;
  • 82) 0.999 999 999 439 782 874 202 240 828 550 676 48 × 2 = 1 + 0.999 999 998 879 565 748 404 481 657 101 352 96;
  • 83) 0.999 999 998 879 565 748 404 481 657 101 352 96 × 2 = 1 + 0.999 999 997 759 131 496 808 963 314 202 705 92;
  • 84) 0.999 999 997 759 131 496 808 963 314 202 705 92 × 2 = 1 + 0.999 999 995 518 262 993 617 926 628 405 411 84;
  • 85) 0.999 999 995 518 262 993 617 926 628 405 411 84 × 2 = 1 + 0.999 999 991 036 525 987 235 853 256 810 823 68;
  • 86) 0.999 999 991 036 525 987 235 853 256 810 823 68 × 2 = 1 + 0.999 999 982 073 051 974 471 706 513 621 647 36;
  • 87) 0.999 999 982 073 051 974 471 706 513 621 647 36 × 2 = 1 + 0.999 999 964 146 103 948 943 413 027 243 294 72;
  • 88) 0.999 999 964 146 103 948 943 413 027 243 294 72 × 2 = 1 + 0.999 999 928 292 207 897 886 826 054 486 589 44;
  • 89) 0.999 999 928 292 207 897 886 826 054 486 589 44 × 2 = 1 + 0.999 999 856 584 415 795 773 652 108 973 178 88;
  • 90) 0.999 999 856 584 415 795 773 652 108 973 178 88 × 2 = 1 + 0.999 999 713 168 831 591 547 304 217 946 357 76;
  • 91) 0.999 999 713 168 831 591 547 304 217 946 357 76 × 2 = 1 + 0.999 999 426 337 663 183 094 608 435 892 715 52;
  • 92) 0.999 999 426 337 663 183 094 608 435 892 715 52 × 2 = 1 + 0.999 998 852 675 326 366 189 216 871 785 431 04;
  • 93) 0.999 998 852 675 326 366 189 216 871 785 431 04 × 2 = 1 + 0.999 997 705 350 652 732 378 433 743 570 862 08;
  • 94) 0.999 997 705 350 652 732 378 433 743 570 862 08 × 2 = 1 + 0.999 995 410 701 305 464 756 867 487 141 724 16;
  • 95) 0.999 995 410 701 305 464 756 867 487 141 724 16 × 2 = 1 + 0.999 990 821 402 610 929 513 734 974 283 448 32;
  • 96) 0.999 990 821 402 610 929 513 734 974 283 448 32 × 2 = 1 + 0.999 981 642 805 221 859 027 469 948 566 896 64;
  • 97) 0.999 981 642 805 221 859 027 469 948 566 896 64 × 2 = 1 + 0.999 963 285 610 443 718 054 939 897 133 793 28;
  • 98) 0.999 963 285 610 443 718 054 939 897 133 793 28 × 2 = 1 + 0.999 926 571 220 887 436 109 879 794 267 586 56;
  • 99) 0.999 926 571 220 887 436 109 879 794 267 586 56 × 2 = 1 + 0.999 853 142 441 774 872 219 759 588 535 173 12;
  • 100) 0.999 853 142 441 774 872 219 759 588 535 173 12 × 2 = 1 + 0.999 706 284 883 549 744 439 519 177 070 346 24;
  • 101) 0.999 706 284 883 549 744 439 519 177 070 346 24 × 2 = 1 + 0.999 412 569 767 099 488 879 038 354 140 692 48;
  • 102) 0.999 412 569 767 099 488 879 038 354 140 692 48 × 2 = 1 + 0.998 825 139 534 198 977 758 076 708 281 384 96;
  • 103) 0.998 825 139 534 198 977 758 076 708 281 384 96 × 2 = 1 + 0.997 650 279 068 397 955 516 153 416 562 769 92;
  • 104) 0.997 650 279 068 397 955 516 153 416 562 769 92 × 2 = 1 + 0.995 300 558 136 795 911 032 306 833 125 539 84;
  • 105) 0.995 300 558 136 795 911 032 306 833 125 539 84 × 2 = 1 + 0.990 601 116 273 591 822 064 613 666 251 079 68;
  • 106) 0.990 601 116 273 591 822 064 613 666 251 079 68 × 2 = 1 + 0.981 202 232 547 183 644 129 227 332 502 159 36;
  • 107) 0.981 202 232 547 183 644 129 227 332 502 159 36 × 2 = 1 + 0.962 404 465 094 367 288 258 454 665 004 318 72;
  • 108) 0.962 404 465 094 367 288 258 454 665 004 318 72 × 2 = 1 + 0.924 808 930 188 734 576 516 909 330 008 637 44;
  • 109) 0.924 808 930 188 734 576 516 909 330 008 637 44 × 2 = 1 + 0.849 617 860 377 469 153 033 818 660 017 274 88;
  • 110) 0.849 617 860 377 469 153 033 818 660 017 274 88 × 2 = 1 + 0.699 235 720 754 938 306 067 637 320 034 549 76;
  • 111) 0.699 235 720 754 938 306 067 637 320 034 549 76 × 2 = 1 + 0.398 471 441 509 876 612 135 274 640 069 099 52;
  • 112) 0.398 471 441 509 876 612 135 274 640 069 099 52 × 2 = 0 + 0.796 942 883 019 753 224 270 549 280 138 199 04;
  • 113) 0.796 942 883 019 753 224 270 549 280 138 199 04 × 2 = 1 + 0.593 885 766 039 506 448 541 098 560 276 398 08;
  • 114) 0.593 885 766 039 506 448 541 098 560 276 398 08 × 2 = 1 + 0.187 771 532 079 012 897 082 197 120 552 796 16;
  • 115) 0.187 771 532 079 012 897 082 197 120 552 796 16 × 2 = 0 + 0.375 543 064 158 025 794 164 394 241 105 592 32;
  • 116) 0.375 543 064 158 025 794 164 394 241 105 592 32 × 2 = 0 + 0.751 086 128 316 051 588 328 788 482 211 184 64;
  • 117) 0.751 086 128 316 051 588 328 788 482 211 184 64 × 2 = 1 + 0.502 172 256 632 103 176 657 576 964 422 369 28;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 000 000 000 000 000 054 210 108 624 274 99(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1110 1100 1(2)

5. Positive number before normalization:

0.000 000 000 000 000 000 054 210 108 624 274 99(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1110 1100 1(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 65 positions to the right, so that only one non zero digit remains to the left of it:


0.000 000 000 000 000 000 054 210 108 624 274 99(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1110 1100 1(2) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1110 1100 1(2) × 20 =


1.1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001(2) × 2-65


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -65


Mantissa (not normalized):
1.1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-65 + 2(11-1) - 1 =


(-65 + 1 023)(10) =


958(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 958 ÷ 2 = 479 + 0;
  • 479 ÷ 2 = 239 + 1;
  • 239 ÷ 2 = 119 + 1;
  • 119 ÷ 2 = 59 + 1;
  • 59 ÷ 2 = 29 + 1;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


958(10) =


011 1011 1110(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001 =


1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1011 1110


Mantissa (52 bits) =
1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001


Decimal number 0.000 000 000 000 000 000 054 210 108 624 274 99 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1011 1110 - 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1111 1101 1001


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100