0.000 000 000 000 000 000 054 210 108 624 276 4 Converted to 64 Bit Double Precision IEEE 754 Binary Floating Point Representation Standard

Convert decimal 0.000 000 000 000 000 000 054 210 108 624 276 4(10) to 64 bit double precision IEEE 754 binary floating point representation standard (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

What are the steps to convert decimal number
0.000 000 000 000 000 000 054 210 108 624 276 4(10) to 64 bit double precision IEEE 754 binary floating point representation (1 bit for sign, 11 bits for exponent, 52 bits for mantissa)

1. First, convert to binary (in base 2) the integer part: 0.
Divide the number repeatedly by 2.

Keep track of each remainder.

We stop when we get a quotient that is equal to zero.


  • division = quotient + remainder;
  • 0 ÷ 2 = 0 + 0;

2. Construct the base 2 representation of the integer part of the number.

Take all the remainders starting from the bottom of the list constructed above.

0(10) =


0(2)


3. Convert to binary (base 2) the fractional part: 0.000 000 000 000 000 000 054 210 108 624 276 4.

Multiply it repeatedly by 2.


Keep track of each integer part of the results.


Stop when we get a fractional part that is equal to zero.


  • #) multiplying = integer + fractional part;
  • 1) 0.000 000 000 000 000 000 054 210 108 624 276 4 × 2 = 0 + 0.000 000 000 000 000 000 108 420 217 248 552 8;
  • 2) 0.000 000 000 000 000 000 108 420 217 248 552 8 × 2 = 0 + 0.000 000 000 000 000 000 216 840 434 497 105 6;
  • 3) 0.000 000 000 000 000 000 216 840 434 497 105 6 × 2 = 0 + 0.000 000 000 000 000 000 433 680 868 994 211 2;
  • 4) 0.000 000 000 000 000 000 433 680 868 994 211 2 × 2 = 0 + 0.000 000 000 000 000 000 867 361 737 988 422 4;
  • 5) 0.000 000 000 000 000 000 867 361 737 988 422 4 × 2 = 0 + 0.000 000 000 000 000 001 734 723 475 976 844 8;
  • 6) 0.000 000 000 000 000 001 734 723 475 976 844 8 × 2 = 0 + 0.000 000 000 000 000 003 469 446 951 953 689 6;
  • 7) 0.000 000 000 000 000 003 469 446 951 953 689 6 × 2 = 0 + 0.000 000 000 000 000 006 938 893 903 907 379 2;
  • 8) 0.000 000 000 000 000 006 938 893 903 907 379 2 × 2 = 0 + 0.000 000 000 000 000 013 877 787 807 814 758 4;
  • 9) 0.000 000 000 000 000 013 877 787 807 814 758 4 × 2 = 0 + 0.000 000 000 000 000 027 755 575 615 629 516 8;
  • 10) 0.000 000 000 000 000 027 755 575 615 629 516 8 × 2 = 0 + 0.000 000 000 000 000 055 511 151 231 259 033 6;
  • 11) 0.000 000 000 000 000 055 511 151 231 259 033 6 × 2 = 0 + 0.000 000 000 000 000 111 022 302 462 518 067 2;
  • 12) 0.000 000 000 000 000 111 022 302 462 518 067 2 × 2 = 0 + 0.000 000 000 000 000 222 044 604 925 036 134 4;
  • 13) 0.000 000 000 000 000 222 044 604 925 036 134 4 × 2 = 0 + 0.000 000 000 000 000 444 089 209 850 072 268 8;
  • 14) 0.000 000 000 000 000 444 089 209 850 072 268 8 × 2 = 0 + 0.000 000 000 000 000 888 178 419 700 144 537 6;
  • 15) 0.000 000 000 000 000 888 178 419 700 144 537 6 × 2 = 0 + 0.000 000 000 000 001 776 356 839 400 289 075 2;
  • 16) 0.000 000 000 000 001 776 356 839 400 289 075 2 × 2 = 0 + 0.000 000 000 000 003 552 713 678 800 578 150 4;
  • 17) 0.000 000 000 000 003 552 713 678 800 578 150 4 × 2 = 0 + 0.000 000 000 000 007 105 427 357 601 156 300 8;
  • 18) 0.000 000 000 000 007 105 427 357 601 156 300 8 × 2 = 0 + 0.000 000 000 000 014 210 854 715 202 312 601 6;
  • 19) 0.000 000 000 000 014 210 854 715 202 312 601 6 × 2 = 0 + 0.000 000 000 000 028 421 709 430 404 625 203 2;
  • 20) 0.000 000 000 000 028 421 709 430 404 625 203 2 × 2 = 0 + 0.000 000 000 000 056 843 418 860 809 250 406 4;
  • 21) 0.000 000 000 000 056 843 418 860 809 250 406 4 × 2 = 0 + 0.000 000 000 000 113 686 837 721 618 500 812 8;
  • 22) 0.000 000 000 000 113 686 837 721 618 500 812 8 × 2 = 0 + 0.000 000 000 000 227 373 675 443 237 001 625 6;
  • 23) 0.000 000 000 000 227 373 675 443 237 001 625 6 × 2 = 0 + 0.000 000 000 000 454 747 350 886 474 003 251 2;
  • 24) 0.000 000 000 000 454 747 350 886 474 003 251 2 × 2 = 0 + 0.000 000 000 000 909 494 701 772 948 006 502 4;
  • 25) 0.000 000 000 000 909 494 701 772 948 006 502 4 × 2 = 0 + 0.000 000 000 001 818 989 403 545 896 013 004 8;
  • 26) 0.000 000 000 001 818 989 403 545 896 013 004 8 × 2 = 0 + 0.000 000 000 003 637 978 807 091 792 026 009 6;
  • 27) 0.000 000 000 003 637 978 807 091 792 026 009 6 × 2 = 0 + 0.000 000 000 007 275 957 614 183 584 052 019 2;
  • 28) 0.000 000 000 007 275 957 614 183 584 052 019 2 × 2 = 0 + 0.000 000 000 014 551 915 228 367 168 104 038 4;
  • 29) 0.000 000 000 014 551 915 228 367 168 104 038 4 × 2 = 0 + 0.000 000 000 029 103 830 456 734 336 208 076 8;
  • 30) 0.000 000 000 029 103 830 456 734 336 208 076 8 × 2 = 0 + 0.000 000 000 058 207 660 913 468 672 416 153 6;
  • 31) 0.000 000 000 058 207 660 913 468 672 416 153 6 × 2 = 0 + 0.000 000 000 116 415 321 826 937 344 832 307 2;
  • 32) 0.000 000 000 116 415 321 826 937 344 832 307 2 × 2 = 0 + 0.000 000 000 232 830 643 653 874 689 664 614 4;
  • 33) 0.000 000 000 232 830 643 653 874 689 664 614 4 × 2 = 0 + 0.000 000 000 465 661 287 307 749 379 329 228 8;
  • 34) 0.000 000 000 465 661 287 307 749 379 329 228 8 × 2 = 0 + 0.000 000 000 931 322 574 615 498 758 658 457 6;
  • 35) 0.000 000 000 931 322 574 615 498 758 658 457 6 × 2 = 0 + 0.000 000 001 862 645 149 230 997 517 316 915 2;
  • 36) 0.000 000 001 862 645 149 230 997 517 316 915 2 × 2 = 0 + 0.000 000 003 725 290 298 461 995 034 633 830 4;
  • 37) 0.000 000 003 725 290 298 461 995 034 633 830 4 × 2 = 0 + 0.000 000 007 450 580 596 923 990 069 267 660 8;
  • 38) 0.000 000 007 450 580 596 923 990 069 267 660 8 × 2 = 0 + 0.000 000 014 901 161 193 847 980 138 535 321 6;
  • 39) 0.000 000 014 901 161 193 847 980 138 535 321 6 × 2 = 0 + 0.000 000 029 802 322 387 695 960 277 070 643 2;
  • 40) 0.000 000 029 802 322 387 695 960 277 070 643 2 × 2 = 0 + 0.000 000 059 604 644 775 391 920 554 141 286 4;
  • 41) 0.000 000 059 604 644 775 391 920 554 141 286 4 × 2 = 0 + 0.000 000 119 209 289 550 783 841 108 282 572 8;
  • 42) 0.000 000 119 209 289 550 783 841 108 282 572 8 × 2 = 0 + 0.000 000 238 418 579 101 567 682 216 565 145 6;
  • 43) 0.000 000 238 418 579 101 567 682 216 565 145 6 × 2 = 0 + 0.000 000 476 837 158 203 135 364 433 130 291 2;
  • 44) 0.000 000 476 837 158 203 135 364 433 130 291 2 × 2 = 0 + 0.000 000 953 674 316 406 270 728 866 260 582 4;
  • 45) 0.000 000 953 674 316 406 270 728 866 260 582 4 × 2 = 0 + 0.000 001 907 348 632 812 541 457 732 521 164 8;
  • 46) 0.000 001 907 348 632 812 541 457 732 521 164 8 × 2 = 0 + 0.000 003 814 697 265 625 082 915 465 042 329 6;
  • 47) 0.000 003 814 697 265 625 082 915 465 042 329 6 × 2 = 0 + 0.000 007 629 394 531 250 165 830 930 084 659 2;
  • 48) 0.000 007 629 394 531 250 165 830 930 084 659 2 × 2 = 0 + 0.000 015 258 789 062 500 331 661 860 169 318 4;
  • 49) 0.000 015 258 789 062 500 331 661 860 169 318 4 × 2 = 0 + 0.000 030 517 578 125 000 663 323 720 338 636 8;
  • 50) 0.000 030 517 578 125 000 663 323 720 338 636 8 × 2 = 0 + 0.000 061 035 156 250 001 326 647 440 677 273 6;
  • 51) 0.000 061 035 156 250 001 326 647 440 677 273 6 × 2 = 0 + 0.000 122 070 312 500 002 653 294 881 354 547 2;
  • 52) 0.000 122 070 312 500 002 653 294 881 354 547 2 × 2 = 0 + 0.000 244 140 625 000 005 306 589 762 709 094 4;
  • 53) 0.000 244 140 625 000 005 306 589 762 709 094 4 × 2 = 0 + 0.000 488 281 250 000 010 613 179 525 418 188 8;
  • 54) 0.000 488 281 250 000 010 613 179 525 418 188 8 × 2 = 0 + 0.000 976 562 500 000 021 226 359 050 836 377 6;
  • 55) 0.000 976 562 500 000 021 226 359 050 836 377 6 × 2 = 0 + 0.001 953 125 000 000 042 452 718 101 672 755 2;
  • 56) 0.001 953 125 000 000 042 452 718 101 672 755 2 × 2 = 0 + 0.003 906 250 000 000 084 905 436 203 345 510 4;
  • 57) 0.003 906 250 000 000 084 905 436 203 345 510 4 × 2 = 0 + 0.007 812 500 000 000 169 810 872 406 691 020 8;
  • 58) 0.007 812 500 000 000 169 810 872 406 691 020 8 × 2 = 0 + 0.015 625 000 000 000 339 621 744 813 382 041 6;
  • 59) 0.015 625 000 000 000 339 621 744 813 382 041 6 × 2 = 0 + 0.031 250 000 000 000 679 243 489 626 764 083 2;
  • 60) 0.031 250 000 000 000 679 243 489 626 764 083 2 × 2 = 0 + 0.062 500 000 000 001 358 486 979 253 528 166 4;
  • 61) 0.062 500 000 000 001 358 486 979 253 528 166 4 × 2 = 0 + 0.125 000 000 000 002 716 973 958 507 056 332 8;
  • 62) 0.125 000 000 000 002 716 973 958 507 056 332 8 × 2 = 0 + 0.250 000 000 000 005 433 947 917 014 112 665 6;
  • 63) 0.250 000 000 000 005 433 947 917 014 112 665 6 × 2 = 0 + 0.500 000 000 000 010 867 895 834 028 225 331 2;
  • 64) 0.500 000 000 000 010 867 895 834 028 225 331 2 × 2 = 1 + 0.000 000 000 000 021 735 791 668 056 450 662 4;
  • 65) 0.000 000 000 000 021 735 791 668 056 450 662 4 × 2 = 0 + 0.000 000 000 000 043 471 583 336 112 901 324 8;
  • 66) 0.000 000 000 000 043 471 583 336 112 901 324 8 × 2 = 0 + 0.000 000 000 000 086 943 166 672 225 802 649 6;
  • 67) 0.000 000 000 000 086 943 166 672 225 802 649 6 × 2 = 0 + 0.000 000 000 000 173 886 333 344 451 605 299 2;
  • 68) 0.000 000 000 000 173 886 333 344 451 605 299 2 × 2 = 0 + 0.000 000 000 000 347 772 666 688 903 210 598 4;
  • 69) 0.000 000 000 000 347 772 666 688 903 210 598 4 × 2 = 0 + 0.000 000 000 000 695 545 333 377 806 421 196 8;
  • 70) 0.000 000 000 000 695 545 333 377 806 421 196 8 × 2 = 0 + 0.000 000 000 001 391 090 666 755 612 842 393 6;
  • 71) 0.000 000 000 001 391 090 666 755 612 842 393 6 × 2 = 0 + 0.000 000 000 002 782 181 333 511 225 684 787 2;
  • 72) 0.000 000 000 002 782 181 333 511 225 684 787 2 × 2 = 0 + 0.000 000 000 005 564 362 667 022 451 369 574 4;
  • 73) 0.000 000 000 005 564 362 667 022 451 369 574 4 × 2 = 0 + 0.000 000 000 011 128 725 334 044 902 739 148 8;
  • 74) 0.000 000 000 011 128 725 334 044 902 739 148 8 × 2 = 0 + 0.000 000 000 022 257 450 668 089 805 478 297 6;
  • 75) 0.000 000 000 022 257 450 668 089 805 478 297 6 × 2 = 0 + 0.000 000 000 044 514 901 336 179 610 956 595 2;
  • 76) 0.000 000 000 044 514 901 336 179 610 956 595 2 × 2 = 0 + 0.000 000 000 089 029 802 672 359 221 913 190 4;
  • 77) 0.000 000 000 089 029 802 672 359 221 913 190 4 × 2 = 0 + 0.000 000 000 178 059 605 344 718 443 826 380 8;
  • 78) 0.000 000 000 178 059 605 344 718 443 826 380 8 × 2 = 0 + 0.000 000 000 356 119 210 689 436 887 652 761 6;
  • 79) 0.000 000 000 356 119 210 689 436 887 652 761 6 × 2 = 0 + 0.000 000 000 712 238 421 378 873 775 305 523 2;
  • 80) 0.000 000 000 712 238 421 378 873 775 305 523 2 × 2 = 0 + 0.000 000 001 424 476 842 757 747 550 611 046 4;
  • 81) 0.000 000 001 424 476 842 757 747 550 611 046 4 × 2 = 0 + 0.000 000 002 848 953 685 515 495 101 222 092 8;
  • 82) 0.000 000 002 848 953 685 515 495 101 222 092 8 × 2 = 0 + 0.000 000 005 697 907 371 030 990 202 444 185 6;
  • 83) 0.000 000 005 697 907 371 030 990 202 444 185 6 × 2 = 0 + 0.000 000 011 395 814 742 061 980 404 888 371 2;
  • 84) 0.000 000 011 395 814 742 061 980 404 888 371 2 × 2 = 0 + 0.000 000 022 791 629 484 123 960 809 776 742 4;
  • 85) 0.000 000 022 791 629 484 123 960 809 776 742 4 × 2 = 0 + 0.000 000 045 583 258 968 247 921 619 553 484 8;
  • 86) 0.000 000 045 583 258 968 247 921 619 553 484 8 × 2 = 0 + 0.000 000 091 166 517 936 495 843 239 106 969 6;
  • 87) 0.000 000 091 166 517 936 495 843 239 106 969 6 × 2 = 0 + 0.000 000 182 333 035 872 991 686 478 213 939 2;
  • 88) 0.000 000 182 333 035 872 991 686 478 213 939 2 × 2 = 0 + 0.000 000 364 666 071 745 983 372 956 427 878 4;
  • 89) 0.000 000 364 666 071 745 983 372 956 427 878 4 × 2 = 0 + 0.000 000 729 332 143 491 966 745 912 855 756 8;
  • 90) 0.000 000 729 332 143 491 966 745 912 855 756 8 × 2 = 0 + 0.000 001 458 664 286 983 933 491 825 711 513 6;
  • 91) 0.000 001 458 664 286 983 933 491 825 711 513 6 × 2 = 0 + 0.000 002 917 328 573 967 866 983 651 423 027 2;
  • 92) 0.000 002 917 328 573 967 866 983 651 423 027 2 × 2 = 0 + 0.000 005 834 657 147 935 733 967 302 846 054 4;
  • 93) 0.000 005 834 657 147 935 733 967 302 846 054 4 × 2 = 0 + 0.000 011 669 314 295 871 467 934 605 692 108 8;
  • 94) 0.000 011 669 314 295 871 467 934 605 692 108 8 × 2 = 0 + 0.000 023 338 628 591 742 935 869 211 384 217 6;
  • 95) 0.000 023 338 628 591 742 935 869 211 384 217 6 × 2 = 0 + 0.000 046 677 257 183 485 871 738 422 768 435 2;
  • 96) 0.000 046 677 257 183 485 871 738 422 768 435 2 × 2 = 0 + 0.000 093 354 514 366 971 743 476 845 536 870 4;
  • 97) 0.000 093 354 514 366 971 743 476 845 536 870 4 × 2 = 0 + 0.000 186 709 028 733 943 486 953 691 073 740 8;
  • 98) 0.000 186 709 028 733 943 486 953 691 073 740 8 × 2 = 0 + 0.000 373 418 057 467 886 973 907 382 147 481 6;
  • 99) 0.000 373 418 057 467 886 973 907 382 147 481 6 × 2 = 0 + 0.000 746 836 114 935 773 947 814 764 294 963 2;
  • 100) 0.000 746 836 114 935 773 947 814 764 294 963 2 × 2 = 0 + 0.001 493 672 229 871 547 895 629 528 589 926 4;
  • 101) 0.001 493 672 229 871 547 895 629 528 589 926 4 × 2 = 0 + 0.002 987 344 459 743 095 791 259 057 179 852 8;
  • 102) 0.002 987 344 459 743 095 791 259 057 179 852 8 × 2 = 0 + 0.005 974 688 919 486 191 582 518 114 359 705 6;
  • 103) 0.005 974 688 919 486 191 582 518 114 359 705 6 × 2 = 0 + 0.011 949 377 838 972 383 165 036 228 719 411 2;
  • 104) 0.011 949 377 838 972 383 165 036 228 719 411 2 × 2 = 0 + 0.023 898 755 677 944 766 330 072 457 438 822 4;
  • 105) 0.023 898 755 677 944 766 330 072 457 438 822 4 × 2 = 0 + 0.047 797 511 355 889 532 660 144 914 877 644 8;
  • 106) 0.047 797 511 355 889 532 660 144 914 877 644 8 × 2 = 0 + 0.095 595 022 711 779 065 320 289 829 755 289 6;
  • 107) 0.095 595 022 711 779 065 320 289 829 755 289 6 × 2 = 0 + 0.191 190 045 423 558 130 640 579 659 510 579 2;
  • 108) 0.191 190 045 423 558 130 640 579 659 510 579 2 × 2 = 0 + 0.382 380 090 847 116 261 281 159 319 021 158 4;
  • 109) 0.382 380 090 847 116 261 281 159 319 021 158 4 × 2 = 0 + 0.764 760 181 694 232 522 562 318 638 042 316 8;
  • 110) 0.764 760 181 694 232 522 562 318 638 042 316 8 × 2 = 1 + 0.529 520 363 388 465 045 124 637 276 084 633 6;
  • 111) 0.529 520 363 388 465 045 124 637 276 084 633 6 × 2 = 1 + 0.059 040 726 776 930 090 249 274 552 169 267 2;
  • 112) 0.059 040 726 776 930 090 249 274 552 169 267 2 × 2 = 0 + 0.118 081 453 553 860 180 498 549 104 338 534 4;
  • 113) 0.118 081 453 553 860 180 498 549 104 338 534 4 × 2 = 0 + 0.236 162 907 107 720 360 997 098 208 677 068 8;
  • 114) 0.236 162 907 107 720 360 997 098 208 677 068 8 × 2 = 0 + 0.472 325 814 215 440 721 994 196 417 354 137 6;
  • 115) 0.472 325 814 215 440 721 994 196 417 354 137 6 × 2 = 0 + 0.944 651 628 430 881 443 988 392 834 708 275 2;
  • 116) 0.944 651 628 430 881 443 988 392 834 708 275 2 × 2 = 1 + 0.889 303 256 861 762 887 976 785 669 416 550 4;

We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit) and at least one integer that was different from zero => FULL STOP (Losing precision - the converted number we get in the end will be just a very good approximation of the initial one).


4. Construct the base 2 representation of the fractional part of the number.

Take all the integer parts of the multiplying operations, starting from the top of the constructed list above:


0.000 000 000 000 000 000 054 210 108 624 276 4(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001(2)

5. Positive number before normalization:

0.000 000 000 000 000 000 054 210 108 624 276 4(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001(2)

6. Normalize the binary representation of the number.

Shift the decimal mark 64 positions to the right, so that only one non zero digit remains to the left of it:


0.000 000 000 000 000 000 054 210 108 624 276 4(10) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001(2) =


0.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0001 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001(2) × 20 =


1.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001(2) × 2-64


7. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

Sign 0 (a positive number)


Exponent (unadjusted): -64


Mantissa (not normalized):
1.0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001


8. Adjust the exponent.

Use the 11 bit excess/bias notation:


Exponent (adjusted) =


Exponent (unadjusted) + 2(11-1) - 1 =


-64 + 2(11-1) - 1 =


(-64 + 1 023)(10) =


959(10)


9. Convert the adjusted exponent from the decimal (base 10) to 11 bit binary.

Use the same technique of repeatedly dividing by 2:


  • division = quotient + remainder;
  • 959 ÷ 2 = 479 + 1;
  • 479 ÷ 2 = 239 + 1;
  • 239 ÷ 2 = 119 + 1;
  • 119 ÷ 2 = 59 + 1;
  • 59 ÷ 2 = 29 + 1;
  • 29 ÷ 2 = 14 + 1;
  • 14 ÷ 2 = 7 + 0;
  • 7 ÷ 2 = 3 + 1;
  • 3 ÷ 2 = 1 + 1;
  • 1 ÷ 2 = 0 + 1;

10. Construct the base 2 representation of the adjusted exponent.

Take all the remainders starting from the bottom of the list constructed above.


Exponent (adjusted) =


959(10) =


011 1011 1111(2)


11. Normalize the mantissa.

a) Remove the leading (the leftmost) bit, since it's allways 1, and the decimal point, if the case.


b) Adjust its length to 52 bits, only if necessary (not the case here).


Mantissa (normalized) =


1. 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001 =


0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001


12. The three elements that make up the number's 64 bit double precision IEEE 754 binary floating point representation:

Sign (1 bit) =
0 (a positive number)


Exponent (11 bits) =
011 1011 1111


Mantissa (52 bits) =
0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001


Decimal number 0.000 000 000 000 000 000 054 210 108 624 276 4 converted to 64 bit double precision IEEE 754 binary floating point representation:

0 - 011 1011 1111 - 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0000 0110 0001


How to convert numbers from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point standard

Follow the steps below to convert a base 10 decimal number to 64 bit double precision IEEE 754 binary floating point:

  • 1. If the number to be converted is negative, start with its the positive version.
  • 2. First convert the integer part. Divide repeatedly by 2 the positive representation of the integer number that is to be converted to binary, until we get a quotient that is equal to zero, keeping track of each remainder.
  • 3. Construct the base 2 representation of the positive integer part of the number, by taking all the remainders from the previous operations, starting from the bottom of the list constructed above. Thus, the last remainder of the divisions becomes the first symbol (the leftmost) of the base two number, while the first remainder becomes the last symbol (the rightmost).
  • 4. Then convert the fractional part. Multiply the number repeatedly by 2, until we get a fractional part that is equal to zero, keeping track of each integer part of the results.
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the multiplying operations, starting from the top of the list constructed above (they should appear in the binary representation, from left to right, in the order they have been calculated).
  • 6. Normalize the binary representation of the number, shifting the decimal mark (the decimal point) "n" positions either to the left, or to the right, so that only one non zero digit remains to the left of the decimal mark.
  • 7. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary, by using the same technique of repeatedly dividing by 2, as shown above:
    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1
  • 8. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal mark, if the case) and adjust its length to 52 bits, either by removing the excess bits from the right (losing precision...) or by adding extra bits set on '0' to the right.
  • 9. Sign (it takes 1 bit) is either 1 for a negative or 0 for a positive number.

Example: convert the negative number -31.640 215 from the decimal system (base ten) to 64 bit double precision IEEE 754 binary floating point:

  • 1. Start with the positive version of the number:

    |-31.640 215| = 31.640 215

  • 2. First convert the integer part, 31. Divide it repeatedly by 2, keeping track of each remainder, until we get a quotient that is equal to zero:
    • division = quotient + remainder;
    • 31 ÷ 2 = 15 + 1;
    • 15 ÷ 2 = 7 + 1;
    • 7 ÷ 2 = 3 + 1;
    • 3 ÷ 2 = 1 + 1;
    • 1 ÷ 2 = 0 + 1;
    • We have encountered a quotient that is ZERO => FULL STOP
  • 3. Construct the base 2 representation of the integer part of the number by taking all the remainders of the previous dividing operations, starting from the bottom of the list constructed above:

    31(10) = 1 1111(2)

  • 4. Then, convert the fractional part, 0.640 215. Multiply repeatedly by 2, keeping track of each integer part of the results, until we get a fractional part that is equal to zero:
    • #) multiplying = integer + fractional part;
    • 1) 0.640 215 × 2 = 1 + 0.280 43;
    • 2) 0.280 43 × 2 = 0 + 0.560 86;
    • 3) 0.560 86 × 2 = 1 + 0.121 72;
    • 4) 0.121 72 × 2 = 0 + 0.243 44;
    • 5) 0.243 44 × 2 = 0 + 0.486 88;
    • 6) 0.486 88 × 2 = 0 + 0.973 76;
    • 7) 0.973 76 × 2 = 1 + 0.947 52;
    • 8) 0.947 52 × 2 = 1 + 0.895 04;
    • 9) 0.895 04 × 2 = 1 + 0.790 08;
    • 10) 0.790 08 × 2 = 1 + 0.580 16;
    • 11) 0.580 16 × 2 = 1 + 0.160 32;
    • 12) 0.160 32 × 2 = 0 + 0.320 64;
    • 13) 0.320 64 × 2 = 0 + 0.641 28;
    • 14) 0.641 28 × 2 = 1 + 0.282 56;
    • 15) 0.282 56 × 2 = 0 + 0.565 12;
    • 16) 0.565 12 × 2 = 1 + 0.130 24;
    • 17) 0.130 24 × 2 = 0 + 0.260 48;
    • 18) 0.260 48 × 2 = 0 + 0.520 96;
    • 19) 0.520 96 × 2 = 1 + 0.041 92;
    • 20) 0.041 92 × 2 = 0 + 0.083 84;
    • 21) 0.083 84 × 2 = 0 + 0.167 68;
    • 22) 0.167 68 × 2 = 0 + 0.335 36;
    • 23) 0.335 36 × 2 = 0 + 0.670 72;
    • 24) 0.670 72 × 2 = 1 + 0.341 44;
    • 25) 0.341 44 × 2 = 0 + 0.682 88;
    • 26) 0.682 88 × 2 = 1 + 0.365 76;
    • 27) 0.365 76 × 2 = 0 + 0.731 52;
    • 28) 0.731 52 × 2 = 1 + 0.463 04;
    • 29) 0.463 04 × 2 = 0 + 0.926 08;
    • 30) 0.926 08 × 2 = 1 + 0.852 16;
    • 31) 0.852 16 × 2 = 1 + 0.704 32;
    • 32) 0.704 32 × 2 = 1 + 0.408 64;
    • 33) 0.408 64 × 2 = 0 + 0.817 28;
    • 34) 0.817 28 × 2 = 1 + 0.634 56;
    • 35) 0.634 56 × 2 = 1 + 0.269 12;
    • 36) 0.269 12 × 2 = 0 + 0.538 24;
    • 37) 0.538 24 × 2 = 1 + 0.076 48;
    • 38) 0.076 48 × 2 = 0 + 0.152 96;
    • 39) 0.152 96 × 2 = 0 + 0.305 92;
    • 40) 0.305 92 × 2 = 0 + 0.611 84;
    • 41) 0.611 84 × 2 = 1 + 0.223 68;
    • 42) 0.223 68 × 2 = 0 + 0.447 36;
    • 43) 0.447 36 × 2 = 0 + 0.894 72;
    • 44) 0.894 72 × 2 = 1 + 0.789 44;
    • 45) 0.789 44 × 2 = 1 + 0.578 88;
    • 46) 0.578 88 × 2 = 1 + 0.157 76;
    • 47) 0.157 76 × 2 = 0 + 0.315 52;
    • 48) 0.315 52 × 2 = 0 + 0.631 04;
    • 49) 0.631 04 × 2 = 1 + 0.262 08;
    • 50) 0.262 08 × 2 = 0 + 0.524 16;
    • 51) 0.524 16 × 2 = 1 + 0.048 32;
    • 52) 0.048 32 × 2 = 0 + 0.096 64;
    • 53) 0.096 64 × 2 = 0 + 0.193 28;
    • We didn't get any fractional part that was equal to zero. But we had enough iterations (over Mantissa limit = 52) and at least one integer part that was different from zero => FULL STOP (losing precision...).
  • 5. Construct the base 2 representation of the fractional part of the number, by taking all the integer parts of the previous multiplying operations, starting from the top of the constructed list above:

    0.640 215(10) = 0.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 6. Summarizing - the positive number before normalization:

    31.640 215(10) = 1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2)

  • 7. Normalize the binary representation of the number, shifting the decimal mark 4 positions to the left so that only one non-zero digit stays to the left of the decimal mark:

    31.640 215(10) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) =
    1 1111.1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 20 =
    1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0(2) × 24

  • 8. Up to this moment, there are the following elements that would feed into the 64 bit double precision IEEE 754 binary floating point representation:

    Sign: 1 (a negative number)

    Exponent (unadjusted): 4

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

  • 9. Adjust the exponent in 11 bit excess/bias notation and then convert it from decimal (base 10) to 11 bit binary (base 2), by using the same technique of repeatedly dividing it by 2, as shown above:

    Exponent (adjusted) = Exponent (unadjusted) + 2(11-1) - 1 = (4 + 1023)(10) = 1027(10) =
    100 0000 0011(2)

  • 10. Normalize mantissa, remove the leading (leftmost) bit, since it's allways '1' (and the decimal sign) and adjust its length to 52 bits, by removing the excess bits, from the right (losing precision...):

    Mantissa (not-normalized): 1.1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100 1010 0

    Mantissa (normalized): 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Conclusion:

    Sign (1 bit) = 1 (a negative number)

    Exponent (8 bits) = 100 0000 0011

    Mantissa (52 bits) = 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100

  • Number -31.640 215, converted from decimal system (base 10) to 64 bit double precision IEEE 754 binary floating point =
    1 - 100 0000 0011 - 1111 1010 0011 1110 0101 0010 0001 0101 0111 0110 1000 1001 1100