how should cd (delete-next-char) and ch (delete-previous-char) behave on a unicode string containing diacritics?
imagine we are editing this string "abcབོད་ཀྱི་སྐད་ཡིག།中文".
cursor is at the end, after 文 character, and control-h is pressed repeatedly.
when deleting diacritics with ch, cursor should stay at the base character (same position). below are the lines demonstrating what should happen at each step:
abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中 abcབོད་ཀྱི་སྐད་ཡིག། abcབོད་ཀྱི་སྐད་ཡིག abcབོད་ཀྱི་སྐད་ཡི abcབོད་ཀྱི་སྐད་ཡ abcབོད་ཀྱི་སྐད་ abcབོད་ཀྱི་སྐད abcབོད་ཀྱི་སྐ abcབོད་ཀྱི་ས abcབོད་ཀྱི་ abcབོད་ཀྱི abcབོད་ཀྱ abcབོད་ཀ abcབོད་ abcབོད abcབོ abcབ abc ab a
however, cd (delete-next) should delete base char with its diacritics completely (until the next base character, i.e. the next non-zero cellwidth character).
demonstration below. cursor is at the beginning, control-d is pressed on each step:
abcབོད་ཀྱི་སྐད་ཡིག།中文 bcབོད་ཀྱི་སྐད་ཡིག།中文 cབོད་ཀྱི་སྐད་ཡིག།中文 བོད་ཀྱི་སྐད་ཡིག།中文 ད་ཀྱི་སྐད་ཡིག།中文 ་ཀྱི་སྐད་ཡིག།中文 ཀྱི་སྐད་ཡིག།中文 ་སྐད་ཡིག།中文 སྐད་ཡིག།中文 ་ཡིག།中文 ཡིག།中文 ག།中文 །中文 中文 文
abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文 abcབོད་ཀྱི་སྐད་ཡིག།中文
C:characters I:integers, `u?C returns unicode code points (indented lines are input, others are output):
`u?"abcབོད་ཀྱི་སྐད་ཡིག།中文" 97 98 99 3926 3964 3921 3851 3904 4017 3954 3851 3942 3984 3921 3851 3937 3954 3906 3853 20013 25991
`u@`u?"abcབོད་ཀྱི་སྐད་ཡིག།中文" abcབོད་ཀྱི་སྐད་ཡིག།中文
wc:0+`c$|/(2!0 32 127 160 768 880 1155 1162 1425 1470 1471 1472 1473 1475 1476 1478 1479 1480 1552 1563 1564 1565 1611 1632 1648 1649 1750 1757 1759 1765 1767 1769 1770 1774 1809 1810 1840 1867 1958 1969 2027 2036 2045 2046 2070 2074 2075 2084 2085 2088 2089 2094 2137 2140 2200 2208 2250 2274 2275 2307 2362 2363 2364 2365 2369 2377 2381 2382 2385 2392 2402 2404 2433 2434 2492 2493 2497 2501 2509 2510 2530 2532 2558 2559 2561 2563 2620 2621 2625 2627 2631 2633 2635 2638 2641 2642 2672 2674 2677 2678 2689 2691 2748 2749 2753 2758 2759 2761 2765 2766 2786 2788 2810 2816 2817 2818 2876 2877 2879 2880 2881 2885 2893 2894 2901 2903 2914 2916 2946 2947 3008 3009 3021 3022 3072 3073 3076 3077 3132 3133 3134 3137 3142 3145 3146 3150 3157 3159 3170 3172 3201 3202 3260 3261 3263 3264 3270 3271 3276 3278 3298 3300 3328 3330 3387 3389 3393 3397 3405 3406 3426 3428 3457 3458 3530 3531 3538 3541 3542 3543 3633 3634 3636 3643 3655 3663 3761 3762 3764 3773 3784 3791 3864 3866 3893 3894 3895 3896 3897 3898 3953 3967 3968 3973 3974 3976 3981 3992 3993 4029 4038 4039 4141 4145 4146 4152 4153 4155 4157 4159 4184 4186 4190 4193 4209 4213 4226 4227 4229 4231 4237 4238 4253 4254 4448 4608 4957 4960 5906 5909 5938 5940 5970 5972 6002 6004 6068 6070 6071 6078 6086 6087 6089 6100 6109 6110 6155 6160 6277 6279 6313 6314 6432 6435 6439 6441 6450 6451 6457 6460 6679 6681 6683 6684 6742 6743 6744 6751 6752 6753 6754 6755 6757 6765 6771 6781 6783 6784 6832 6863 6912 6916 6964 6965 6966 6971 6972 6973 6978 6979 7019 7028 7040 7042 7074 7078 7080 7082 7083 7086 7142 7143 7144 7146 7149 7150 7151 7154 7212 7220 7222 7224 7376 7379 7380 7393 7394 7401 7405 7406 7412 7413 7416 7418 7616 7680 8203 8208 8234 8239 8288 8293 8294 8304 8400 8433 11503 11506 11647 11648 11744 11776 12330 12334 12441 12443 42607 42611 42612 42622 42654 42656 42736 42738 43010 43011 43014 43015 43019 43020 43045 43047 43052 43053 43204 43206 43232 43250 43263 43264 43302 43310 43335 43346 43392 43395 43443 43444 43446 43450 43452 43454 43493 43494 43561 43567 43569 43571 43573 43575 43587 43588 43596 43597 43644 43645 43696 43697 43698 43701 43703 43705 43710 43712 43713 43714 43756 43758 43766 43767 44005 44006 44008 44009 44013 44014 55216 55239 55243 55292 64286 64287 65024 65040 65056 65072 65279 65280 65529 65532 66045 66046 66272 66273 66422 66427 68097 68100 68101 68103 68108 68112 68152 68155 68159 68160 68325 68327 68900 68904 69291 69293 69373 69376 69446 69457 69506 69510 69633 69634 69688 69703 69744 69745 69747 69749 69759 69762 69811 69815 69817 69819 69826 69827 69888 69891 69927 69932 69933 69941 70003 70004 70016 70018 70070 70079 70089 70093 70095 70096 70191 70194 70196 70197 70198 70200 70206 70207 70209 70210 70367 70368 70371 70379 70400 70402 70459 70461 70464 70465 70502 70509 70512 70517 70712 70720 70722 70725 70726 70727 70750 70751 70835 70841 70842 70843 70847 70849 70850 70852 71090 71094 71100 71102 71103 71105 71132 71134 71219 71227 71229 71230 71231 71233 71339 71340 71341 71342 71344 71350 71351 71352 71453 71456 71458 71462 71463 71468 71727 71736 71737 71739 71995 71997 71998 71999 72003 72004 72148 72152 72154 72156 72160 72161 72193 72203 72243 72249 72251 72255 72263 72264 72273 72279 72281 72284 72330 72343 72344 72346 72752 72759 72760 72766 72767 72768 72850 72872 72874 72881 72882 72884 72885 72887 73009 73015 73018 73019 73020 73022 73023 73030 73031 73032 73104 73106 73109 73110 73111 73112 73459 73461 73472 73474 73526 73531 73536 73537 73538 73539 78896 78913 78919 78934 92912 92917 92976 92983 94031 94032 94095 94099 94180 94181 113821 113823 113824 113828 118528 118574 118576 118599 119143 119146 119155 119171 119173 119180 119210 119214 119362 119365 121344 121399 121403 121453 121461 121462 121476 121477 121499 121504 121505 121520 122880 122887 122888 122905 122907 122914 122915 122917 122918 122923 123023 123024 123184 123191 123566 123567 123628 123632 124140 124144 125136 125143 125252 125259 917505 917506 917536 917632 917760 918000';2*~2!4352 4448 8986 8988 9001 9003 9193 9197 9200 9201 9203 9204 9725 9727 9748 9750 9800 9812 9855 9856 9875 9876 9889 9890 9898 9900 9917 9919 9924 9926 9934 9935 9940 9941 9962 9963 9970 9972 9973 9974 9978 9979 9981 9982 9989 9990 9994 9996 10024 10025 10060 10061 10062 10063 10067 10070 10071 10072 10133 10136 10160 10161 10175 10176 11035 11037 11088 11089 11093 11094 11904 11930 11931 12020 12032 12246 12272 12330 12334 12351 12353 12439 12443 12544 12549 12592 12593 12687 12688 12772 12783 12831 12832 42125 42128 42183 43360 43389 44032 55204 63744 64110 64112 64218 65040 65050 65072 65107 65108 65127 65128 65132 65281 65377 65504 65511 94176 94180 94192 94194 94208 100344 100352 101590 101632 101641 110576 110580 110581 110588 110589 110591 110592 110883 110898 110899 110928 110931 110933 110934 110948 110952 110960 111356 126980 126981 127183 127184 127374 127375 127377 127387 127488 127491 127504 127548 127552 127561 127568 127570 127584 127590 127744 127777 127789 127798 127799 127869 127870 127892 127904 127947 127951 127956 127968 127985 127988 127989 127992 128063 128064 128065 128066 128253 128255 128318 128331 128335 128336 128360 128378 128379 128405 128407 128420 128421 128507 128592 128640 128710 128716 128717 128720 128723 128725 128728 128732 128736 128747 128749 128756 128765 128992 129004 129008 129009 129292 129339 129340 129350 129351 129536 129648 129661 129664 129673 129680 129726 129727 129734 129742 129756 129760 129769 129776 129785 131072 173792 173824 177978 177984 178206 178208 183970 183984 191457 191472 192094 194560 195102 196608 201547 201552 205744')@\:!1114112 wc@`u?"abcབོད་ཀྱི་སྐད་ཡིག།中文" 111101110011011101122
0,+\wc@`u?"abcབོད་ཀྱི་སྐད་ཡིག།中文" 0 1 2 3 4 4 5 6 7 7 7 8 9 9 10 11 12 12 13 14 16 18
`u@I gives back the string from codepoints.
wc@I returns cellwidth-of-unicodecodepoint.
notice, wc returned 0 for diacritics and 2 for the last two hanzi.
cd (delete-next)x:3 /cursor position s:`u?"abcབོད་ཀྱི་སྐད་ཡིག།中文" s 97 98 99 3926 3964 3921 3851 3904 4017 3954 3851 3942 3984 3921 3851 3937 3954 3906 3853 20013 25991 #s 21 x<!#s 000011111111111111111 0<wc@s 111101110011011101111 (0<wc@s)&x<!#s 000001110011011101111 &(0<wc@s)&x<!#s 5 6 7 10 11 13 14 15 17 18 19 20 (#s),&(0<wc@s)&x<!#s 21 5 6 7 10 11 13 14 15 17 18 19 20 e:&/(#s),&(0<wc@s)&x<!#s e 5 (x#s),e_s 97 98 99 3921 3851 3904 4017 3954 3851 3942 3984 3921 3851 3937 3954 3906 3853 20013 25991 `u@(x#s),e_s abcད་ཀྱི་སྐད་ཡིག།中文 `u@s abcབོད་ཀྱི་སྐད་ཡིག།中文
"བོ" (3926 3964~`u?"བོ") code points are gone. yay!
ch (delete-prev)x:5 /cursor is right before "ད". ((x-1)#s),x_s abcབད་ཀྱི་སྐད་ཡིག།中文diacritic " ོ" is gone and base character is naked: "བ". yay!
cb (move back) and cf (move forward)
x:5 /cb should skip the 4 (diacritic) and be 3.
!#s
!21
x>!#s
111110000000000000000
wc@s
111101110011011101122
0<wc@s
111101110011011101111
(0<wc@s)&x>!#s
111100000000000000000
&(0<wc@s)&x>!#s
0123
x:|/0,&(0<wc@s)&x>!#s
x
3
and for cf we use min of count-s joinedwith where min-of-0-lessthan-widths and x-lessthan-til-count-s:
x:3 /cf should skip the 4 (diacritic) and be 5.
x<!#s
000000111111111111111
x:&/(#s),&(0<wc@s)&x<!#s
x
5
running it again should give 6 since that is the next base character (no diacritics):
x:&/(#s),&(0<wc@s)&x<!#s x 6
unicode is kinda messy.
the width of the characters are put here and there so we have to define large index arrays and search etc (i can't see a pattern, if you do please teach me, see footer).
programming string manipulation with unicode characters is not easy (for me).
also this complexity leads to buggy implementations.
for example rlwrap (version 0.48) does one extra deletion when we do the same ch demo.
also it does not handle individual deletion of diacritics (it deals base chars with their diacritics at all cases).
i thought about this but i am still not confident on how should the k primitives behave.
for example should the count-of-unicode-string return sum-of-cell-widths or just the count of codepoints.
in line editing, both counts are needed. maybe count should be codepoint and there should be a separate width counter? e.g. 3=#"བོད" and 2=`w?"བོད". but not sure how take, drop.. primitives should behave. again, for example truncating a string for visual purposes take and drop should act on cellwidth count but for other purposes it should be codepoints.. i still feel foggy on these questions.
please send mistakes if there are any to a at fall