data(hanja): a lone consonant is a reading too
Every Korean keyboard's 한자 key answers a lone consonant with the KS X 1001 symbol palette, and has since 한글 워드프로세서: ㅁ for ※ ○ △ ㈜, ㄴ for the brackets, ㄹ for the units, ㅇ for the circled numbers. strans sends that key to the same search as a syllable -- Khanja is Ctrl+H at strans.c:830, and startsearch seeds the query with whatever ko.c left pending -- but every one of hanja.dict's 187286 readings is a syllable, so the popup came up with a query in it and nothing to pick: ㅁ: 0 candidates ㄴ: 0 candidates ㄹ: 0 candidates 한: 99 candidates 韓 漢 寒 限 閑 恨 旱 汗 翰 邯 罕 悍 澣 閒 瀚 libhangul ships that palette beside the Hanja table already imported here: data/hanja/mssymbol.txt, same commit, same author, same BSD-3 terms, same key:value:comment format -- and keyed by the compatibility jamo ko.c already holds, U+3141 for ㅁ. So the engine does not change at all; the same dictlookup on the same trie now finds something: ㅁ: 75 candidates # & * @ § ※ ☆ ★ ○ ● ◎ ◇ ◆ □ ■ △ ▲ ▽ ▼ ㄴ: 23 candidates " ( ) [ ] { } ‘ ’ “ ” 〔 〕 〈 〉 《 》 「 」 ㄹ: 94 candidates $ % ₩ F ′ ″ ℃ Å ¢ £ ¥ ¤ ℉ ‰ € ㎕ ㎖ ㎗ ℓ 한: 99 candidates 韓 漢 寒 限 閑 恨 旱 汗 翰 邯 罕 悍 澣 閒 瀚 Both scripts widen by one rule -- a syllable reading gives Hanja, a jamo reading gives a symbol -- and hanja.src regenerates byte for byte as it was, because upstream's own non-syllable readings are words like ㄱ자집 whose values were never Hanja and still fall out. 985 of mssymbol.txt's 987 rows survive: its ideographic space and its soft hyphen do not, since a candidate the popup cannot draw is not a candidate, and the row format separates candidates with a space besides. The two keyspaces cannot collide -- one is syllables, one is single jamo -- so the 187286 existing rows are unchanged, byte for byte, and 18 rows join them. mkhanja takes a source list as mkemoji already does, and keeps each upstream header, which is why the licence text now appears twice. 89 unit, check-live, check-stress and valgrind all clean. The five new assertions were checked by breaking the change five ways: dropping mssymbol.src from SOURCES, letting issymbol keep a formatting character, letting a jamo reading keep Hanja, widening isjamo to the vowels, and making mkhanja reject jamo readings. Each fails only the tests that exist for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
49
map/README
49
map/README
@@ -1,42 +1,53 @@
|
||||
# Dictionary data
|
||||
|
||||
## Korean Hanja data
|
||||
## Korean Hanja and symbol data
|
||||
|
||||
`hanja.src` is the tracked, reviewable source for Korean Hanja conversion.
|
||||
Every data row is Hanja, a tab, and its modern Hangul reading, a character
|
||||
or a word:
|
||||
`hanja.src` and `mssymbol.src` are the tracked, reviewable sources for what
|
||||
the Hanja search converts. Every data row is a result, a tab, and its
|
||||
Hangul reading. A Hanja reads as a syllable or a word; a symbol reads as
|
||||
the lone consonant a Korean keyboard's 한자 key offers it under:
|
||||
|
||||
```
|
||||
漢 한
|
||||
漢字 한자
|
||||
※ ㅁ
|
||||
```
|
||||
|
||||
`mkhanja` validates that contract and groups rows by their reading to
|
||||
produce the runtime dictionary. Candidate order follows source order. The
|
||||
source keeps all retained pairs for review; the generated dictionary stores
|
||||
the first 128 candidates per reading because that is the engine's lookup
|
||||
limit.
|
||||
`mkhanja` validates that contract and groups the rows of both sources by
|
||||
their reading to produce the runtime dictionary. The two keyspaces do not
|
||||
meet: a reading is either syllables or one jamo. Candidate order follows
|
||||
source order. The sources keep all retained pairs for review; the
|
||||
generated dictionary stores the first 128 candidates per reading because
|
||||
that is the engine's lookup limit.
|
||||
|
||||
The source is derived from libhangul release tag `libhangul-0.2.0`. The
|
||||
The sources are derived from libhangul release tag `libhangul-0.2.0`. The
|
||||
annotated tag object is `20afc38922e3595ee3ed5b186f2ea05afe663763`, its
|
||||
peeled commit is `41c702f5d3581325b646ef6249f1f641b0427ae0`, and
|
||||
`data/hanja/hanja.txt` has blob ID
|
||||
`199cfd70c4b306257ac2e018714a1285f7cf0ed3`. Import and regenerate it from
|
||||
the repository root with:
|
||||
`data/hanja/hanja.txt` and `data/hanja/mssymbol.txt` have blob IDs
|
||||
`199cfd70c4b306257ac2e018714a1285f7cf0ed3` and
|
||||
`31c4e63d74293b6a759d4a2005cd1e7333746631`. Import and regenerate them
|
||||
from the repository root with:
|
||||
|
||||
```sh
|
||||
curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/hanja.txt -o upstream
|
||||
printf '%s %s\n' dd44dcc856cf542b1022d0f39c2e9b9f8805fdcc5923be80f04849ed97ce0996 upstream | sha256sum -c -
|
||||
map/libhangul2hanja upstream >map/hanja.src
|
||||
curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/mssymbol.txt -o upstream
|
||||
printf '%s %s\n' b685a4ebe2716b25eb29c42f6da6493716ecd78b604b44a33420987d21948e99 upstream | sha256sum -c -
|
||||
map/libhangul2hanja upstream >map/mssymbol.src
|
||||
map/mkhanja >map/hanja.dict
|
||||
```
|
||||
|
||||
`libhangul2hanja` retains only rows whose reading is modern Hangul syllables
|
||||
throughout and whose value is Hanja the popup can draw: U+3400–U+4DBF,
|
||||
U+4E00–U+9FFF, or U+F900–U+FAFF. Thus jamo readings, mixed values, and
|
||||
supplementary-plane ideographs are omitted. The import preserves the upstream license header and row order.
|
||||
The retained data is BSD 3-Clause licensed by Choe Hwanjin; the complete text
|
||||
is in `LICENSES/BSD-3-Clause-libhangul-hanja.txt`.
|
||||
`libhangul2hanja` keeps a syllable reading only when its value is Hanja the
|
||||
popup can draw: U+3400–U+4DBF, U+4E00–U+9FFF, or U+F900–U+FAFF. Thus mixed
|
||||
values and supplementary-plane ideographs are omitted. It keeps a jamo
|
||||
reading, U+3131–U+314E, only when its value is one rune the popup can draw
|
||||
and a candidate row can carry — so `mssymbol.txt`'s ideographic space and
|
||||
soft hyphen are omitted, leaving 985 of its 987 rows, and a Hanja under a
|
||||
jamo reading is omitted as noise. The import preserves each upstream
|
||||
license header and row order. The retained data is BSD 3-Clause licensed
|
||||
by Choe Hwanjin; the complete text is in
|
||||
`LICENSES/BSD-3-Clause-libhangul-hanja.txt`.
|
||||
|
||||
## Emoji and symbol data
|
||||
|
||||
|
||||
Reference in New Issue
Block a user